Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief
Provider pricing hub

Kimi API Pricing (August 2026)

Last synced 24 confirmed releases in the last 30 daysPreview Kimi monitoringSources attached

Moonshot's official Kimi K3 pricing page lists $3.00 cache-miss input, $0.30 cache-hit input, and $15.00 output per million tokens. Kimi K3 also raises the context ceiling to 1M. The previous flagship tier remains at $0.95/$4.00 (K2.6 and its coding sibling K2.7 Code) and a value tier at $0.60/$3.00 (K2.5). The older rows keep 256K context. The table below owns the current Kimi price comparison; each rate is tied to Moonshot's platform documentation rather than a third-party router.

Start with the Kimi K3 model profile for current benchmark coverage, then compare the family in the Chinese model rankings, the price-vs-performance view, or the head-to-head model comparison.

Every Kimi API price per 1M tokens

Kimi API prices per one million tokens, by model.
ModelInput $/MCached input $/MOutput $/MContext
Kimi K3$3$0.3$151.05M
Kimi 2.6$0.95$4256K
Kimi K2.7 Code$0.95$4256K
Kimi K2.5$0.6$3256K
Kimi K2$0.6$2.5128K

The calculator uses real-time rates. Moonshot's Batch API charges 60% of standard price for K2.7 Code, K2.6, and K2.5, but the official Batch pricing page does not list Kimi K3. Cache-hit input is $0.30/M for Kimi K3, $0.19/M for K2.7 Code, and $0.15/M for Kimi K2.

Estimate your monthly Kimi API bill

Estimated monthly Kimi API cost for the workload above, by model.
ModelEst. monthly cost
Kimi K3$60
Kimi 2.6$17.5
Kimi K2.7 Code$17.5
Kimi K2.5$12
Kimi K2$11

Estimates use Kimi's standard per-token rates from the table above. For token-level presets and cross-provider comparison, use the full LLM API pricing calculator.

How Moonshot prices the Kimi API

Kimi K3 is the new premium API tier at $3.00 input / $15.00 output per million tokens, with cache-hit input at $0.30. The previous flagship rate of $0.95/$4.00 still covers both K2.6 (the general model) and kimi-k2.7-code (the coding endpoint), so choosing the specialist costs nothing extra. K2.5 sits at $0.60/$3.00, and Moonshot explicitly documents that its thinking and non-thinking modes run under the same model at the same price — no separate reasoning surcharge to model.

The discount Moonshot publishes is cache-hit input: $0.30 per million tokens on Kimi K3, $0.19 on K2.7 Code, and $0.15 on the older Kimi K2. For coding agents that resend large repository context on every turn, that cache lane is where most of the real savings live. Moonshot's Batch pricing page also charges 60% of standard price for K2.7 Code, K2.6, and K2.5. Kimi K3 is not on the supported Batch model list, so the calculator above stays on comparable real-time rates instead of applying the discount to every row.

Kimi vs DeepSeek and OpenAI pricing

Against its closest rival, the price ordering is now one-directional: at $0.435/$0.87, DeepSeek V4 Pro is about 54% cheaper than K2.6 ($0.95/$4.00) on input and 78% cheaper on output. That does not settle quality, latency, or deployment fit. For pure volume work, DeepSeek V4 Flash ($0.14/$0.28) undercuts every Kimi tier by a wider margin.

Kimi K3's $15 output rate is 25% above GPT-5.6 Terra's $12 and half GPT-5.6 Sol's $30 rate. K2.6 remains the cheaper Kimi endpoint at $4 output. See the full DeepSeek API pricing and OpenAI API pricing tables, or line models up head-to-head on the compare hub.

Which Kimi model to use

Kimi K3 is the capability pick and the only Kimi row with a 1M-token context window. Its model profile carries the current overall, coding, and agentic evidence rather than freezing those scores on a pricing page. If your workload is coding agents specifically, K2.7 Code is the same $0.95/$4.00 with a published $0.19 cache-hit input rate; its profile keeps specialist benchmark coverage separate from overall ranking eligibility.

K2.5 ($0.60/$3.00) is the value tier with 256K context. The legacy Kimi K2($0.60/$2.50, 128K) remains the cheapest listed output lane in the family. Kimi K3's full weights are now published, but a 2.8T-parameter model is not a casual self-host. Read the Kimi K3 release review and run the self-host calculator before treating zero API spend as zero infrastructure cost.

Kimi API pricing FAQ

How much does the Kimi API cost?

Kimi K3 costs $3.00 input and $15.00 output per million tokens, with cache-hit input at $0.30. Kimi K2.6 and Kimi K2.7 Code both cost $0.95/$4.00, while Kimi K2.5 costs $0.60/$3.00. Batch jobs cost 60% of the standard rate for K2.7 Code, K2.6, and K2.5; Kimi K3 is not currently listed for Batch.

Is the Kimi K2 API free?

No — Moonshot's hosted Kimi API is pay-per-token. The genuinely free path is self-hosting: K2.6 is an open-weight release, so teams with their own GPUs can run it without per-token fees. Hardware, operations, and power still have a cost. Use the self-host calculator and open-source model rankings before choosing that route.

What is Moonshot AI's API pricing per million tokens?

Kimi K3 is $3 input and $15 output per million tokens. K2.6 and K2.7 Code cost $0.95/$4, and K2.5 costs $0.60/$3. Moonshot lists cache-hit input at $0.30 for Kimi K3 and $0.19 for K2.7 Code. Eligible Batch jobs on K2.7 Code, K2.6, and K2.5 cost 60% of the corresponding real-time rate.

How does Kimi pricing compare to DeepSeek?

DeepSeek V4 Pro ($0.435/$0.87) is about 54% cheaper than Kimi K2.6 ($0.95/$4.00) on input and 78% cheaper on output. Token price therefore favors DeepSeek at current direct rates, while model quality, latency, support, and deployment requirements can still favor Kimi. For high-volume pipelines DeepSeek V4 Flash ($0.14/$0.28) is cheaper again — see the DeepSeek pricing hub for the full table.

Keep comparing