Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
Weekly brief and archive

Sponsored partner brief

100 million free tokens to test Mercury 2.5

Sponsored by Inception.

Mercury 2.5 raises quality at a lower price

A coding agent repeats supporting steps throughout a session: search, route requests, compact context. The time and cost of those steps accumulate before the user sees the answer. Inception’s September 8 release puts Mercury 2.5 on that path. The company reports a 40% increase in intelligence over Mercury 2, while keeping the same low-latency, low-cost serving profile. Its stated output speed is 1,107 tokens per second.

Explore Mercury 2.5 →

What the launch changes

1,107
Output tokens per second
260K
Context window in tokens
$0.20 / $0.75
Standard input / output price per million tokens
$0.04 / $0.15
Launch input / output price per million tokens · 80% off

Prices are per million tokens. Figures reported by Inception. Mercury 2.5 supports tunable reasoning, parallel tool calls, and schema-aligned JSON. Inception describes its quality as comparable to GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. That comparison is Inception’s assessment.

The limit

These are existing Mercury production examples. The launch page does not identify them as Mercury 2.5 measurements. We would test the new model on one repeated step, keep the task fixed, and measure quality, time, and cost before changing a production route. The figures and customer results in this brief come from Inception’s launch post; we have not independently reproduced them.

Three places to measure the wait

Search

A search request can involve query rewriting, reranking, fact extraction, and answer checking. Inception says several leading search-infrastructure companies already run Mercury in production.

Voice

Inception reports median model response latency close to 170 milliseconds on OpenCall’s live customer-call workload. That measures the model response, not the entire voice conversation.

Coding

Augment Code uses Mercury for context compaction, model routing, and MCP tool search. Inception reports an 82% latency reduction for compaction, from roughly 150 seconds to 27 seconds, and a 90% cost reduction while maintaining quality. Tool-search summaries return in under a second.

Give it one call to earn

Mercury models are available through the Inception API, Baseten, and OpenRouter. Inception’s launch offer includes 100 million free API tokens.

Partner links

The links below are part of this sponsored edition.

Archive copy preserves the sponsor disclosure, provider attribution, limits, and pricing stated when this partner brief was sent.

Want the next issue?

The signup form and another real sample are on the weekly brief page.

Subscribe to the weekly brief