# 100 million free tokens to test Mercury 2.5

Sent September 8, 2026

Sponsored by Inception.

## Mercury 2.5 raises quality at a lower price

A coding agent repeats supporting steps throughout a session: search, route requests, compact context. The time and cost of those steps accumulate before the user sees the answer. Inception’s September 8 release puts Mercury 2.5 on that path. The company reports a 40% increase in intelligence over Mercury 2, while keeping the same low-latency, low-cost serving profile. Its stated output speed is 1,107 tokens per second.

[Explore Mercury 2.5 →](https://www.inceptionlabs.ai/models) (Sponsored)

## What the launch changes

- Output tokens per second: 1,107
- Context window in tokens: 260K
- Standard input / output price per million tokens: $0.20 / $0.75
- Launch input / output price per million tokens · 80% off: $0.04 / $0.15

Prices are per million tokens. Figures reported by Inception. Mercury 2.5 supports tunable reasoning, parallel tool calls, and schema-aligned JSON. Inception describes its quality as comparable to GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. That comparison is Inception’s assessment.

## The limit

These are existing Mercury production examples. The launch page does not identify them as Mercury 2.5 measurements. We would test the new model on one repeated step, keep the task fixed, and measure quality, time, and cost before changing a production route. The figures and customer results in this brief come from Inception’s launch post; we have not independently reproduced them.

## Three places to measure the wait

- Search: A search request can involve query rewriting, reranking, fact extraction, and answer checking. Inception says several leading search-infrastructure companies already run Mercury in production.
- Voice: Inception reports median model response latency close to 170 milliseconds on OpenCall’s live customer-call workload. That measures the model response, not the entire voice conversation.
- Coding: Augment Code uses Mercury for context compaction, model routing, and MCP tool search. Inception reports an 82% latency reduction for compaction, from roughly 150 seconds to 27 seconds, and a 90% cost reduction while maintaining quality. Tool-search summaries return in under a second.

## Give it one call to earn

Mercury models are available through the Inception API, Baseten, and OpenRouter. Inception’s launch offer includes 100 million free API tokens.

## Partner links

The links below are part of this sponsored edition.

- [Mercury Voice](https://www.inceptionlabs.ai/models): Preview. Mercury Voice is a dedicated diffusion LLM for voice agents, with reported time-to-first-token under 170 milliseconds.
- [Mercury Router](https://www.inceptionlabs.ai/models): Preview. Mercury Router uses a diffusion LLM to route prompts across open and closed models based on quality, speed, and cost.
- [Try Mercury 2.5 in chat →](https://chat.inceptionlabs.ai/): 
- [Start with the API docs →](https://docs.inceptionlabs.ai/get-started/get-started): 

Archive copy preserves the sponsor disclosure, provider attribution, limits, and pricing stated when this partner brief was sent.

Canonical page: https://benchlm.ai/newsletter/issues/2026-09-08-mercury-2-5-sponsored
