Capability
Unranked
field median 56.3
Not eligible for a public rank
Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.
Follow model changesReleased Sep 8, 2026 — see all recent releases
Data as of September 8, 2026 · How the score is built
Share or export
Instruction Following ranks #53. A well-rounded choice across a range of tasks.
5 published rows leave some tracked benchmark slots empty. Independent runtime speed has not been measured.
Each value carries a field reference instead of floating alone. Markers compare this model with the current ranked and priced catalog; they are not absolute quality thresholds.
Capability
Unranked
field median 56.3
Not eligible for a public rank
Price
$0.040input / $0.15 output
input median $1
cached $0.004 · blended $0.095
Speed
1107tok/s
No field comparison available
Provider-reported · NVIDIA GPUs · TTFT not published
Context
260Ktokens
field median 256,000
Maximum output length is tracked separately
Coverage is split by category so a strong number never hides a thin evidence base. Verified means the row is tied to a published source; provisional rows remain visible but separate.
Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.
Scores and ranks appear only where published evidence can be displayed. The table keeps the score, weight, cohort, and evidence state together.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticRank #91 of 152Percentile 40thWeight 22%2 benchmarksVerified | 44.5 | #91 of 152 | 40th | 22% | 2 benchmarks | Verified |
| CodingRank #73 of 151Percentile 52ndWeight 20%1 benchmarkVerified | 47.8 | #73 of 151 | 52nd | 20% | 1 benchmark | Verified |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeRank #94 of 182Percentile 49thWeight 12%1 benchmarkVerified | 48.3 | #94 of 182 | 49th | 12% | 1 benchmark | Verified |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingRank #53 of 121Percentile 57thWeight 5%1 benchmarkVerified | 80.1 | #53 of 121 | 57th | 5% | 1 benchmark | Verified |
Coding opens by default. The marker compares each value with the best source-verified result in the catalog; provisional leaders do not set the reference. Expand the remaining categories for every published row.
| Benchmark | Score | Versus best verified row | Gap | Weight | Evidence |
|---|---|---|---|---|---|
| SciCodeScientific Code Benchmark | Score38% | Versus best verified row Best verified: Sakana Fugu · 60.1% | Gap22.1 behind | WeightWeighted 10% | Provider exact |
| Benchmark | Score | Versus best verified row | Gap | Weight | Evidence |
|---|---|---|---|---|---|
| τ³-bench resultsτ³-Bench Tool-Agent-User Evaluation | Score96.0% | Versus best verified row Best verified: Mercury 2.5 · 96.0% | GapBest verified | WeightDisplay only | Provider exact |
| DeepSearchQA | Score34.0% | Versus best verified row Best verified: Claude Opus 5 · 95.0% | Gap61 behind | WeightDisplay only | Provider exact |
| Benchmark | Score | Versus best verified row | Gap | Weight | Evidence |
|---|---|---|---|---|---|
| GPQA-DGPQA Diamond | Score79.0% | Versus best verified row Best verified: GPT-6 Astra · 96.0% | Gap17 behind | WeightDisplay only | Provider exact |
| Benchmark | Score | Versus best verified row | Gap | Weight | Evidence |
|---|---|---|---|---|---|
| IFBenchInstruction Following Benchmark | Score77% | Versus best verified row Best verified: MAI-Thinking-1 · 85% | Gap8 behind | WeightWeighted 70% | Provider exact |
The sequence follows explicit supersedes links. A successor's displayed score stays at least 0.1 points above its predecessor; raw benchmark rows do not move. Scores and prices remain blank when the corresponding public row or first-party rate is unavailable.
Feb 24, 2026
Mercury 2Score 42.3 · $0.25 / $0.75
Sep 8, 2026 · you are here
Mercury 2.5Not publicly ranked · $0.04 / $0.15
Base entry
The visual layer above carries the decisions. These notes preserve the model, ranking, coverage, and family context behind the numbers.
We track Mercury 2.5, but the public leaderboard excludes this profile until enough non-generated benchmark coverage is available. Only published rows appear above.
Mercury 2.5 is a proprietary model with a 260K context window. It uses an explicit reasoning mode, which can improve complex problem solving while adding latency and token use.
Available through the Inception API as `mercury-2.5`, with 260K context, tunable reasoning, parallel tool calls, and schema-aligned JSON. The launch also lists Baseten and OpenRouter access.
Inception released Mercury 2.5 on September 8, 2026. The launch chart supplies provider-reported GPQA Diamond, IFBench, long-context reasoning, SciCode, DeepSearchQA at ten tool calls, Omniscience accuracy and non-hallucination, and Tau3Bench Telecom results. TerminalBench lacks a version; the GDPval bar uses an unexplained percentage scale. Neither is mapped. Benchmark values are separate from Mercury 2 and Mercury 2.5 Preview.
Mercury 2.5 sits in the Mercury 2.5 family with Mercury 2.5 Preview. Its explicit predecessor is Mercury 2. 5 of 427 tracked benchmark slots currently have displayable evidence. Missing categories stay blank.
Its strongest eligible category is Instruction Following at #53, while its lowest eligible position is Knowledge at #94. a well-rounded choice across a range of tasks.
The models page lists standard rates of $0.20 per million input tokens, $0.02 per million cached input tokens, and $0.75 per million output tokens. The 80% launch discount reduces them to $0.04, $0.004, and $0.15. Inception does not state an end date on that page.
The first-party quick start uses the API model ID mercury-2.5. Mercury 2.5 supports adjustable reasoning, parallel tool calls, and JSON constrained to a schema. Mercury 2 remains supported for existing customers.
Inception reports 79% on GPQA Diamond, 77% on IFBench, and 38% on SciCode. The chart also reports 34% on DeepSearchQA with a ten-tool-call budget. These values belong to Mercury 2.5 and are not copied to Mercury 2 or the earlier Preview record.
The chart’s 96% Tau3Bench result is labeled Telecom in its source note. It should not be compared directly with the older Mercury 2 Banking result. Its TerminalBench bar does not identify a benchmark version, so we leave that result outside the versioned ledger until the setup is documented.
Inception reports 1,107 output tokens per second on NVIDIA GPUs. That is a provider throughput claim, separate from the hosted runtime measurements. It does not establish time to first token or the full response latency of a voice application.
Mercury Voice and Mercury Router were announced as separate previews. Their latency and routing claims do not become Mercury 2.5 specifications.
Inception · Model release
Radar confirmed these at the source. Use Mercury 2.5 in your work? Explore Radar to follow supported changes and choose your alerts.
Related resources
Last updated September 8, 2026. Runtime fields remain blank until a sourced snapshot exists.
Get one weekly email when material rank, price, availability, or benchmark evidence changes are worth revisiting.
Read a sample issueJoin 2,000+ readers.