Model comparison
Muse Spark vs Ornith-1.0-397B
Head-to-head evidence from 4 shared benchmark results across 2 categories. Overall scores shown here use BenchLM's provisional ranking lane.
Evidence parity. Muse Spark and Ornith-1.0-397B share 4 comparable benchmark results. 2 of 8 categories are comparable. 37 results are unique to Muse Spark; 3 to Ornith-1.0-397B.
Updated July 14, 2026- Shared results
- 4
- Muse Spark only
- 37
- Ornith-1.0-397B only
- 3
- Comparable categories
- 2 / 8
Pick Ornith-1.0-397B if you want the stronger benchmark profile. Muse Spark only becomes the better choice if you need the larger 262K context window.
Confidence note. This is a partial-evidence comparison with 4 shared benchmark results across 2 evidence categories; 2 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
Ornith-1.0-397B is clearly ahead on the provisional aggregate, 73 to 63. The gap is large enough that you do not need to squint at the spreadsheet to see the difference.
Ornith-1.0-397B's sharpest advantage is in agentic, where it averages 77.5 against 59. The single biggest benchmark swing on the page is Terminal-Bench 2.0, 59% to 77.5%.
Muse Spark gives you the larger context window at 262K, compared with 256K for Ornith-1.0-397B.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | Muse Spark | Δ | Ornith-1.0-397B |
|---|---|---|---|
| Agentic | Muse Spark59.0 | Margin→ 18.5 | Ornith-1.0-397B77.5 |
| Coding | Muse Spark61.7 | Margin→ 8.0 | Ornith-1.0-397B69.7 |
| Reasoning | Muse Spark42.5 | MarginNo overlap | Ornith-1.0-397BNot measured |
| Knowledge | Muse Spark50.4 | MarginNo overlap | Ornith-1.0-397BNot measured |
| Math | Muse Spark32.9 | MarginNo overlap | Ornith-1.0-397BNot measured |
| Multimodal | Muse Spark82.5 | MarginNo overlap | Ornith-1.0-397BNot measured |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
Terminal-Bench 2.0
AgenticA 59%B 77.5%Winner: Ornith-1.0-397BΔ 18.5Terminal-Bench 2.0: Muse Spark scored 59%; Ornith-1.0-397B scored 77.5%. Ornith-1.0-397B wins this benchmark. - Source ↗
SWE-bench Pro
CodingA 52.4%B 62.2%Winner: Ornith-1.0-397BΔ 9.8SWE-bench Pro: Muse Spark scored 52.4%; Ornith-1.0-397B scored 62.2%. Ornith-1.0-397B wins this benchmark. - Source ↗
SWE-bench Verified
CodingA 77.4%B 82.4%Winner: Ornith-1.0-397BΔ 5SWE-bench Verified: Muse Spark scored 77.4%; Ornith-1.0-397B scored 82.4%. Ornith-1.0-397B wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | Muse Spark | Ornith-1.0-397B | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | Muse SparkNot available | Ornith-1.0-397B$0 input / $0 output | A complete price comparison is not available. |
| Generation speedtokens per second | Muse SparkNot available | Ornith-1.0-397BNot available | A complete speed comparison is not available. |
| First-answer latencyseconds to first token | Muse SparkNot available | Ornith-1.0-397BNot available | A complete latency comparison is not available. |
| Context windowmaximum listed tokens | Muse Spark262K | Ornith-1.0-397B256K | Muse Spark lists the larger context window. |
Benchmark Deep Dive
AgenticOrnith-1.0-397B wins8 benchmarks
| Benchmark | Muse Spark | Ornith-1.0-397B | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 59% | 77.5% | Ornith-1.0-397B leads |
| Tau2-TelecomSource | 91.5% | — | Not comparable |
| DeepSearchQASource | 74.8% | — | Not comparable |
| CyberGymSource | 43.5% | — | Not comparable |
| Claw-EvalSource | 63.8% | 77.1% | Ornith-1.0-397B leads |
| AA Agentic IndexSource | 28.7% | — | Not comparable |
| GDPval-AASource | 32.2% | — | Not comparable |
| GDPval-AASource | 1144 | — | Not comparable |
CodingOrnith-1.0-397B wins10 benchmarks
| Benchmark | Muse Spark | Ornith-1.0-397B | Result |
|---|---|---|---|
| SWE-bench VerifiedSource | 77.4% | 82.4% | Ornith-1.0-397B leads |
| SWE-bench ProSource | 52.4% | 62.2% | Ornith-1.0-397B leads |
| LiveCodeBench ProSource | 80.0% | — | Not comparable |
| Vibe Code BenchSource | 19.67% | — | Not comparable |
| AA Coding IndexSource | 58.6% | — | Not comparable |
| Terminal-Bench HardSource | 45.5% | — | Not comparable |
| AA-SciCodeSource | 51.5% | — | Not comparable |
| SWE MultilingualSource | — | 78.9% | Not comparable |
| NL2RepoSource | — | 48.2% | Not comparable |
| Terminal-Bench 2.0Source | — | 77.5% | Not comparable |
Reasoning3 benchmarks
Knowledge11 benchmarks
| Benchmark | Muse Spark | Ornith-1.0-397B | Result |
|---|---|---|---|
| GPQA-DSource | 89.5% | — | Not comparable |
| HLESource | 50.4% | — | Not comparable |
| HLE w/o toolsSource | 42.8% | — | Not comparable |
| HealthBench HardSource | 42.8% | — | Not comparable |
| MedXpertQA (Text)Source | 52.6% | — | Not comparable |
| Artificial Analysis Intelligence IndexSource | 43.1% | — | Not comparable |
| AA-GPQA DiamondSource | 88.4% | — | Not comparable |
| AA-HLESource | 39.9% | — | Not comparable |
| AA-Omniscience IndexSource | 4.1% | — | Not comparable |
| AA-Omniscience AccuracySource | 44.6% | — | Not comparable |
| AA-Omniscience Hallucination RateSource | 73.2% | — | Not comparable |
Math2 benchmarks
Multimodal9 benchmarks
| Benchmark | Muse Spark | Ornith-1.0-397B | Result |
|---|---|---|---|
| CharXivSource | 86.4% | — | Not comparable |
| MMMU-ProSource | 80.4% | — | Not comparable |
| ERQASource | 64.7% | — | Not comparable |
| SimpleVQASource | 71.3% | — | Not comparable |
| ScreenSpot ProSource | 84.1% | — | Not comparable |
| ZeroBenchSource | 33.0% | — | Not comparable |
| MedXpertQA (MM)Source | 78.4% | — | Not comparable |
| GDPval-AASource | 1444 | — | Not comparable |
| AA-MMMU-ProSource | 80.5% | — | Not comparable |
Inst. Following1 benchmarks
| Benchmark | Muse Spark | Ornith-1.0-397B | Result |
|---|---|---|---|
| AA-IFBenchSource | 75.9% | — | Not comparable |
Frequently Asked Questions (3)
Which is better, Muse Spark or Ornith-1.0-397B?
Ornith-1.0-397B is ahead on BenchLM's provisional leaderboard, 73 to 63. The biggest single separator in this matchup is Terminal-Bench 2.0, where the scores are 59% and 77.5%.
Which is better for coding, Muse Spark or Ornith-1.0-397B?
Ornith-1.0-397B has the edge for coding in this comparison, averaging 69.7 versus 61.7. Inside this category, SWE-bench Pro is the benchmark that creates the most daylight between them.
Which is better for agentic tasks, Muse Spark or Ornith-1.0-397B?
Ornith-1.0-397B has the edge for agentic tasks in this comparison, averaging 77.5 versus 59. Inside this category, Terminal-Bench 2.0 is the benchmark that creates the most daylight between them.
Related Comparisons
Explore More
The AI models change fast. We track them for you.
A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.
Free. No spam. Unsubscribe anytime.