Model comparison
Nemotron 3 Ultra vs Ornith-1.0-35B
Head-to-head evidence from 4 shared benchmark results across 2 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.
Verified leaderboard positions: Nemotron 3 Ultra #24; Ornith-1.0-35B unranked
Evidence parity. Nemotron 3 Ultra and Ornith-1.0-35B share 4 comparable benchmark results. 2 of 8 categories are comparable. 36 results are unique to Nemotron 3 Ultra; 3 to Ornith-1.0-35B.
Updated July 16, 2026- Shared results
- 4
- Nemotron 3 Ultra only
- 36
- Ornith-1.0-35B only
- 3
- Comparable categories
- 2 / 8
Pick Ornith-1.0-35B if you want the stronger benchmark profile. Nemotron 3 Ultra only becomes the better choice if you need the larger 1M context window.
Confidence note. This is a partial-evidence comparison with 4 shared benchmark results across 2 evidence categories; 2 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
Ornith-1.0-35B finishes one point ahead on BenchLM's provisional leaderboard, 62 to 61. That is enough to call, but not enough to treat as a blowout. This matchup comes down to a few meaningful edges rather than one model dominating the board.
Ornith-1.0-35B's sharpest advantage is in agentic, where it averages 64.2 against 51.3. The single biggest benchmark swing on the page is Terminal-Bench 2.0, 56.4% to 64.2%.
Nemotron 3 Ultra gives you the larger context window at 1M, compared with 256K for Ornith-1.0-35B.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | Nemotron 3 Ultra | Δ | Ornith-1.0-35B |
|---|---|---|---|
| Agentic | Nemotron 3 Ultra51.3 | Margin→ 12.9 | Ornith-1.0-35B64.2 |
| Coding | Nemotron 3 Ultra58.3 | Margin→ 7.6 | Ornith-1.0-35B65.9 |
| Reasoning | Nemotron 3 Ultra61.9 | MarginNo overlap | Ornith-1.0-35BNot measured |
| Knowledge | Nemotron 3 Ultra54.2 | MarginNo overlap | Ornith-1.0-35BNot measured |
| Multilingual | Nemotron 3 Ultra83.0 | MarginNo overlap | Ornith-1.0-35BNot measured |
| Inst. Following | Nemotron 3 Ultra81.7 | MarginNo overlap | Ornith-1.0-35BNot measured |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
Terminal-Bench 2.0
AgenticA 56.4%B 64.2%Winner: Ornith-1.0-35BΔ 7.8Terminal-Bench 2.0: Nemotron 3 Ultra scored 56.4%; Ornith-1.0-35B scored 64.2%. Ornith-1.0-35B wins this benchmark. - Source ↗
SWE-bench Verified
CodingA 71.9%B 75.6%Winner: Ornith-1.0-35BΔ 3.7SWE-bench Verified: Nemotron 3 Ultra scored 71.9%; Ornith-1.0-35B scored 75.6%. Ornith-1.0-35B wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | Nemotron 3 Ultra | Ornith-1.0-35B | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | Nemotron 3 Ultra$0 input / $0 output | Ornith-1.0-35B$0 input / $0 output | Listed prices are equal. |
| Generation speedtokens per second | Nemotron 3 UltraNot available | Ornith-1.0-35BNot available | A complete speed comparison is not available. |
| First-answer latencyseconds to first token | Nemotron 3 UltraNot available | Ornith-1.0-35BNot available | A complete latency comparison is not available. |
| Context windowmaximum listed tokens | Nemotron 3 Ultra1M | Ornith-1.0-35B256K | Nemotron 3 Ultra lists the larger context window. |
Benchmark Deep Dive
AgenticOrnith-1.0-35B wins14 benchmarks
| Benchmark | Nemotron 3 Ultra | Ornith-1.0-35B | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 56.4% | 64.2% | Ornith-1.0-35B leads |
| PinchBenchSource | 90.0% | — | Not comparable |
| BrowseCompSource | 44.4% | — | Not comparable |
| τ³-bench resultsSource | 70.9% | — | Not comparable |
| GDPval-AASource | 33.2% | — | Not comparable |
| HLE w/ toolsSource | 37.4% | — | Not comparable |
| AA Agentic IndexSource | 27.4% | — | Not comparable |
| τ²-bench resultsSource | 83.3% | — | Not comparable |
| GDPval-AASource | 1164 | — | Not comparable |
| AA BriefcaseSource | 870 | — | Not comparable |
| AA EnterpriseOps-GymSource | 28.9% | — | Not comparable |
| AA Harvey LABSource | 3.3% | — | Not comparable |
| AA Tau3 BankingSource | 13.8% | — | Not comparable |
| Claw-EvalSource | — | 69.8% | Not comparable |
CodingOrnith-1.0-35B wins10 benchmarks
| Benchmark | Nemotron 3 Ultra | Ornith-1.0-35B | Result |
|---|---|---|---|
| SWE-bench VerifiedSource | 71.9% | 75.6% | Ornith-1.0-35B leads |
| SWE MultilingualSource | 67.7% | 69.3% | Ornith-1.0-35B leads |
| LiveCodeBench v6Source | 89.0% | — | Not comparable |
| SciCodeSource | 44.6% | — | Not comparable |
| Terminal-Bench 2.0Source | 56.4% | 64.2% | Ornith-1.0-35B leads |
| AA Coding IndexSource | 49.3% | — | Not comparable |
| Terminal-Bench HardSource | 36.4% | — | Not comparable |
| AA-SciCodeSource | 39.9% | — | Not comparable |
| SWE-bench ProSource | — | 50.4% | Not comparable |
| NL2RepoSource | — | 34.6% | Not comparable |
Reasoning3 benchmarks
Knowledge12 benchmarks
| Benchmark | Nemotron 3 Ultra | Ornith-1.0-35B | Result |
|---|---|---|---|
| GPQASource | 87% | — | Not comparable |
| GPQA-DSource | 87.0% | — | Not comparable |
| HLESource | 26.7% | — | Not comparable |
| HLE w/o toolsSource | 26.7% | — | Not comparable |
| MMLU-ProSource | 86.8% | — | Not comparable |
| AA-Omniscience AccuracySource | 21.6% | — | Not comparable |
| Artificial Analysis Intelligence IndexSource | 37.8% | — | Not comparable |
| AA-GPQA DiamondSource | 86.7% | — | Not comparable |
| AA-HLESource | 26.6% | — | Not comparable |
| AA-Omniscience IndexSource | -0.8% | — | Not comparable |
| AA-Omniscience Hallucination RateSource | 28.5% | — | Not comparable |
| AA Openness IndexSource | 83.3% | — | Not comparable |
Multilingual1 benchmarks
| Benchmark | Nemotron 3 Ultra | Ornith-1.0-35B | Result |
|---|---|---|---|
| MMLU-ProXSource | 83% | — | Not comparable |
Multimodal1 benchmarks
| Benchmark | Nemotron 3 Ultra | Ornith-1.0-35B | Result |
|---|---|---|---|
| Design Arena WebsiteSource | 1132 | — | Not comparable |
Frequently Asked Questions (3)
Which is better, Nemotron 3 Ultra or Ornith-1.0-35B?
Ornith-1.0-35B is ahead on BenchLM's provisional leaderboard, 62 to 61. The biggest single separator in this matchup is Terminal-Bench 2.0, where the scores are 56.4% and 64.2%.
Which is better for coding, Nemotron 3 Ultra or Ornith-1.0-35B?
Ornith-1.0-35B has the edge for coding in this comparison, averaging 65.9 versus 58.3. Inside this category, Terminal-Bench 2.0 is the benchmark that creates the most daylight between them.
Which is better for agentic tasks, Nemotron 3 Ultra or Ornith-1.0-35B?
Ornith-1.0-35B has the edge for agentic tasks in this comparison, averaging 64.2 versus 51.3. Inside this category, Terminal-Bench 2.0 is the benchmark that creates the most daylight between them.
Related Comparisons
Explore More
The AI models change fast. We track them for you.
A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.
Free. No spam. Unsubscribe anytime.