Skip to main content

Model comparison

Ornith-1.0-397B vs ZAYA1-74B-Preview

Data verified

Head-to-head evidence from 1 shared benchmark result across 1 category. Overall scores shown here use the public BenchAlign v5 ranking lane.

DeepReinforce AI
N/A
No comparison
1 category wins0 category wins

Evidence parity. Ornith-1.0-397B and ZAYA1-74B-Preview share 1 comparable benchmark result. 1 of 8 categories are comparable. 6 results are unique to Ornith-1.0-397B; 6 to ZAYA1-74B-Preview.

Updated July 18, 2026
Shared results
1
Ornith-1.0-397B only
6
ZAYA1-74B-Preview only
6
Comparable categories
1 / 8

Treat this as a split decision. Ornith-1.0-397B makes more sense if coding is the priority; ZAYA1-74B-Preview is the better fit if its strengths line up with your actual workload.

Confidence note. This is a partial-evidence comparison with 1 shared benchmark result across 1 evidence category; 1 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.

Why this result

Ornith-1.0-397B and ZAYA1-74B-Preview finish on the same BenchAlign overall score, so this is less about a single winner and more about where the edge shows up. The BenchAlign headline says tie; the benchmark table is where the real choice happens.

Category breakdown

Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.

Category scores and score margins for Ornith-1.0-397B and ZAYA1-74B-Preview
CategoryOrnith-1.0-397BΔZAYA1-74B-Preview
CodingOrnith-1.0-397B74.6Margin 21.4ZAYA1-74B-Preview53.2
AgenticOrnith-1.0-397B77.5MarginNo overlapZAYA1-74B-PreviewNot measured
KnowledgeOrnith-1.0-397BNot measuredMarginNo overlapZAYA1-74B-Preview66.1
MathOrnith-1.0-397BNot measuredMarginNo overlapZAYA1-74B-Preview76.4

Decisive benchmark drivers

The largest measured benchmark gaps in this matchup, with exact reported values.

More
A · Ornith-1.0-397BB · ZAYA1-74B-Preview
  1. SWE-bench Verified

    Coding
    Source ↗
    A 82.4%B 53.2%
    Winner: Ornith-1.0-397BΔ 29.2
    SWE-bench Verified: Ornith-1.0-397B scored 82.4%; ZAYA1-74B-Preview scored 53.2%. Ornith-1.0-397B wins this benchmark.

Operational comparison

Runtime and commercial metrics are compared only when both models have a complete sourced value.

MetricOrnith-1.0-397BZAYA1-74B-PreviewComparison
Input / output priceUSD per 1M tokensOrnith-1.0-397B$0 input / $0 outputZAYA1-74B-Preview$0 input / $0 outputListed prices are equal.
Generation speedtokens per secondOrnith-1.0-397BNot availableZAYA1-74B-PreviewNot availableA complete speed comparison is not available.
First-answer latencyseconds to first tokenOrnith-1.0-397BNot availableZAYA1-74B-PreviewNot availableA complete latency comparison is not available.
Context windowmaximum listed tokensOrnith-1.0-397B256KZAYA1-74B-Preview256KListed context windows are equal.

Benchmark Deep Dive

Agentic
BenchmarkOrnith-1.0-397BZAYA1-74B-PreviewResult
Terminal-Bench 2.0Source 77.5%Not comparable
Claw-EvalSource 77.1%Not comparable
τ²-bench AirlineSource 56.1%Not comparable
CodingOrnith-1.0-397B wins
BenchmarkOrnith-1.0-397BZAYA1-74B-PreviewResult
SWE-bench VerifiedSource 82.4%53.2%Ornith-1.0-397B leads
SWE-bench ProSource 62.2%Not comparable
SWE MultilingualSource 78.9%Not comparable
NL2RepoSource 48.2%Not comparable
Terminal-Bench 2.0Source 77.5%Not comparable
LiveCodeBench v6Source 65.7%Not comparable
Knowledge
BenchmarkOrnith-1.0-397BZAYA1-74B-PreviewResult
MMLU-ProSource 68.1%Not comparable
GPQASource 57.3%Not comparable
GPQA-DSource 57.3%Not comparable
Math
BenchmarkOrnith-1.0-397BZAYA1-74B-PreviewResult
AIME26Source 76.4%Not comparable
Frequently Asked Questions (2)

Which is better, Ornith-1.0-397B or ZAYA1-74B-Preview?

Ornith-1.0-397B and ZAYA1-74B-Preview are tied on the BenchAlign overall score, so the right pick depends on which category matters most for your use case.

Which is better for coding, Ornith-1.0-397B or ZAYA1-74B-Preview?

Ornith-1.0-397B has the edge for coding in this comparison, averaging 74.6 versus 53.2. Inside this category, SWE-bench Verified is the benchmark that creates the most daylight between them.

Related Comparisons

Last updated: July 18, 2026

The AI models change fast. We track them for you.

A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.

Free. No spam. Unsubscribe anytime.