# Trinity-Large-Thinking Benchmark Scores & Performance

> Trinity-Large-Thinking by Arcee AI scores 34.23/100 overall, ranking #148 out of 514 AI models.

Canonical page: https://benchlm.ai/models/trinity-large-thinking

Last updated: September 29, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Arcee AI |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 512K |
| Overall Score | 34.23/100 |
| Overall Rank | #148 of 514 |

## Family & Coverage

- Family: Trinity Large
- Variant: thinking
- Benchmarks covered: 21 of 486
- Sibling models: [Trinity-Large-Preview](/models/trinity-large-preview)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [τ²-bench results](/benchmarks/tau2-bench) | 90.1% |
| [Gert Labs](/benchmarks/gertlabs) | 32.55% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 0.0% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 1.2% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 512 |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [SWE-bench Verified*](/benchmarks/sweverifiedarcee) | 63.2% |
| [AA-SciCode](/benchmarks/aascicode) | 40.6% |
| [AA Coding Index](/benchmarks/aacodingindex) | 25.8% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Design Arena Website](/benchmarks/designarenawebsite) | 1145 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 38.0% |
| [CritPt](/benchmarks/critpt) | 0.9% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA-D](/benchmarks/gpqa-diamond) | 76.3% |
| [MMLU-Pro (Arcee)](/benchmarks/mmluproarcee) | 83.4% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 10.8% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 75.2% |
| [AA-HLE](/benchmarks/aahle) | 15.8% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -44.1% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 22.5% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 85.9% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-IFBench](/benchmarks/aaifbench) | 56.3% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AIME25 (Arcee)](/benchmarks/aime2025arcee) | 96.3% |

## Other Arcee AI Models

- [Trinity-Large-Preview](/models/trinity-large-preview) - Score: not computed
