# GPT-5.6 Sol Benchmark Scores & Performance

> GPT-5.6 Sol by OpenAI scores 78.49/100 overall, ranking #7 out of 507 AI models.

Canonical page: https://benchlm.ai/models/gpt-5-6-sol

Last updated: September 24, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | OpenAI |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1.05M |
| Overall Score | 78.49/100 |
| Overall Rank | #7 of 507 |

## Family & Coverage

- Family: GPT-5.6
- Variant: sol (sol)
- Benchmarks covered: 74 of 483
- Sibling models: [GPT-5.6 Terra](/models/gpt-5-6-terra), [GPT-5.6 Luna](/models/gpt-5-6-luna), [GPT-5.6 Cyber](/models/gpt-5-6-cyber)
- Related earlier model: [GPT-5.5](/models/gpt-5-5)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 3.0](/benchmarks/terminal-bench-3) | 34.6% |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 91.9% |
| [BrowseComp](/benchmarks/browsecomp) | 92.2% |
| [OSWorld 2.0](/benchmarks/osworld2) | 62.6% |
| [CyberGym](/benchmarks/cybergym) | 84.5% |
| [ExploitGym](/benchmarks/exploitgym) | 33.7% |
| [Toolathlon](/benchmarks/toolathlon) | 58% |
| [τ²-bench results](/benchmarks/tau2-bench) | 85.1% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 54.4% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1735 |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1487 |
| [AA ITBench](/benchmarks/aaitbench) | 56.2% |
| [AA Tau3 Banking](/benchmarks/aatau3banking) | 44.3% |
| [AA AutomationBench](/benchmarks/aaautomationbench) | 60.1% |
| [AA Harvey LAB](/benchmarks/aaharveylab) | 87.2% |
| [terminalBenchHard](/benchmarks/terminal-bench-hard) | 65.9% |
| [AA EnterpriseOps-Gym](/benchmarks/aaenterpriseopsgym) | 42.9% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 85.8% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 50.5% |
| [GDP.pdf](/benchmarks/aagdppdf) | 27.2% |
| [AA-AnalystAgent](/benchmarks/aaanalystagent) | 47.5% |
| [ApprenticeBench](/benchmarks/apprenticebench) | 26% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Bug Hunt Bench](/benchmarks/bug-hunt-bench) | 42 fixes |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 64.6% |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 91.9% |
| [DeepSWE](/benchmarks/deepswe) | 72.7% |
| [FrontierCode 1.1 Extended](/benchmarks/frontiercode11extended) | 60.6% |
| [FrontierSWE v2](/benchmarks/frontierswev2) | 32.2% |
| [CursorBench 3.2](/benchmarks/cursorbench32) | 67.2% |
| [VulcanBench v3](/benchmarks/vulcanbench) | 87.0% |
| [AA-SciCode](/benchmarks/aascicode) | 57.1% |
| [VulcanBench CII v1](/benchmarks/vulcanciiv1) | 86.5% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 82.6% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 96.2% |
| [AA Coding Index](/benchmarks/aacodingindex) | 77.4% |
| [CursorBench 4.0](/benchmarks/cursorbench40) | 41.7% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMMU-Pro](/benchmarks/mmmu-pro) | 83% |
| [MMMU-Pro w/ Python](/benchmarks/mmmupropython) | 84.6% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 83.4% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [ARC-AGI-2](/benchmarks/arc-agi-2) | 92.5% |
| [ARC-AGI-3](/benchmarks/arcagi3) | 7.8% |
| [GeneBench-Pro](/benchmarks/genebenchpro) | 28.7% |
| [AA-LCR](/benchmarks/lcr) | 84.0% |
| [CritPt](/benchmarks/critpt) | 32.3% |
| [MLCR-AA](/benchmarks/aamlcr) | 26.1% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 94.6% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 94.6% |
| [HLE-Verified](/benchmarks/hleverified) | 54.5% |
| [LABBench2](/benchmarks/labbench2) | 82.1% |
| [HealthBench Professional](/benchmarks/healthbenchprofessional) | 60.5% |
| [HealthBench Hard](/benchmarks/healthbench-hard) | 33.1% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 58.9% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 94.1% |
| [AA-HLE](/benchmarks/aahle) | 49.5% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 22.0% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 59.4% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 92.2% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 95.2% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 89.1% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-IFBench](/benchmarks/aaifbench) | 72.7% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [FrontierMath (legacy)](/benchmarks/frontiermath) | 89% |
| [FrontierMath v2 (Tiers 1-3)](/benchmarks/frontiermathv2tiers13) | 89.000% |
| [FrontierMath v2 (Tier 4)](/benchmarks/frontiermathv2tier4) | 83.000% |

## Other OpenAI Models

- [GPT-6 Astra](/models/gpt-6-astra) - Score: 88.69
- [GPT-6 Sol](/models/gpt-6-sol) - Score: 81.28
- [GPT-5.6 Terra](/models/gpt-5-6-terra) - Score: 72.58
- [GPT-5.5 Pro](/models/gpt-5-5-pro) - Score: 72.42
- [GPT-5.4 Pro](/models/gpt-5-4-pro) - Score: 70.88
- [GPT-5.5](/models/gpt-5-5) - Score: 69.13
- [GPT-5.4](/models/gpt-5-4) - Score: 68.51
- [GPT-5.2 Pro](/models/gpt-5-2-pro) - Score: 66.71
- [GPT-6 Luna](/models/gpt-6-luna) - Score: 66.59
- [GPT-5.6 Luna](/models/gpt-5-6-luna) - Score: 65.6
- [GPT-5.3 Codex](/models/gpt-5-3-codex) - Score: 62.18
- [GPT-5.2](/models/gpt-5-2) - Score: 61.37
- [GPT-5.1](/models/gpt-5-1) - Score: 58.28
- [GPT-5.4 mini](/models/gpt-5-4-mini) - Score: 55.01
- [GPT-5.4 nano](/models/gpt-5-4-nano) - Score: 50.81
- [GPT-5.2-Codex](/models/gpt-5-2-codex) - Score: 50.15
- [o3-pro](/models/o3-pro) - Score: 47.25
- [o3](/models/o3) - Score: 47.04
- [GPT-5 (medium)](/models/gpt-5-medium) - Score: 44.92
- [GPT-5.1-Codex](/models/gpt-5-1-codex) - Score: 44.47
- [GPT-5 mini](/models/gpt-5-mini) - Score: 44.05
- [o3-mini](/models/o3-mini) - Score: 40.16
- [GPT-4.1](/models/gpt-4-1) - Score: 39.79
- [GPT-OSS 120B](/models/gpt-oss-120b) - Score: 37.76
- [o1](/models/o1) - Score: 37.58
- [o1-preview](/models/o1-preview) - Score: 36.32
- [o1-pro](/models/o1-pro) - Score: 35.65
- [GPT-5 nano](/models/gpt-5-nano) - Score: 35.01
- [GPT-OSS 20B](/models/gpt-oss-20b) - Score: 33.59
- [GPT-4o](/models/gpt-4o) - Score: 30.77
- [GPT-4.1 mini](/models/gpt-4-1-mini) - Score: 29.13
- [GPT-4o mini](/models/gpt-4o-mini) - Score: 26.88
- [GPT-4.1 nano](/models/gpt-4-1-nano) - Score: 24.87
- [GPT-4 Turbo](/models/gpt-4-turbo) - Score: 21.64
- [GPT-5.1-Codex-Max](/models/gpt-5-1-codex-max) - Score: not computed
- [o4-mini (high)](/models/o4-mini-high) - Score: not computed
- [GPT-5.2 Instant](/models/gpt-5-2-instant) - Score: not computed
- [GPT-5.3 Instant](/models/gpt-5-3-instant) - Score: not computed
- [GPT-5 (high)](/models/gpt-5-high) - Score: not computed
- [GPT-5.3-Codex-Spark](/models/gpt-5-3-codex-spark) - Score: not computed
- [GPT-Live-1](/models/gpt-live-1) - Score: not computed
- [GPT-5.6 Cyber](/models/gpt-5-6-cyber) - Score: not computed
- [GPT Live Transcribe](/models/gpt-live-transcribe) - Score: not computed
- [GPT Transcribe](/models/gpt-transcribe) - Score: not computed
- [GPT Realtime 2.1](/models/gpt-realtime-2-1) - Score: not computed
- [GPT Realtime 2.1 Mini](/models/gpt-realtime-2-1-mini) - Score: not computed
- [GPT Audio 1.5](/models/gpt-audio-1-5) - Score: not computed
- [GPT Realtime 1.5](/models/gpt-realtime-1-5) - Score: not computed
- [GPT-4o mini TTS](/models/gpt-4o-mini-tts) - Score: not computed
- [GPT Realtime](/models/gpt-realtime) - Score: not computed
- [GPT Realtime 2](/models/gpt-realtime-2) - Score: not computed
- [GPT Realtime mini](/models/gpt-realtime-mini) - Score: not computed
- [GPT-4o Audio](/models/gpt-4o-audio) - Score: not computed
- [GPT-4o mini Audio](/models/gpt-4o-mini-audio) - Score: not computed
