# Gemini 3.1 Pro Benchmark Scores & Performance

> Gemini 3.1 Pro by Google scores 70.14/100 overall, ranking #16 out of 491 AI models.

Canonical page: https://benchlm.ai/models/gemini-3-1-pro

Last updated: September 18, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Google |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Overall Score | 70.14/100 |
| Overall Rank | #16 of 491 |

## Family & Coverage

- Family: Gemini 3.1 Pro
- Variant: base
- Benchmarks covered: 46 of 446
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Claw-Eval](/benchmarks/claw-eval) | 57.8% |
| [DeepSearchQA](/benchmarks/deepsearchqa) | 69.7% |
| [τ²-bench results](/benchmarks/tau2-bench) | 95.6% |
| [APEX-Agents-AA](/benchmarks/apexagentsaa) | 32.0% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 20.2% |
| [Gert Labs](/benchmarks/gertlabs) | 56.87% |
| [ResearchClawBench](/benchmarks/researchclawbench) | 13.3% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 70.8% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 10.3% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 904 |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [LiveCodeBench Pro](/benchmarks/livecodebench-pro) | 82.9% |
| [React Native Evals](/benchmarks/reactnativeevals) | 78.9% |
| [Vibe Code Bench](/benchmarks/vibecodebench) | 32.03% |
| [AA-SciCode](/benchmarks/aascicode) | 58.7% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 88.5% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 78.8% |
| [AA Coding Index](/benchmarks/aacodingindex) | 68.8% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMMU-Pro](/benchmarks/mmmu-pro) | 83.9% |
| [CharXiv](/benchmarks/charxiv) | 80.2% |
| [ERQA](/benchmarks/erqa) | 69.4% |
| [SimpleVQA](/benchmarks/simplevqa) | 72.4% |
| [ScreenSpot Pro](/benchmarks/screenspot-pro) | 84.4% |
| [ZeroBench](/benchmarks/zerobench) | 29.0% |
| [MedXpertQA (MM)](/benchmarks/medxpertqamm) | 81.3% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 82.4% |
| [Design Arena Website](/benchmarks/designarenawebsite) | 1265 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [ARC-AGI-2](/benchmarks/arc-agi-2) | 77.1% |
| [ARC-AGI-3](/benchmarks/arcagi3) | 0.4% |
| [AA-LCR](/benchmarks/lcr) | 82.0% |
| [CritPt](/benchmarks/critpt) | 17.7% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA-D](/benchmarks/gpqa-diamond) | 94.3% |
| [HLE w/o tools](/benchmarks/hlenotools) | 45.4% |
| [HealthBench Hard](/benchmarks/healthbench-hard) | 20.6% |
| [MedXpertQA (Text)](/benchmarks/medxpertqatext) | 71.5% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 30.4% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 94.1% |
| [AA-HLE](/benchmarks/aahle) | 47.0% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 31.9% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 54.9% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 50.9% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 95.5% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 91.0% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-IFBench](/benchmarks/aaifbench) | 77.1% |

## Multilingual Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA Global-MMLU-Lite](/benchmarks/aaglobalmmlulite) | 93.2% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [FrontierMath v2 (Tiers 1-3)](/benchmarks/frontiermathv2tiers13) | 36.900% |
| [FrontierMath v2 (Tier 4)](/benchmarks/frontiermathv2tier4) | 16.700% |

## Other Google Models

- [Gemini 3.8 Flash](/models/gemini-3-8-flash) - Score: 75.89
- [Gemini 3.7 Flash](/models/gemini-3-7-flash) - Score: 68.52
- [Gemini 3.6 Flash](/models/gemini-3-6-flash) - Score: 68.42
- [Gemini 3.5 Flash](/models/gemini-3-5-flash) - Score: 68.02
- [Gemini 3 Pro](/models/gemini-3-pro) - Score: 66.36
- [Gemini 3 Flash](/models/gemini-3-flash) - Score: 61.94
- [Gemini 3 Pro Deep Think](/models/gemini-3-pro-deep-think) - Score: 59.32
- [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) - Score: 58.78
- [Gemini 2.5 Pro](/models/gemini-2-5-pro) - Score: 57.31
- [Gemini 3.1 Flash-Lite](/models/gemini-3-1-flash-lite) - Score: 56.32
- [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) - Score: 55
- [Gemma 4 31B](/models/gemma-4-31b) - Score: 52.4
- [Gemini 2.5 Flash](/models/gemini-2-5-flash) - Score: 51.24
- [Gemma 4 12B](/models/gemma-4-12b) - Score: 42.99
- [Gemma 4 E4B](/models/gemma-4-e4b) - Score: 40.18
- [Gemma 4 E2B](/models/gemma-4-e2b) - Score: 39.58
- [Gemma 3 27B](/models/gemma-3-27b) - Score: 35.73
- [Gemini 1.5 Pro](/models/gemini-1-5-pro) - Score: 33.46
- [Gemini 1.0 Pro](/models/gemini-1-0-pro) - Score: 16.32
- [Gemini 3.8 Live](/models/gemini-3-8-live) - Score: not computed
- [Gemini 3.8 Live Extended Thinking](/models/gemini-3-8-live-extended-thinking) - Score: not computed
- [Gemini 3.8 Flash Cyber](/models/gemini-3-8-flash-cyber) - Score: not computed
- [Gemini 3.5 Transcribe](/models/gemini-3-5-transcribe) - Score: not computed
- [Gemini 3.5 Flash Cyber](/models/gemini-3-5-flash-cyber) - Score: not computed
- [Gemini 3.1 Flash TTS Preview](/models/gemini-3-1-flash-tts-preview) - Score: not computed
- [Gemini 2.5 Flash Native Audio Preview (12-2025)](/models/gemini-2-5-flash-native-audio-preview-12-2025) - Score: not computed
- [Gemini 2.5 Flash TTS Preview](/models/gemini-2-5-flash-tts-preview) - Score: not computed
- [Gemini 2.5 Pro TTS Preview](/models/gemini-2-5-pro-tts-preview) - Score: not computed
- [Gemini 3.1 Flash Live Preview](/models/gemini-3-1-flash-live-preview) - Score: not computed
