Model profile
Gemini 3 Pro Deep Think
Evidence coverage
2 of 321 tracked benchmarks are published. 1 is verified and 1 provisional. 1 of 8 categories are measured.
- Published / tracked
- 2 / 321
- Verified
- 1
- Provisional
- 1
- Categories with evidence
- 1 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding0 benchmarksNot measured
- Reasoning2 benchmarksMixed evidence
- Knowledge0 benchmarksNot measured
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following0 benchmarksNot measured
Gemini 3 Pro Deep Think ranks #41 out of 200 models on the public leaderboard with an overall score of 61.31/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
Gemini 3 Pro Deep Think is a proprietary model with a 2M token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
Gemini 3 Pro Deep Think sits inside the Gemini 3 Pro family alongside Gemini 3 Pro. This profile currently has 2 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 60.56–61.31
- Gemini 3 Pro Deep ThinkCurrent modelGoogle#4161.31Gemini 3 Pro Deep Think is #41 with a score of 61.31.
- GLM-4.7Z.AICompare#4261.16GLM-4.7 is #42 with a score of 61.16.
- Gemma 4 31BGoogleCompare#4361.08Gemma 4 31B is #43 with a score of 61.08.
- GPT-5.4 ProOpenAICompare#4460.89GPT-5.4 Pro is #44 with a score of 60.89.
- Qwen3.5-27BAlibabaCompare#4560.7Qwen3.5-27B is #45 with a score of 60.7.
- DeepSeek V4 ProDeepSeekCompare#4660.66DeepSeek V4 Pro is #46 with a score of 60.66.
- Qwen3.5-122B-A10BAlibabaCompare#4760.56Qwen3.5-122B-A10B is #47 with a score of 60.56.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingWeight 20%0 benchmarksNot measured | Not measured | Not ranked | Not available | 20% | 0 benchmarks | Not measured |
| ReasoningRank Not rankedWeight 17%2 benchmarksMixed sources | 58.3 | Not ranked | Not available | 17% | 2 benchmarks | Mixed sources |
| KnowledgeWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1486 | ±3.9 | 41,631 |
| Coding | 1519 | ±7.2 | 8,622 |
| Math | 1479 | ±11.6 | 2,680 |
| Instruction Following | 1474 | ±6.5 | 11,259 |
| Creative Writing | 1485 | ±8.4 | 6,359 |
| Multi-turn | 1495 | ±8.0 | 6,774 |
| Hard Prompts | 1504 | ±5.1 | 22,573 |
| Hard Prompts (English) | 1503 | ±6.6 | 10,776 |
| Longer Query | 1492 | ±6.8 | 10,592 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Reasoning2 benchmarks
Abstraction and Reasoning Corpus for AGI v2
Critical Physics Tasks
Frequently Asked Questions
How does Gemini 3 Pro Deep Think perform overall in AI benchmarks?
Gemini 3 Pro Deep Think has 2 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is Gemini 3 Pro Deep Think good for reasoning and logic?
Gemini 3 Pro Deep Think has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.
Which sibling models are related to Gemini 3 Pro Deep Think?
Gemini 3 Pro Deep Think belongs to the Gemini 3 Pro family. Related variants on BenchLM include Gemini 3 Pro.
Does Gemini 3 Pro Deep Think have full benchmark coverage on BenchLM?
Not yet. Gemini 3 Pro Deep Think currently has 2 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of Gemini 3 Pro Deep Think?
Gemini 3 Pro Deep Think has a published context window of 2M, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.