Model profile
GPT-4 Turbo
Evidence coverage
4 of 321 tracked benchmarks are published. 0 are verified and 4 provisional. 2 of 8 categories are measured.
- Published / tracked
- 4 / 321
- Verified
- 0
- Provisional
- 4
- Categories with evidence
- 2 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding2 benchmarksReported
- Reasoning0 benchmarksNot measured
- Knowledge2 benchmarksReported
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following0 benchmarksNot measured
GPT-4 Turbo ranks #188 out of 200 models on the public leaderboard with an overall score of 27.44/100. It also ranks #87 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
GPT-4 Turbo is a proprietary model with a 128K token context window. It processes queries without explicit chain-of-thought reasoning, offering faster response times and lower token usage.
This profile currently has 4 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Coding (#118). This performance profile makes it particularly well-suited for software development and code generation tasks.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 21.39–27.44
- GPT-4 TurboCurrent modelOpenAI#18827.44GPT-4 Turbo is #188 with a score of 27.44.
- Kimi K2Moonshot AICompare#18927.19Kimi K2 is #189 with a score of 27.19.
- MiniMax M1 80kMiniMaxCompare#19025.12MiniMax M1 80k is #190 with a score of 25.12.
- Llama 4 MaverickMetaCompare#19123.49Llama 4 Maverick is #191 with a score of 23.49.
- Phi-4MicrosoftCompare#19222.69Phi-4 is #192 with a score of 22.69.
- Gemini 1.0 ProGoogleCompare#19321.78Gemini 1.0 Pro is #193 with a score of 21.78.
- Claude 3 HaikuAnthropicCompare#19421.39Claude 3 Haiku is #194 with a score of 21.39.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Coding3%Eligible cohort rank #118 of 122Category score 30.7
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingRank #118 of 122Percentile 3rdWeight 20%2 benchmarksReported | 30.7 | #118 of 122 | 3rd | 20% | 2 benchmarks | Reported |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeRank Not rankedWeight 12%2 benchmarksReported | 54.4 | Not ranked | Not available | 12% | 2 benchmarks | Reported |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1160 | Not available | Not available |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Coding2 benchmarks
Artificial Analysis Coding Index
Artificial Analysis SciCode
Knowledge2 benchmarks
Artificial Analysis Humanity's Last Exam
Frequently Asked Questions
How does GPT-4 Turbo perform overall in AI benchmarks?
GPT-4 Turbo has 4 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is GPT-4 Turbo good for knowledge and understanding?
GPT-4 Turbo has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Is GPT-4 Turbo good for coding and programming?
GPT-4 Turbo ranks #118 out of 122 models in coding and programming benchmarks with an average score of 30.7. There are stronger options in this category.
Does GPT-4 Turbo have full benchmark coverage on BenchLM?
Not yet. GPT-4 Turbo currently has 4 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of GPT-4 Turbo?
GPT-4 Turbo has a published context window of 128K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.