Benchmark profile
InferenceEval
An internal scientific-inference evaluation reported in the Grok 4.6 model card.
Data verified 26 confirmed releases in the last 30 daysStart free briefBenchmark score on InferenceEval — August 12, 2026
BenchLM mirrors the published score view for InferenceEval. Grok 4.6 leads the public snapshot at 46.9% , followed by GPT-5.6 Sol (44.1%) and Grok 4.5 (41.3%). We do not use these results to rank models overall.
Grok 4.6
xAI
grok-4-6
GPT-5.6 Sol
OpenAI
gpt-5-6-sol
Grok 4.5
xAI
grok-4-5
Benchmark score table (6 models)
ScoreThe published InferenceEval snapshot places Grok 4.6 first at 46.9%. The third row is 5.6 points behind. The broader top-10 range is 8.0 points, so many of the published results sit in a relatively narrow band.
6 models have been evaluated on InferenceEval. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. InferenceEval is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About InferenceEval
Year
2026
Tasks
Internal scientific-inference tasks
Format
Accuracy
Difficulty
Frontier scientific reasoning
No matching public benchmark repository or result page was found. BenchLM keeps the card's exact comparison rows display-only.
BenchLM freshness & provenance
Version
InferenceEval 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does InferenceEval measure?
An internal scientific-inference evaluation reported in the Grok 4.6 model card.
Which model scores highest on InferenceEval?
Grok 4.6 by xAI currently leads with a score of 46.9% on InferenceEval.
How many models are evaluated on InferenceEval?
6 AI models have been evaluated on InferenceEval on BenchLM.
Compare Top Models on InferenceEval
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.