# Phi-4 Benchmark Scores & Performance

> Phi-4 by Microsoft scores 32.91/100 overall, ranking #209 out of 411 AI models.

Canonical page: https://benchlm.ai/models/phi-4

Last updated: September 4, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Microsoft |
| Source Type | Open Weight |
| Reasoning Type | Non-Reasoning |
| Context Window | 16K |
| Overall Score | 32.91/100 |
| Overall Rank | #209 of 411 |

## Family & Coverage

- Family: Phi-4
- Variant: base
- Benchmarks covered: 11 of 422
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [τ²-bench results](/benchmarks/tau2-bench) | 0% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-SciCode](/benchmarks/aascicode) | 26.0% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 0.0% |
| [CritPt](/benchmarks/critpt) | 0.0% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 4.5% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 57.5% |
| [AA-HLE](/benchmarks/aahle) | 3.8% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -55.7% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 14.1% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 81.2% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-IFBench](/benchmarks/aaifbench) | 23.5% |

## Other Microsoft Models

- [MAI-Thinking-1](/models/mai-thinking-1) - Score: 52.27
- [MAI-Voice-2](/models/mai-voice-2) - Score: not computed
- [Phi-4 Multimodal Instruct](/models/phi-4-multimodal-instruct) - Score: not computed
