# Artificial Analysis ITBench-AA (AA ITBench)

> An independently evaluated IT-operations benchmark from Artificial Analysis.

Canonical page: https://benchlm.ai/benchmarks/aaitbench

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About AA ITBench

- Year: 2026
- Tasks: IT incident-response tasks
- Format: Task success rate
- Difficulty: Enterprise IT operations
- Paper: [Artificial Analysis ITBench-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/itbench-aa)

Stored as a display-only agentic row.

AA ITBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (8 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 56.2% |
| 2 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 51.0% |
| 3 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 47.7% |
| 4 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 46.7% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 45.8% |
| 6 | [GLM-5.2](/models/glm-5-2) | Z.AI | 42.7% |
| 7 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 42.5% |
| 8 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 40.3% |

## FAQ

### What does AA ITBench measure?

An independently evaluated IT-operations benchmark from Artificial Analysis.

### Which model scores highest on AA ITBench?

GPT-5.6 Sol by OpenAI currently leads with a score of 56.2% on AA ITBench.

### How many models are evaluated on AA ITBench?

8 AI models have been evaluated on AA ITBench on BenchLM.

### Does AA ITBench affect BenchLM's overall score?

Not directly. AA ITBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on AA ITBench

- [GPT-5.6 Sol vs GPT-5.6 Terra](/compare/gpt-5-6-sol-vs-gpt-5-6-terra)
- [GPT-5.6 Terra vs Kimi K3](/compare/gpt-5-6-terra-vs-kimi-k3)
- [Kimi K3 vs Claude Opus 4.7 (Adaptive)](/compare/claude-opus-4-7-adaptive-vs-kimi-k3)
- [Claude Opus 4.7 (Adaptive) vs GPT-5.5](/compare/claude-opus-4-7-adaptive-vs-gpt-5-5)
