# Terminal-Bench Hard

> A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.

Canonical page: https://benchlm.ai/benchmarks/terminal-bench-hard

- Category: [Coding](/coding)
- Last updated: September 27, 2026

## About Terminal-Bench Hard

- Year: 2026
- Tasks: Agentic coding and terminal tasks
- Format: Task success rate
- Difficulty: Professional software engineering
- Paper: [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3)

BenchLM stores Terminal-Bench Hard separately from Terminal-Bench 2.0 because OpenRouter and Artificial Analysis publish it as a distinct benchmark card.

Terminal-Bench Hard is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does Terminal-Bench Hard measure?

A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.

### Which model scores highest on Terminal-Bench Hard?

No models have been evaluated on Terminal-Bench Hard yet.

### How many models are evaluated on Terminal-Bench Hard?

0 AI models have been evaluated on Terminal-Bench Hard on BenchLM.

### Does Terminal-Bench Hard affect BenchLM's overall score?

Not directly. Terminal-Bench Hard is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
