# TIR-Bench

> A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.

Canonical page: https://benchlm.ai/benchmarks/tirbench

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About TIR-Bench

- Year: 2026
- Tasks: Visual agent and interface reasoning
- Format: Screenshot-grounded task reasoning
- Difficulty: Computer-use visual reasoning
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

TIR-Bench appears in Qwen's launch tables as a visual-agent benchmark with separate submetrics. BenchLM tracks it as a display-only row while preserving the exact values published by providers.

TIR-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does TIR-Bench measure?

A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.

### Which model scores highest on TIR-Bench?

No models have been evaluated on TIR-Bench yet.

### How many models are evaluated on TIR-Bench?

0 AI models have been evaluated on TIR-Bench on BenchLM.

### Does TIR-Bench affect BenchLM's overall score?

Not directly. TIR-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
