# τ²-Bench Airline Domain (τ²-bench Airline)

> τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.

Canonical page: https://benchlm.ai/benchmarks/tau2airline

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About τ²-bench Airline

- Year: 2025
- Tasks: Airline customer-service tasks
- Format: Domain success under a published trial policy
- Difficulty: Policy-constrained airline support workflows
- Paper: [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982)

This lane owns results explicitly labeled for the airline domain. It stays separate from telecom, retail, aggregates, the archived original TAU-bench release, and newer τ³-bench runs.

τ²-bench Airline is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [ZAYA1-74B-Preview](/models/zaya1-74b-preview) | Zyphra | 56.1% |

## FAQ

### What does τ²-bench Airline measure?

τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.

### Which model scores highest on τ²-bench Airline?

ZAYA1-74B-Preview by Zyphra currently leads with a score of 56.1% on τ²-bench Airline.

### How many models are evaluated on τ²-bench Airline?

1 AI models have been evaluated on τ²-bench Airline on BenchLM.

### Does τ²-bench Airline affect BenchLM's overall score?

Not directly. τ²-bench Airline is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
