# Toloka Arena

> An independent agentic-intelligence evaluation from Toloka using private simulated workflows and a pass^5 metric.

Canonical page: https://benchlm.ai/benchmarks/tolokaarena

- Category: [Agentic](/agentic)
- Last updated: June 3, 2026

## About Toloka Arena

- Year: 2026
- Tasks: Private simulated enterprise workflows
- Format: pass^5 arena score
- Difficulty: Agentic workflow reliability
- Paper: [Toloka Arena](https://toloka.ai/arena)

Toloka Arena evaluates agents on private simulated workflows with tools, databases, policies, and multi-turn tasks. BenchLM tracks it as source metadata only until a stable public leaderboard data feed is available.

Toloka Arena is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does Toloka Arena measure?

An independent agentic-intelligence evaluation from Toloka using private simulated workflows and a pass^5 metric.

### Which model leads the published Toloka Arena snapshot?

No models have been evaluated on Toloka Arena yet.

### How many models are evaluated on Toloka Arena?

The June 3, 2026 contains 0 AI models.

### Does Toloka Arena affect BenchLM's overall score?

Not directly. Toloka Arena is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
