Skip to main content
BenchLM

Scale Labs AudioMC – Text Output (AudioMC – Text Output)

We show this table for reference; we do not rank on it.

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Scale score on AudioMC – Text Output — September 29, 2026 snapshot

We mirror the published scale score view for AudioMC – Text Output. Inkling (Thinking) leads the public snapshot at 56.6%, followed by Inkling-small\n (54.9%) and gemini-3-pro-preview (Thinking) (54.6%). We do not use these results to rank models overall.

18 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 29, 2026 snapshot

Scale score table (18 models)

Score
1
Inkling (Thinking)Thinkingmachines
56.6%
2
Inkling-small\nThinkingmachines
54.9%
6
GPT Realtime 1.5OpenAI · Closed
29.9%
8
Gemini 2.5 FlashGoogle · Closed
26.1%
10
GPT RealtimeOpenAI · Closed
23.4%
13
GPT Realtime miniOpenAI · Closed
16.6%
14
Phi-4 Multimodal InstructMicrosoft · Open weight
15.5%
15
15.5%
17
13.7%
18
Qwen2.5-Omni 7BAlibaba · Open weight
11.9%

How to read this leaderboard

Compare the published configurations as complete evaluation systems. The source can combine a base model, agent scaffold, tools, budget, and inference setting in each result.

Operator receipt: 18 sourced rows are currently displayable on this page; the leading published row is Inkling (Thinking) at 56.6%.

Honest limit: This Scale table is display-only context, not benchmark provenance or a weighted model-only comparison.

How BenchLM shows AudioMC – Text Output

BenchLM mirrors 18 published rows from Scale Labs’ public AudioMC – Text Output leaderboard, captured on September 29, 2026 snapshot.

The table is display only. It is useful context for a published agent or model configuration, but it does not enter BenchLM’s overall or category rankings.

Snapshot

18 published rowsScale Labs sourceDisplay only

The published AudioMC – Text Output snapshot places Inkling (Thinking) first at 56.6%. The third row is 2.0 points behind. The broader top-10 range is 33.2 points, so the table still separates the published systems.

18 models have been evaluated on AudioMC – Text Output. The benchmark falls in the Multimodal & Grounded category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. AudioMC – Text Output is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AudioMC – Text Output

Year

2026

Tasks

18 published rows

Format

Published Scale leaderboard score

Difficulty

External agent and model evaluation

BenchLM mirrors 18 published rows from the AudioMC – Text Output public table captured on September 29, 2026 snapshot.

Freshness and provenance

Version

AudioMC – Text Output 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AudioMC – Text Output measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Which model leads the published AudioMC – Text Output snapshot?

Inkling (Thinking) currently leads the published AudioMC – Text Output snapshot with 56.6% scale score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on AudioMC – Text Output?

The September 29, 2026 snapshot snapshot contains 18 AI models.

Last updated: September 29, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.