Skip to main content
BenchLM

Berkeley Function Calling Leaderboard v4 (BFCL v4)

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

Top models on BFCL v4 — September 27, 2026

As of September 27, 2026, BTL-3 leads the BFCL v4 leaderboard with 88.5% , followed by Atria Dawn Preview (77.0%) and Qwen3.7 Max (75.0%).

22 modelsAgentic3% of Agentic reference weightCurrentUpdated September 27, 2026

Leaderboard (22 models)

Score
1
BTL-3Bad Theory Labs · Open weight
88.5%
2
Atria Dawn PreviewShanghai Artificial Intelligence Laboratory · Open weight
77.0%
3
Qwen3.7 MaxAlibaba · Closed
75.0%
4
BTL-4Bad Theory Labs · Open weight
73.5%
5
Ling 3.0 FlashInclusionAI · Open weight
73.0%
6
Qwen3.7 PlusAlibaba · Closed
72.9%
7
Pokee-Isaac 28BPokee AI · Closed
70.9%
8
MiniCPM5-2BOpenBMB · Open weight
66.6%
9
Granite 4.2 30BIBM · Open weight
61.4%
10
LLaDA2.2-flashInclusionAI · Open weight
60.8%
11
LFM2.5-2.6BLiquidAI · Open weight
56.9%
12
Granite 4.2 3BIBM · Open weight
52.4%
13
Granite 4.2 8BIBM · Open weight
52.4%
14
LFM2.5-8B-A1BLiquidAI · Open weight
49.7%
15
LLaDA2.2-miniInclusionAI · Open weight
47.7%
16
Mellum2-12B-A2.5B-ThinkingJetBrains · Open weight
45.6%
17
Mellum2-12B-A2.5B-InstructJetBrains · Open weight
44.2%
18
ZAYA1-8BZyphra · Open weight
39.2%
19
LFM2.5-VL-3BLiquidAI · Open weight
32.5%
20
MiniCPM5-1BOpenBMB · Open weight
25.1%
21
LFM2.5-VL-450MLiquidAI · Open weight
21.1%
22
LFM2.5-230MLiquidAI · Open weight
21.0%

According to BenchLM.ai, BTL-3 leads the BFCL v4 benchmark with a score of 88.5%, followed by Atria Dawn Preview (77.0%) and Qwen3.7 Max (75.0%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

22 models have been evaluated on BFCL v4. The benchmark falls in the Agentic category. BenchAlign v5.7 gives BFCL v4 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About BFCL v4

Year

2026

Tasks

Function-calling tasks

Format

Tool invocation and schema evaluation

Difficulty

Advanced tool use

BenchLM stores BFCL v4 as a function-calling benchmark.

Freshness and provenance

Version

BFCL v4 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does BFCL v4 measure?

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

Which model scores highest on BFCL v4?

BTL-3 by Bad Theory Labs currently leads with a score of 88.5% on BFCL v4.

How many models are evaluated on BFCL v4?

22 AI models have been evaluated on BFCL v4 on BenchLM.

Last updated: September 27, 2026 · BenchLM version BFCL v4 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.