Berkeley Function Calling Leaderboard v3 (BFCL v3)
We show this table for reference; we do not rank on it.
A function-calling benchmark for tool selection, schema adherence, and argument correctness, covering single-turn, parallel, irrelevance and multi-turn subsets.
Benchmark score on BFCL v3 — September 27, 2026
We compile the BFCL v3 rows from provider self-reports. Ternary Bonsai 2 27B leads the table at 74.9%. We do not use these results to rank models overall.
1 modelAgenticCurrentDisplay onlyUpdated September 27, 2026
Benchmark score table (1 model)
ScoreAbout BFCL v3
Year
2025
Tasks
Function-calling tasks
Format
Tool invocation and schema evaluation
Difficulty
Advanced tool use
BenchLM stores BFCL v3 as a display-only function-calling reference outside the current weighted core schema, and keeps it on its own key so a v3 result never overwrites the separate BFCL v4 lane. Providers differ on whether they run prompt-based or native function calling, so the row note should record which mode produced the value.
Freshness and provenance
Version
BFCL v3 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does BFCL v3 measure?
A function-calling benchmark for tool selection, schema adherence, and argument correctness, covering single-turn, parallel, irrelevance and multi-turn subsets.
Which model scores highest on BFCL v3?
Ternary Bonsai 2 27B by Prism ML currently leads with a score of 74.9% on BFCL v3.
How many models are evaluated on BFCL v3?
1 AI models have been evaluated on BFCL v3 on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.