# Berkeley Function Calling Leaderboard v4 (BFCL v4)

> A function-calling benchmark for tool selection, schema adherence, and argument correctness.

Canonical page: https://benchlm.ai/benchmarks/bfcl-v4

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About BFCL v4

- Year: 2026
- Tasks: Function-calling tasks
- Format: Tool invocation and schema evaluation
- Difficulty: Advanced tool use
- Paper: [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking)

BenchLM stores BFCL v4 as a function-calling benchmark.

BenchAlign v5.7 gives BFCL v4 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Leaderboard (22 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [BTL-3](/models/btl-3) | Bad Theory Labs | 88.5% |
| 2 | [Atria Dawn Preview](/models/atria-dawn-preview) | Shanghai Artificial Intelligence Laboratory | 77.0% |
| 3 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 75.0% |
| 4 | [BTL-4](/models/btl-4) | Bad Theory Labs | 73.5% |
| 5 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 73.0% |
| 6 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 72.9% |
| 7 | [Pokee-Isaac 28B](/models/pokee-isaac-28b) | Pokee AI | 70.9% |
| 8 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 66.6% |
| 9 | [Granite 4.2 30B](/models/granite-4-2-30b) | IBM | 61.4% |
| 10 | [LLaDA2.2-flash](/models/llada2-2-flash) | InclusionAI | 60.8% |
| 11 | [LFM2.5-2.6B](/models/lfm2-5-2-6b) | LiquidAI | 56.9% |
| 12 | [Granite 4.2 3B](/models/granite-4-2-3b) | IBM | 52.4% |
| 13 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 52.4% |
| 14 | [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b) | LiquidAI | 49.7% |
| 15 | [LLaDA2.2-mini](/models/llada2-2-mini) | InclusionAI | 47.7% |
| 16 | [Mellum2-12B-A2.5B-Thinking](/models/mellum2-12b-a2-5b-thinking) | JetBrains | 45.6% |
| 17 | [Mellum2-12B-A2.5B-Instruct](/models/mellum2-12b-a2-5b-instruct) | JetBrains | 44.2% |
| 18 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 39.2% |
| 19 | [LFM2.5-VL-3B](/models/lfm2-5-vl-3b) | LiquidAI | 32.5% |
| 20 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 25.1% |
| 21 | [LFM2.5-VL-450M](/models/lfm2-5-vl-450m) | LiquidAI | 21.1% |
| 22 | [LFM2.5-230M](/models/lfm2-5-230m) | LiquidAI | 21.0% |

## FAQ

### What does BFCL v4 measure?

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

### Which model scores highest on BFCL v4?

BTL-3 by Bad Theory Labs currently leads with a score of 88.5% on BFCL v4.

### How many models are evaluated on BFCL v4?

22 AI models have been evaluated on BFCL v4 on BenchLM.

### Does BFCL v4 affect BenchLM's overall score?

Yes. BenchAlign v5.7 gives BFCL v4 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Compare Top Models on BFCL v4

- [BTL-3 vs Atria Dawn Preview](/compare/atria-dawn-preview-vs-btl-3)
- [Atria Dawn Preview vs Qwen3.7 Max](/compare/atria-dawn-preview-vs-qwen3-7-max)
- [Qwen3.7 Max vs BTL-4](/compare/btl-4-vs-qwen3-7-max)
- [BTL-4 vs Ling 3.0 Flash](/compare/btl-4-vs-ling-3-0-flash)
