# Berkeley Function Calling Leaderboard v3 (BFCL v3)

> A function-calling benchmark for tool selection, schema adherence, and argument correctness, covering single-turn, parallel, irrelevance and multi-turn subsets.

Canonical page: https://benchlm.ai/benchmarks/bfclv3

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About BFCL v3

- Year: 2025
- Tasks: Function-calling tasks
- Format: Tool invocation and schema evaluation
- Difficulty: Advanced tool use
- Paper: [The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models](https://gorilla.cs.berkeley.edu/leaderboard.html)

BenchLM stores BFCL v3 as a display-only function-calling reference outside the current weighted core schema, and keeps it on its own key so a v3 result never overwrites the separate BFCL v4 lane. Providers differ on whether they run prompt-based or native function calling, so the row note should record which mode produced the value.

BFCL v3 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Ternary Bonsai 2 27B](/models/ternary-bonsai-2-27b) | Prism ML | 74.9% |

## FAQ

### What does BFCL v3 measure?

A function-calling benchmark for tool selection, schema adherence, and argument correctness, covering single-turn, parallel, irrelevance and multi-turn subsets.

### Which model scores highest on BFCL v3?

Ternary Bonsai 2 27B by Prism ML currently leads with a score of 74.9% on BFCL v3.

### How many models are evaluated on BFCL v3?

1 AI models have been evaluated on BFCL v3 on BenchLM.

### Does BFCL v3 affect BenchLM's overall score?

Not directly. BFCL v3 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
