← Benchmarks
BFCL-V4
Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple, parallel, and nested function calls across diverse programming scenarios.
id bfcl-v4 · max 1 · 22 models reported
| # | Model | Score |
|---|---|---|
| 1 | Atria Dawn Preview Shanghai AI Laboratory · open | 0.77 |
| 2 | Qwen3.7 Max Alibaba Cloud / Qwen Team | 0.75 |
| 3 | Ling 3.0 Flash InclusionAI | 0.73 |
| 4 | Qwen3.5-397B-A17B Alibaba Cloud / Qwen Team · open | 0.73 |
| 5 | Qwen3.7-Plus Alibaba Cloud / Qwen Team | 0.73 |
| 6 | Qwen3.5-122B-A10B Alibaba Cloud / Qwen Team · open | 0.72 |
