all the models — AI benchmark observatory
← Benchmarks

BFCL-V4

Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple, parallel, and nested function calls across diverse programming scenarios.

id bfcl-v4 · max 1 · 22 models reported

#ModelScore
1Atria Dawn Preview
Shanghai AI Laboratory · open
0.77
2Qwen3.7 Max
Alibaba Cloud / Qwen Team
0.75
3Ling 3.0 Flash
InclusionAI
0.73
4Qwen3.5-397B-A17B
Alibaba Cloud / Qwen Team · open
0.73
5Qwen3.7-Plus
Alibaba Cloud / Qwen Team
0.73
6Qwen3.5-122B-A10B
Alibaba Cloud / Qwen Team · open
0.72