all the models — AI benchmark observatory
← Benchmarks

BFCL-v3

Berkeley Function Calling Leaderboard v3 (BFCL-v3) is an advanced benchmark that evaluates large language models' function calling capabilities through multi-turn and multi-step interactions. It introduces extended conversational exchanges where models must retain contextual information across turns and execute multiple internal function calls for complex user requests. The benchmark includes 1000 test cases across domains like vehicle control, trading bots, travel booking, and file system management, using state-based evaluation to verify both system state changes and execution path correctness.

id bfcl-v3 · max 1 · 20 models reported

#ModelScore
1GLM-4.5
Zhipu AI · open
0.78
2GLM-4.5-Air
Zhipu AI · open
0.76
3LongCat-Flash-Thinking
Meituan · open
0.74
4MAI-Thinking-1
Microsoft
0.72
5Qwen3-Next-80B-A3B-Thinking
Alibaba Cloud / Qwen Team · open
0.72
6Qwen3 VL 235B A22B Thinking
Alibaba Cloud / Qwen Team · open
0.72
7Qwen3-235B-A22B-Thinking-2507
Alibaba Cloud / Qwen Team · open
0.72
8Qwen3 VL 32B Thinking
Alibaba Cloud / Qwen Team · open
0.72
9Qwen3-235B-A22B-Instruct-2507
Alibaba Cloud / Qwen Team · open
0.71
10Qwen3-Next-80B-A3B-Instruct
Alibaba Cloud / Qwen Team · open
0.70
11Qwen3 VL 32B Instruct
Alibaba Cloud / Qwen Team · open
0.70