← Benchmarks
ComplexFuncBench
ComplexFuncBench is a benchmark designed to evaluate large language models' capabilities in handling complex function calling scenarios. It encompasses multi-step and constrained function calling tasks that require long-parameter filling, parameter value reasoning, and managing contexts up to 128k tokens. The benchmark includes 1,000 samples across five real-world scenarios.
id complexfuncbench · max 1 · 7 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
