← Benchmarks
Nexus
NexusRaven benchmark for evaluating function calling capabilities of large language models in zero-shot scenarios across cybersecurity tools and API interactions
id nexus · max 1 · 4 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
