all the models — AI benchmark observatory
← Benchmarks

Search and Function-Calling

Search and Function-Calling is an OpenAI internal production benchmark measuring reliable search-tool use and function calling in agentic workflows, reported as a pass rate.

id openai-search-function-calling · max 1 · 3 models reported

#ModelScore
1GPT-5.6 Terra
OpenAI
0.95
2GPT-5.6 Sol
OpenAI
0.91
3GPT-5.6 Luna
OpenAI
0.90