all the models — AI benchmark observatory
← Benchmarks

MCP-Universe

MCP-Universe evaluates LLMs on complex multi-step agentic tasks using Model Context Protocol (MCP) tools across diverse interactive environments, testing planning, tool orchestration, and task completion.

id mcp-universe · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.