all the models — AI benchmark observatory
← Benchmarks

QwenWorldBench

QwenWorldBench is Qwen's internal benchmark for evaluating LLMs as world models that simulate agentic environments across Terminal, SWE, MCP, Search, OS, Android, and Web domains.

id qwenworldbench · max 1 · 2 models reported

#ModelScore

No scores for this benchmark yet.