all the models — AI benchmark observatory
← Benchmarks

OSWorld 2.0

OSWorld 2.0 is a benchmark of 108 long-horizon, real-world computer-use workflows spanning everyday and professional tasks. Each task is an end-to-end workflow that takes human users a median of about 1.6 hours, scored with a binary-completion metric, and targets challenges such as dynamic environments, cross-source reasoning, and implicit-state inference.

id osworld-2.0 · max 1 · 14 models reported

#ModelScore
1Claude Opus 5.5
Anthropic
0.82
2Claude Fable 5.1
Anthropic
0.78
3GPT-6 Astra
OpenAI
0.73
4Claude Opus 5
Anthropic
0.71
OSWorld 2.0 Leaderboard · all the models