← Benchmarks
OSWorld 2.0
OSWorld 2.0 is a benchmark of 108 long-horizon, real-world computer-use workflows spanning everyday and professional tasks. Each task is an end-to-end workflow that takes human users a median of about 1.6 hours, scored with a binary-completion metric, and targets challenges such as dynamic environments, cross-source reasoning, and implicit-state inference.
id osworld-2.0 · max 1 · 14 models reported
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5.5 Anthropic | 0.82 |
| 2 | Claude Fable 5.1 Anthropic | 0.78 |
| 3 | GPT-6 Astra OpenAI | 0.73 |
| 4 | Claude Opus 5 Anthropic | 0.71 |
