all the models — AI benchmark observatory
← Benchmarks

OSWorld Extended

OSWorld is a scalable, real computer environment benchmark for evaluating multimodal agents on open-ended tasks across Ubuntu, Windows, and macOS. It comprises 369 computer tasks involving real web and desktop applications, OS file I/O, and multi-application workflows. The benchmark evaluates agents' ability to interact with computer interfaces using screenshots and actions in realistic computing environments.

id osworld-extended · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

OSWorld Extended Leaderboard · all the models