all the models — AI benchmark observatory
← Benchmarks

Android Control Low_EM

Android control benchmark evaluating autonomous agents on mobile device interaction tasks with low exact match scoring criteria

id android-control-low-em · max 1 · 3 models reported

#ModelScore
1Qwen2.5 VL 72B Instruct
Alibaba Cloud / Qwen Team · open
0.94
2Qwen2.5 VL 32B Instruct
Alibaba Cloud / Qwen Team · open
0.93