all the models — AI benchmark observatory
← Benchmarks

Vending-Bench 2

Vending-Bench 2 tests longer horizon planning capabilities by evaluating how well AI models can manage a simulated vending machine business over extended periods. The benchmark measures a model's ability to maintain consistent tool usage and decision-making for a full simulated year of operation, driving higher returns without drifting off task.

id vending-bench-2 · max 1 · 4 models reported

#ModelScore
1Claude Opus 4.6
Anthropic
8017.59
2GLM-5.1
Zhipu AI · open
5634.41
3Gemini 3 Pro
Google
5478.16
4Gemini 3 Flash
Google
3635.00
Vending-Bench 2 Leaderboard · all the models