← Benchmarks
Vending-Bench 2
Vending-Bench 2 tests longer horizon planning capabilities by evaluating how well AI models can manage a simulated vending machine business over extended periods. The benchmark measures a model's ability to maintain consistent tool usage and decision-making for a full simulated year of operation, driving higher returns without drifting off task.
id vending-bench-2 · max 1 · 4 models reported
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 Anthropic | 8017.59 |
| 2 | GLM-5.1 Zhipu AI · open | 5634.41 |
| 3 | Gemini 3 Pro Google | 5478.16 |
| 4 | Gemini 3 Flash Google | 3635.00 |
