all the models — AI benchmark observatory
← Benchmarks

Graphwalks BFS <128k

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length under 128k tokens, returning nodes reachable at specified depths.

id graphwalks-bfs-<128k · max 1 · 12 models reported

#ModelScore
1GPT-5.2
OpenAI
0.94
2GPT-5.4
OpenAI
0.93
Graphwalks BFS <128k Leaderboard · all the models