all the models — AI benchmark observatory
← Benchmarks

Graphwalks BFS >128k

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length over 128k tokens, testing long-context reasoning capabilities.

id graphwalks-bfs->128k · max 1 · 11 models reported

#ModelScore

No scores for this benchmark yet.

Graphwalks BFS >128k Leaderboard · all the models