← Benchmarks
Graphwalks BFS >128k
A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length over 128k tokens, testing long-context reasoning capabilities.
id graphwalks-bfs->128k · max 1 · 11 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
