← Benchmarks
GraphWalks
GraphWalks is a synthetic multi-hop long-context reasoning benchmark in which a model is given an edge-list representation of a graph and must traverse it to find neighboring nodes (via breadth-first search) or parent nodes for a given start node. Performance is reported as F1 of the model-predicted answer set versus the ground truth.
id graphwalks · max 1 · 3 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
