← Benchmarks
BrowseComp Long Context 128k
A challenging benchmark for evaluating web browsing agents' ability to persistently navigate the internet and find hard-to-locate, entangled information. Comprises 1,266 questions requiring strategic reasoning, creative search, and interpretation of retrieved content, with short and easily verifiable answers.
id browsecomp-long-128k · max 1 · 5 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
