all the models — AI benchmark observatory
← Benchmarks

BrowseComp Long Context 128k

A challenging benchmark for evaluating web browsing agents' ability to persistently navigate the internet and find hard-to-locate, entangled information. Comprises 1,266 questions requiring strategic reasoning, creative search, and interpretation of retrieved content, with short and easily verifiable answers.

id browsecomp-long-128k · max 1 · 5 models reported

#ModelScore

No scores for this benchmark yet.