all the models — AI benchmark observatory
← Benchmarks

NIH/Multi-needle

Multi-needle in a haystack benchmark for evaluating long-context comprehension capabilities of language models by testing retrieval of multiple target pieces of information from extended documents

id nih-multi-needle · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

NIH/Multi-needle Leaderboard · all the models