← Benchmarks
NIH/Multi-needle
Multi-needle in a haystack benchmark for evaluating long-context comprehension capabilities of language models by testing retrieval of multiple target pieces of information from extended documents
id nih-multi-needle · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
