← Benchmarks
FACTS Grounding
A benchmark evaluating language models' ability to generate factually accurate and well-grounded responses based on long-form input context, comprising 1,719 examples with documents up to 32k tokens requiring detailed responses that are fully grounded in provided documents
id facts-grounding · max 1 · 13 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
