← Benchmarks
LongFact
LongFact evaluates factual precision over long-form generations containing many individual claims. Each claim is extracted and verified, and the model is scored on claim-level precision, measuring whether extended responses introduce unsupported or false statements.
id longfact · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
