all the models — AI benchmark observatory
← Benchmarks

LongFact

LongFact evaluates factual precision over long-form generations containing many individual claims. Each claim is extracted and verified, and the model is scored on claim-level precision, measuring whether extended responses introduce unsupported or false statements.

id longfact · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.