all the models — AI benchmark observatory
← Benchmarks

TruthfulQA

TruthfulQA is a benchmark to measure whether language models are truthful in generating answers to questions. It comprises 817 questions that span 38 categories, including health, law, finance and politics. The questions are crafted such that some humans would answer falsely due to a false belief or misconception, testing models' ability to avoid generating false answers learned from human texts.

id truthfulqa · max 1 · 18 models reported

#ModelScore

No scores for this benchmark yet.