← Benchmarks
Hallusion Bench
A comprehensive benchmark designed to evaluate image-context reasoning in large visual-language models (LVLMs) by challenging models with 346 images and 1,129 carefully crafted questions to assess language hallucination and visual illusion
id hallusion-bench · max 1 · 20 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
