all the models — AI benchmark observatory
← Benchmarks

Hallusion Bench

A comprehensive benchmark designed to evaluate image-context reasoning in large visual-language models (LVLMs) by challenging models with 346 images and 1,129 carefully crafted questions to assess language hallucination and visual illusion

id hallusion-bench · max 1 · 20 models reported

#ModelScore

No scores for this benchmark yet.

Hallusion Bench Leaderboard · all the models