all the models — AI benchmark observatory
← Benchmarks

POPE

Polling-based Object Probing Evaluation (POPE) is a benchmark for evaluating object hallucination in Large Vision-Language Models (LVLMs). POPE addresses the problem where LVLMs generate objects inconsistent with target images by using a polling-based query method that asks yes/no questions about object presence in images, providing more stable and flexible evaluation of object hallucination.

id pope · max 1 · 3 models reported

#ModelScore
1LFM2.5-VL-3B
Liquid AI · open
0.89
2Phi-3.5-vision-instruct
Microsoft · open
0.86
3Phi-4-multimodal-instruct
Microsoft · open
0.86