all the models — AI benchmark observatory
← Benchmarks

Vibe-Eval

VIBE-Eval is a hard evaluation suite for measuring progress of multimodal language models, consisting of 269 visual understanding prompts with gold-standard responses authored by experts. The benchmark has dual objectives: vibe checking multimodal chat models for day-to-day tasks and rigorously testing frontier models, with the hard set containing >50% questions that all frontier models answer incorrectly.

id vibe-eval · max 1 · 8 models reported

#ModelScore

No scores for this benchmark yet.