← Benchmarks
MTVQA
MTVQA (Multilingual Text-Centric Visual Question Answering) is the first benchmark featuring high-quality human expert annotations across 9 diverse languages, consisting of 6,778 question-answer pairs across 2,116 images. It addresses visual-textual misalignment problems in multilingual text-centric VQA.
id mtvqa · max 1 · 2 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
