all the models — AI benchmark observatory
← Benchmarks

Multi-IF

Multi-IF benchmarks LLMs on multi-turn and multilingual instruction following. It expands upon IFEval by incorporating multi-turn sequences and translating English prompts into 7 other languages, resulting in 4,501 multilingual conversations with three turns each. The benchmark reveals that current leading LLMs struggle with maintaining accuracy in multi-turn instructions and shows higher error rates for non-Latin script languages.

id multi-if · max 1 · 25 models reported

#ModelScore

No scores for this benchmark yet.