all the models — AI benchmark observatory
← Benchmarks

SIFO-Multiturn

SIFO-Multiturn evaluates instruction following capabilities in multi-turn conversational settings, testing how well models maintain context and follow instructions across multiple exchanges.

id sifo-multiturn · max 100 · 1 models reported

#ModelScore
1Qwen3 VL 235B A22B Thinking
Alibaba Cloud / Qwen Team · open
0.71