← Benchmarks
SIFO-Multiturn
SIFO-Multiturn evaluates instruction following capabilities in multi-turn conversational settings, testing how well models maintain context and follow instructions across multiple exchanges.
id sifo-multiturn · max 100 · 1 models reported
| # | Model | Score |
|---|---|---|
| 1 | Qwen3 VL 235B A22B Thinking Alibaba Cloud / Qwen Team · open | 0.71 |
