all the models — AI benchmark observatory
← Models

Phi-3.5-vision-instruct

Microsoft · open weight · phi-3.5-vision-instruct

Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image comparison, summarization, and even video analysis. The model underwent safety post-training for improved instruction-following, alignment, and robust handling of visual and text inputs, and is released under the MIT license.

ScienceQAPOPEMMMU

Benchmark scores

BenchmarkScore
ScienceQA0.91
POPE0.86
MMMU0.43

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Phi-3.5-vision-instruct Benchmarks · all the models