all the models — AI benchmark observatory
← Models

North Micro Vision Instruct

Cohere · open weight · north-micro-vision-instruct

Compact open-weight vision-language model (2.4B) with native-resolution image support for VQA, captioning, grounding, OCR, charts, and documents. Custom 400M SigLIP 2-based vision encoder + 2B North Micro LLM backbone (Command A+ style). Multilingual and multi-image. LM context 128K; multimodal training validated to 8K. Not a reasoning/tool-calling model. Intended for prototyping and fine-tuning.

DocVQAIFEvalMMLUMMLU-Pro

Benchmark scores

BenchmarkScore
DocVQA0.92
IFEval0.75
MMLU0.50
MMLU-Pro0.31

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.