Vision

Apache 2.0Open weights where releasedUpdated August 2026

Pixtral 12B

Pixtral 12B is a Mistral-family model worth evaluating for vision-language local and hosted experiments.

Mistral AI · Mistral

Editorial review

Reviewed byOpenSourcesAI EditorialLast updatedAugust 2026SourcesExact model card

Model checkpoints, context windows, provider support, local runtime compatibility, and license terms can change quickly. Verify the exact model card before production or commercial use.

Best for

Vision-language local and hosted experiments

Who should use it

  • Vision-language local and hosted experiments
  • Builders who want local or self-hosted testing options.

Common workflows

  • Vision-language local and hosted experiments
  • vision workflows
  • local workflows
  • open weights workflows

Deployment and hardware notes

Roughly 7.5 GB at Q4_K_M puts it within reach of a 12 GB card, though image tokens consume context quickly on high-resolution inputs.

License and usage notes

Apache 2.0. Open weights where released. Verify the exact model card and license terms for the checkpoint or hosted provider you use.

Strengths

  • Open weights where released model option for Mistral workflows.
  • Vision-language local and hosted experiments
  • The most permissively licensed capable vision model on this page — the usual choice when a commercial product rules out the Llama, Gemma and research licences.

Limitations

  • Apache 2.0 and genuinely open, but the weights ship in Mistral's own format for vLLM and mistral-common rather than as a standard Transformers checkpoint, so the runtime choice is narrower than the licence suggests.
  • Roughly 7.5 GB at Q4_K_M puts it within reach of a 12 GB card, though image tokens consume context quickly on high-resolution inputs.
  • Context window and limits: 131,072 tokens.
  • Verify the exact model card, provider docs, license, and serving support before production use.

Local workflow notes

The most permissively licensed capable vision model on this page — the usual choice when a commercial product rules out the Llama, Gemma and research licences.

Local runtimes: Ollama where supported, LM Studio where supported, llama.cpp where supported, Transformers

Platforms: Windows, macOS, Linux

Vision spec

Memory~7.5 GB at Q4_K_M (~12.4B parameters derived from the published config: a 12B text tower plus a ~0.3B vision encoder)Image inputVariable-size images at 1024 pixels, 16-pixel patches; multi-imageContext131,072 tokens

Mistral's first vision-language release: a 12B text model paired with a 400M-class vision encoder trained from scratch, taking images at their native size rather than a fixed square, and interleaving several images with text in one prompt.

Sources to verify

Related resources

Continue with model source notes, local tools, and implementation guides related to this model.

Hardware~7.5 GB at Q4_K_M (~12.4B parameters derived from the published config: a 12B text tower plus a ~0.3B vision encoder)RuntimeOllama or LM Studio where supported, llama.cpp, Transformers, vLLMContext131,072 tokensLast updated2026
Exact model card →

Model ecosystem connections

Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.