Vision
Pixtral 12B
Pixtral 12B is a Mistral-family model worth evaluating for vision-language local and hosted experiments.
Mistral AI · Mistral
Editorial review
Model checkpoints, context windows, provider support, local runtime compatibility, and license terms can change quickly. Verify the exact model card before production or commercial use.
Best for
Vision-language local and hosted experiments
Who should use it
- Vision-language local and hosted experiments
- Builders who want local or self-hosted testing options.
Common workflows
- Vision-language local and hosted experiments
- vision workflows
- local workflows
- open weights workflows
Deployment and hardware notes
Roughly 7.5 GB at Q4_K_M puts it within reach of a 12 GB card, though image tokens consume context quickly on high-resolution inputs.
License and usage notes
Apache 2.0. Open weights where released. Verify the exact model card and license terms for the checkpoint or hosted provider you use.
Strengths
- Open weights where released model option for Mistral workflows.
- Vision-language local and hosted experiments
- The most permissively licensed capable vision model on this page — the usual choice when a commercial product rules out the Llama, Gemma and research licences.
Limitations
- Apache 2.0 and genuinely open, but the weights ship in Mistral's own format for vLLM and mistral-common rather than as a standard Transformers checkpoint, so the runtime choice is narrower than the licence suggests.
- Roughly 7.5 GB at Q4_K_M puts it within reach of a 12 GB card, though image tokens consume context quickly on high-resolution inputs.
- Context window and limits: 131,072 tokens.
- Verify the exact model card, provider docs, license, and serving support before production use.
Local workflow notes
The most permissively licensed capable vision model on this page — the usual choice when a commercial product rules out the Llama, Gemma and research licences.
Local runtimes: Ollama where supported, LM Studio where supported, llama.cpp where supported, Transformers
Platforms: Windows, macOS, Linux
Vision spec
Mistral's first vision-language release: a 12B text model paired with a 400M-class vision encoder trained from scratch, taking images at their native size rather than a fixed square, and interleaving several images with text in one prompt.
Sources to verify
Related resources
Continue with model source notes, local tools, and implementation guides related to this model.
Model ecosystem connections
Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.
Recommended runtimes and tools
Setup and deployment
Related model pages
Guides, stacks, and comparisons
Ready to run this model locally?
Find a compatible interface in our Local AI Tools directory →