Vision
Pixtral Large
Pixtral Large is a Mistral-family model worth evaluating for vision-language assistant workflows.
Mistral AI · Mistral
Editorial review
Model checkpoints, context windows, provider support, local runtime compatibility, and license terms can change quickly. Verify the exact model card before production or commercial use.
Best for
Vision-language assistant workflows
Who should use it
- Vision-language assistant workflows
- Builders who want local or self-hosted testing options.
Common workflows
- Vision-language assistant workflows
- vision workflows
- local workflows
- open weights workflows
Deployment and hardware notes
Not a single-GPU model: plan for multiple data-centre cards or a server-class host, and treat any single-card claim about it with suspicion.
License and usage notes
Mistral Research License (non-commercial). Open weights where released. Verify the exact model card and license terms for the checkpoint or hosted provider you use.
Strengths
- Open weights where released model option for Mistral workflows.
- Vision-language assistant workflows
- For local work, pixtral-12b is the same lineage and the same image handling under a licence you can actually ship.
Limitations
- The Mistral Research Licence is non-commercial, which rules the model out of most products regardless of hardware; a separate commercial licence must be negotiated. At roughly 75 GB even at Q4_K_M, it is a multi-GPU or server deployment, not a local one.
- Not a single-GPU model: plan for multiple data-centre cards or a server-class host, and treat any single-card claim about it with suspicion.
- Context window and limits: 131,072 tokens.
- Verify the exact model card, provider docs, license, and serving support before production use.
Local workflow notes
For local work, pixtral-12b is the same lineage and the same image handling under a licence you can actually ship.
Local runtimes: Ollama where supported, LM Studio where supported, llama.cpp where supported, Transformers
Platforms: Windows, macOS, Linux
Vision spec
Pixtral's frontier-scale sibling: the Mistral Large 2 text tower with a 40-layer vision encoder on top. Mistral publishes no parameter count, so the ~123B figure here is derived from the checkpoint's own published configuration rather than quoted.
Sources to verify
Related resources
Continue with model source notes, local tools, and implementation guides related to this model.
Model ecosystem connections
Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.
Recommended runtimes and tools
Setup and deployment
Related model pages
Guides, stacks, and comparisons
Ready to run this model locally?
Find a compatible interface in our Local AI Tools directory →