Vision

Mistral Research License (non-commercial)Open weights where releasedUpdated August 2026

Pixtral Large

Pixtral Large is a Mistral-family model worth evaluating for vision-language assistant workflows.

Mistral AI · Mistral

Editorial review

Reviewed byOpenSourcesAI EditorialLast updatedAugust 2026SourcesExact model card

Model checkpoints, context windows, provider support, local runtime compatibility, and license terms can change quickly. Verify the exact model card before production or commercial use.

Best for

Vision-language assistant workflows

Who should use it

  • Vision-language assistant workflows
  • Builders who want local or self-hosted testing options.

Common workflows

  • Vision-language assistant workflows
  • vision workflows
  • local workflows
  • open weights workflows

Deployment and hardware notes

Not a single-GPU model: plan for multiple data-centre cards or a server-class host, and treat any single-card claim about it with suspicion.

License and usage notes

Mistral Research License (non-commercial). Open weights where released. Verify the exact model card and license terms for the checkpoint or hosted provider you use.

Strengths

  • Open weights where released model option for Mistral workflows.
  • Vision-language assistant workflows
  • For local work, pixtral-12b is the same lineage and the same image handling under a licence you can actually ship.

Limitations

  • The Mistral Research Licence is non-commercial, which rules the model out of most products regardless of hardware; a separate commercial licence must be negotiated. At roughly 75 GB even at Q4_K_M, it is a multi-GPU or server deployment, not a local one.
  • Not a single-GPU model: plan for multiple data-centre cards or a server-class host, and treat any single-card claim about it with suspicion.
  • Context window and limits: 131,072 tokens.
  • Verify the exact model card, provider docs, license, and serving support before production use.

Local workflow notes

For local work, pixtral-12b is the same lineage and the same image handling under a licence you can actually ship.

Local runtimes: Ollama where supported, LM Studio where supported, llama.cpp where supported, Transformers

Platforms: Windows, macOS, Linux

Vision spec

Memory~75 GB at Q4_K_M (~123B parameters derived from the published config: 88 layers at 12,288 hidden, the Mistral Large 2 shape, plus a ~1B vision encoder) — multi-GPU or server classImage inputVariable-size images at 1024 pixels, 16-pixel patches; multi-imageContext131,072 tokens

Pixtral's frontier-scale sibling: the Mistral Large 2 text tower with a 40-layer vision encoder on top. Mistral publishes no parameter count, so the ~123B figure here is derived from the checkpoint's own published configuration rather than quoted.

Sources to verify

Related resources

Continue with model source notes, local tools, and implementation guides related to this model.

Hardware~75 GB at Q4_K_M (~123B parameters derived from the published config: 88 layers at 12,288 hidden, the Mistral Large 2 shape, plus a ~1B vision encoder) — multi-GPU or server classRuntimeOllama or LM Studio where supported, llama.cpp, Transformers, vLLMContext131,072 tokensLast updated2026
Exact model card →

Model ecosystem connections

Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.