Multimodal

Llama 4 Community LicenseOpen weightsUpdated August 2026Frontier 2026

Llama 4 Maverick

Open-weight Llama 4 model positioned for multimodal, reasoning, and general assistant workflows.

Meta · Llama

Editorial review

Reviewed byOpenSourcesAI EditorialLast updatedAugust 2026SourcesExact model card

Model checkpoints, context windows, provider support, local runtime compatibility, and license terms can change quickly. Verify the exact model card before production or commercial use.

Best for

Builders comparing current Llama-family models for assistant, multimodal, and reasoning-oriented workflows.

Who should use it

  • Builders comparing current Llama-family models for assistant, multimodal, and reasoning-oriented workflows.
  • Builders who want local or self-hosted testing options.

Common workflows

  • Multimodal, reasoning, and general assistant workflows
  • multimodal workflows
  • reasoning workflows
  • assistant workflows
  • open weights workflows

Deployment and hardware notes

Full deployments may need high-memory GPUs or hosted inference; local testing depends on available quantized builds.

License and usage notes

Llama 4 Community License. Open weights. Verify the exact model card and license terms for the checkpoint or hosted provider you use.

Strengths

  • Open weights model option for Llama workflows.
  • Builders comparing current Llama-family models for assistant, multimodal, and reasoning-oriented workflows.
  • Evaluate local fit with the exact checkpoint and quantization available for your runtime.
  • Tracked as Frontier 2026 in the OpenSourcesAI model directory.

Limitations

  • Review license terms, availability, runtime support, and hardware requirements for the exact release before standardizing.
  • Full deployments may need high-memory GPUs or hosted inference; local testing depends on available quantized builds.
  • Context window and limits: 1,048,576 tokens (~1M), confirmed from the model's published config.
  • Verify the exact model card, provider docs, license, and serving support before production use.

Local workflow notes

Evaluate local fit with the exact checkpoint and quantization available for your runtime.

Local runtimes: Transformers, vLLM where supported

Platforms: Windows, macOS, Linux, Self-hosted servers

Vision spec

Memory~243 GB at Q4_K_M (401.6B total parameters, 128-expert MoE) — multi-GPU or server classImage input336x336 tiles, 14-pixel patches; Meta states testing to 5 input imagesContext1,048,576 tokens (~1M), confirmed from the model's published config

Meta's first natively multimodal Llama: image and text tokens are fused in the backbone from pretraining rather than bridged by an adapter after the fact, which is the architectural break from Llama 3.2 Vision. The tile geometry and the 34-layer vision encoder are read from the checkpoint's own published configuration; the five-image figure is Meta's own stated testing limit, not a hard cap.

Frontier-model verification note

This page is written to stay accurate as of the latest available 2026 public model information. Availability, licenses, context windows, API support, pricing, benchmark standing, and local-serving support can change quickly. Verify the official model card, provider docs, and license before using this model in production or commercial workflows.

Sources to verify

Related resources

Continue with model source notes, local tools, and implementation guides related to this model.

Hardware~243 GB at Q4_K_M (401.6B total parameters, 128-expert MoE) — multi-GPU or server classRuntimevLLM, Transformers, hosted providers where supportedContext1,048,576 tokens (~1M), confirmed from the model's published configLast updated2026
Exact model card →

Model ecosystem connections

Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.