Model family

MetaUpdated 2026ChatMultimodalLocalSafety

Llama Models

Meta Llama models are widely supported open-weight options for local AI stacks, multimodal workflows, assistant prototypes, and guardrail experiments.

Best for

Chat

Use this family hub to compare Llama variants for chat workflows, then open the detail page for deeper deployment notes.

Multimodal

Use this family hub to compare Llama variants for multimodal workflows, then open the detail page for deeper deployment notes.

Local

Use this family hub to compare Llama variants for local workflows, then open the detail page for deeper deployment notes.

Safety

Use this family hub to compare Llama variants for safety workflows, then open the detail page for deeper deployment notes.

Variants

Llama models grouped by workflow

Latest / flagship

Coding

Vision / multimodal

Safety / guardrails

Local-friendly

Compare

All Llama models in the directory

ModelTypeBest forLocal runner notesLicenseDetail
Llama 3 70BChatBuilders who want a widely supported open-weight chat model with broad runtime compatibility.Commonly used in local workflows through quantized builds, but 70B-class models are best with high-memory GPUs or workstation/server hardware.Llama 3 Community LicenseOpen
Llama 4 ScoutMultimodalTeams evaluating Llama-family models for multimodal assistant, long-context, and application workflows.Evaluate local fit with the exact checkpoint and quantization available for your runtime.Llama license / check exact model cardOpen
Llama 4 MaverickMultimodalBuilders comparing current Llama-family models for assistant, multimodal, and reasoning-oriented workflows.Evaluate local fit with the exact checkpoint and quantization available for your runtime.Llama license / check exact model cardOpen
Llama 3.3 70B InstructChatGeneral assistant workflows, app prototypes, and Llama-family baseline comparisons.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
Llama 3.1 405B InstructChatServer-class assistant evaluation and comparisons against smaller Llama variants.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
Llama 3.1 70B InstructChatTeams comparing widely supported Llama-family 70B-class models.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
Llama 3.1 8B InstructEdgeLocal prototypes, small assistants, and lower-resource evaluation.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
Llama 3 8B InstructEdgeLocal baseline comparisons and lightweight app prototypes.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
Llama Guard 3SafetySafety checks, guardrail experiments, and policy classification workflows.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
Llama 3.2 VisionVisionVision-language experiments, screenshot reasoning, and multimodal app prototypes.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
Llama 3 8BChatLlama 3 8B is Meta's original Llama 3 checkpoint, the direct predecessor to the more widely used Llama 3.1 revision.Widely supported across runtimes and tools. GQA reduces KV cache vs older Llama 2. Good baseline for chat and coding on 8 GB VRAM.Llama 3 Community LicenseOpen
Llama 3.3 70BChatLlama 3.3 70B is Meta's refreshed 70B release, improving instruction-following over the earlier Llama 3.1 70B checkpoint while keeping the same parameter count and 128K context window.December 2024 Llama release with improved instruction following over 3.1 70B at the same 70B parameter count and 128K context. Q4 requires ~41 GB VRAM — suited for RTX 6000 Ada (48 GB), multi-GPU setups, or Mac with 64 GB+ unified memory.Llama 3.3 Community LicenseOpen