- Models
- DeepSeek
Model family
DeepSeek Models
DeepSeek models cover reasoning, coding, distilled local variants, and open-weight assistant workflows.
Best for
Reasoning
Use this family hub to compare DeepSeek variants for reasoning workflows, then open the detail page for deeper deployment notes.
Code
Use this family hub to compare DeepSeek variants for code workflows, then open the detail page for deeper deployment notes.
Distilled
Use this family hub to compare DeepSeek variants for distilled workflows, then open the detail page for deeper deployment notes.
Local
Use this family hub to compare DeepSeek variants for local workflows, then open the detail page for deeper deployment notes.
Source box
DeepSeek family pages should distinguish V4, R1, coder, and distilled releases. Exact license, context, and hardware requirements must be read from the current checkpoint card.
Verified through: June 2026
Model licenses, context windows, release names, and provider terms can vary by checkpoint. Verify the exact model card before production or commercial use.
Jump to
Variants
DeepSeek models grouped by workflow
Latest / flagship
DeepSeek-V4-Pro
DeepSeek · DeepSeek
Best for: Teams comparing frontier-style open-weight reasoning and coding models against hosted closed models.
DeepSeek R1 Distill Qwen 14B
DeepSeek · DeepSeek
Best for: Reasoning experiments, local-friendly evaluation, and distilled model comparisons.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek-R1 Distill 14B
DeepSeek · DeepSeek-R1
Best for: DeepSeek-R1 Distill 14B is a Qwen2.5-based distillation of DeepSeek's R1 reasoning model, sized to bring stronger step-by-step reasoning to hardware that can't run the full-size R1.
Local: Qwen2.5 distillation of R1. Strong reasoning and coding at 14B. Q4_K_M weights measured at 8.99 GB, confirming the 9.0 GB estimate this entry already carried. Q4 tight in 12 GB; comfortable in 16 GB VRAM. On an RTX 3080 (10 GB) 10% of resident bytes spilled to host and generation ran 30.51 tok/s, 36% of the card's bandwidth ceiling (measured 2026-08-07, Ollama 0.32.5).
Coding
DeepSeek R1
DeepSeek · DeepSeek
Best for: Reasoning experiments, coding workflows, and evaluating open reasoning models against closed alternatives.
Local: Can be tested locally through smaller distilled or quantized variants; the full model is better suited to server-class hardware.
DeepSeek Coder V2
DeepSeek · DeepSeek
Best for: Developers comparing open coding models for IDE assistants and coding agents.
Local: Can be tested locally with smaller or quantized builds; larger variants need serious GPU memory.
DeepSeek V3
DeepSeek · DeepSeek
Best for: General DeepSeek-family comparisons and assistant workflow evaluation.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek Coder V2 Lite
DeepSeek · DeepSeek
Best for: Coding assistant tests where full-size coder models are too heavy.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek Coder 33B
DeepSeek · DeepSeek
Best for: Coding-model comparisons and legacy baseline testing.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek Coder 6.7B
DeepSeek · DeepSeek
Best for: Local coding experiments and lightweight code generation baselines.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek-R1 8B (0528 Qwen3)
DeepSeek · DeepSeek-R1
Best for: DeepSeek-R1 8B (0528 Qwen3) is a distilled reasoning model that compresses DeepSeek's R1-0528 update into an 8B Qwen3-based checkpoint small enough to run on a single consumer GPU.
Local: Qwen3-8B distillation of the R1-0528 update — Ollama's deepseek-r1:8b tag now points at this checkpoint, replacing the original Llama-3 distill. Improved math, coding, and logic over the first-release distills. Thinking chains visible in output. Q4 fits in 8 GB VRAM.
Reasoning
DeepSeek R1 Distill Llama 70B
DeepSeek · DeepSeek
Best for: Reasoning experiments, local-friendly evaluation, and distilled model comparisons.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek R1 Distill Llama 8B
DeepSeek · DeepSeek
Best for: Reasoning experiments, local-friendly evaluation, and distilled model comparisons.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek R1 Distill Qwen 32B
DeepSeek · DeepSeek
Best for: Reasoning experiments, local-friendly evaluation, and distilled model comparisons.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek R1 Distill Qwen 7B
DeepSeek · DeepSeek
Best for: Reasoning experiments, local-friendly evaluation, and distilled model comparisons.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
DeepSeek R1 Distill Qwen 1.5B
DeepSeek · DeepSeek
Best for: Reasoning experiments, local-friendly evaluation, and distilled model comparisons.
Local: Use the exact checkpoint and quantization that matches your hardware and latency target.
Compare
All DeepSeek models in the directory
| Model | Type | Best for | Local runner notes | License | Detail |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro | Reasoning | Teams comparing frontier-style open-weight reasoning and coding models against hosted closed models. | Server-class only for full weights; use smaller DeepSeek distills or hosted endpoints for everyday evaluation. | MIT / check exact model card | Open |
| DeepSeek R1 | Reasoning | Reasoning experiments, coding workflows, and evaluating open reasoning models against closed alternatives. | Can be tested locally through smaller distilled or quantized variants; the full model is better suited to server-class hardware. | MIT | Open |
| DeepSeek Coder V2 | Code | Developers comparing open coding models for IDE assistants and coding agents. | Can be tested locally with smaller or quantized builds; larger variants need serious GPU memory. | DeepSeek License | Open |
| DeepSeek V3 | Chat | General DeepSeek-family comparisons and assistant workflow evaluation. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek R1 Distill Llama 70B | Reasoning | Reasoning experiments, local-friendly evaluation, and distilled model comparisons. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek R1 Distill Llama 8B | Reasoning | Reasoning experiments, local-friendly evaluation, and distilled model comparisons. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek R1 Distill Qwen 32B | Reasoning | Reasoning experiments, local-friendly evaluation, and distilled model comparisons. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek R1 Distill Qwen 14B | Reasoning | Reasoning experiments, local-friendly evaluation, and distilled model comparisons. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek R1 Distill Qwen 7B | Reasoning | Reasoning experiments, local-friendly evaluation, and distilled model comparisons. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek R1 Distill Qwen 1.5B | Reasoning | Reasoning experiments, local-friendly evaluation, and distilled model comparisons. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek Coder V2 Lite | Code | Coding assistant tests where full-size coder models are too heavy. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek Coder 33B | Code | Coding-model comparisons and legacy baseline testing. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek Coder 6.7B | Code | Local coding experiments and lightweight code generation baselines. | Use the exact checkpoint and quantization that matches your hardware and latency target. | Check exact model card | Open |
| DeepSeek-R1 8B (0528 Qwen3) | Reasoning | DeepSeek-R1 8B (0528 Qwen3) is a distilled reasoning model that compresses DeepSeek's R1-0528 update into an 8B Qwen3-based checkpoint small enough to run on a single consumer GPU. | Qwen3-8B distillation of the R1-0528 update — Ollama's deepseek-r1:8b tag now points at this checkpoint, replacing the original Llama-3 distill. Improved math, coding, and logic over the first-release distills. Thinking chains visible in output. Q4 fits in 8 GB VRAM. | MIT | Open |
| DeepSeek-R1 Distill 14B | Reasoning | DeepSeek-R1 Distill 14B is a Qwen2.5-based distillation of DeepSeek's R1 reasoning model, sized to bring stronger step-by-step reasoning to hardware that can't run the full-size R1. | Qwen2.5 distillation of R1. Strong reasoning and coding at 14B. Q4_K_M weights measured at 8.99 GB, confirming the 9.0 GB estimate this entry already carried. Q4 tight in 12 GB; comfortable in 16 GB VRAM. On an RTX 3080 (10 GB) 10% of resident bytes spilled to host and generation ran 30.51 tok/s, 36% of the card's bandwidth ceiling (measured 2026-08-07, Ollama 0.32.5). | MIT | Open |