Qwen3-Coder 30B (A3B)
30.5B parameter open-weight model. The coding variant of the 30B sparse MoE - same 128-expert, ~3B-active architecture tuned for agentic coding, with native 256K context for large repos. Q4_K_M is ~19 GB; measured 31.4 tok/s on an RTX 3080 (10 GB) at 46% VRAM residency (2026-08-01), so it stays responsive even when split. Apache 2.0. The natural upgrade path from Qwen2.5-Coder.
Alibaba · Qwen3 Coder
Editorial review
VRAM figures are empirical estimates. Actual usage varies by runtime, context length, and system configuration. Verify on your specific hardware before production use.
Will Qwen3-Coder 30B (A3B) run on your machine?
Qwen3-Coder 30B (A3B) is 30.5B parameters and needs 20.5 GB of VRAM at Q4_K_M — 19 GB of weights plus 1.5 GB of runtime overhead for the inference server itself.
VRAM by quantization
| Quantization | Weights | Needs (with overhead) | Quality |
|---|---|---|---|
| Q4_K_M | 19 GB | 20.5 GB | good |
Fit on common hardware at Q4_K_M
| Hardware | Memory the model can use | System RAM | Verdict |
|---|---|---|---|
| CPU Only | None (CPU only) | 16 GB | Too large |
| RTX 4060 Laptop | 8 GB | 16 GB | Too large |
| RTX 3060 (12GB) | 12 GB | 32 GB | CPU offload |
| RTX 4060 Ti (16GB) | 16 GB | 32 GB | CPU offload |
| RTX 3090 | 24 GB | 64 GB | Comfortable |
| Apple Silicon (Unified Memory) 36 GB | 27 GB of 36 GB | 36 GB | Comfortable |
| RTX 5090 | 32 GB | 64 GB | Comfortable |
Comfortable means VRAM clears the requirement by 2 GB or more. Tight means it covers the requirement with no margin. CPU offload means the model does not fit in VRAM but system RAM is at least 1.6× the weights, so it will run at reduced speed — expect roughly 1–5 tokens per second. Figures are weights plus a fixed runtime overhead and exclude KV-cache growth, which scales with context length.
Apple Silicon shares one pool of memory between the system and the GPU, so a model cannot use all of it. These rows apply the same 75% usable fraction the Compatibility Checker uses, which is why a 36 GB Mac is graded on less than 36 GB.
Ready to run this model locally?
Find a compatible interface in our Local AI Tools directory →