Hardware index · 24 guides
Local AI Hardware Guide: Match Your Budget to Real Model Limits
Choosing hardware for local AI is not about buying the biggest GPU on a spec sheet. It is about matching your budget to physical memory limits, bandwidth, model size, and the type of work you actually want to run. A small private chatbot, a coding assistant, a document RAG system, and a 70B research model all place different pressure on VRAM, system RAM, memory bandwidth, and runtime support.
VRAM is the binding constraint
Every guide is organized around what fits in memory, not spec-sheet marketing.
NVIDIA and Apple Silicon covered
CUDA-tier cards and unified-memory Macs are compared on the same practical basis.
Check before you buy
Use the compatibility checker to confirm fit before purchasing or renting hardware.
Hardware tiers
Hardware tier quick reference
VRAM is the binding constraint. Every row below reflects the model sizes that fit reliably in that memory class — bandwidth determines how fast they run, not whether they fit.
| Memory tier | What fits | Practical ceiling | Speed range | Position |
|---|---|---|---|---|
| 8GB | 1B–4B any quant · 7B/8B Q4 | 13B–14B not a comfortable full-GPU workload; buying new should start at 12GB+ | Comfortable 7B/8B Q4 | RTX 4060 / RX 6600 / Arc A750 class — entry / existing-card tier |
| 10GB | 7B Q4/Q5 · some 7B Q8 | 13B–14B Q4 is tight; buying new should start at 12GB+ | Comfortable 7B use | RTX 3080 10GB class — existing-card tier, not a 2026 buying target |
| 12GB | 7B Q8 · 14B Q4 | 14B Q4 | 15–40 t/s | Practical entry point |
| 16GB | 7B FP16 · 14B Q8 | 14B Q8 | 20–50 t/s | Quality upgrade tier |
| 24GB | 7B FP16 · 14B Q8 · 30B Q4 | 30B Q4 | 35–80 t/s | Consumer sweet spot |
| 32GB | 30B–32B Q4 with headroom · some higher quant | 70B is a stretch (offload / sub-4-bit), not a full-GPU fit | High-headroom over 24GB | RTX 5090 class — high-headroom consumer, below 48GB workstation |
| 48GB | 30B Q8 · 70B Q4 | 70B Q4 | 10–35 t/s | Workstation class |
| 64GB unified (Apple) | 14B FP16 · 32B Q8 · 70B Q4 | 70B Q4 | 10–18 t/s on 70B | macOS / Metal only |
| CPU-only | 1B–7B at Q4 | 7B Q4 practical | 1–12 t/s | No GPU required |
GPU and Apple Silicon
Hardware guides by card
NVIDIA GPUs offer the broadest runtime support for local AI — CUDA acceleration in Ollama, LM Studio, vLLM, llama.cpp, and TensorRT-LLM. Apple Silicon replaces discrete VRAM with a unified memory pool, enabling 70B models on a consumer Mac at the cost of lower tokens-per-second and macOS-only runtime support. CPU-only setups require no GPU but carry strict model size and speed limits.
NVIDIA · Entry / existing-owner tier · 8GB VRAM
RTX 4060 8GB
Compact models at full precision and many 7B/8B Q4 workloads, with limited context headroom. A capable entry card if you own one — not the preferred new local-AI purchase; start at 12GB+.
NVIDIA · Existing-owner tier · 10GB VRAM
RTX 3080 10GB
Comfortable 7B and 8B use on fast Ampere bandwidth. 13B–14B Q4 fits only with tight context headroom. A capable card if you own one — but new buyers should start at 12GB+.
NVIDIA · Budget · 12GB VRAM
RTX 3060 12GB
7B at Q8, 14B at Q4. Budget entry into the 12GB CUDA tier. Same model ceiling as higher-end 12GB cards, lower bandwidth.
NVIDIA · Entry · 16GB VRAM
RTX 4060 Ti 16GB / RTX 5060 Ti 16GB
The budget way into the 16GB tier — 7B at FP16, 13B–14B at Q8 with modest headroom. Lower bandwidth than the 4070 Ti Super or 5080 at the same VRAM.
NVIDIA · Prosumer · 12GB VRAM
RTX 4070 Ti 12GB
7B and 8B at Q8, 14B at Q4. Fastest Ada Lovelace 12GB option. 40% more tokens-per-second than the RTX 3060 at the same model.
NVIDIA · Prosumer · 16GB VRAM
RTX 4080 / 4080 Super 16GB
7B at FP16, 13B and 14B at Q8. The 16GB tier upgrade — full precision on 7B and near-lossless quality on 13B.
NVIDIA · Blackwell · 16GB VRAM
RTX 5080 16GB
The fastest current 16GB card at 960 GB/s. Comfortable 14B at Q8, tight on 24B at Q4 — the same practical ceiling as other 16GB cards, delivered faster.
NVIDIA · Enthusiast · 24GB VRAM · Used market
RTX 3090 24GB
7B at FP16, 14B at Q8, 30B at Q4. Same model ceiling as the RTX 4090 at lower cost. Last consumer card with NVLink support.
NVIDIA · Enthusiast · 24GB VRAM
RTX 4090 24GB
7B at FP16, 14B at Q8, 30B at Q4. Fastest consumer GPU for local AI — 1008 GB/s bandwidth, best tokens-per-second below workstation class.
AMD · RDNA 3 · 24GB VRAM
AMD RX 7900 XTX 24GB
Same 24GB model ceiling as the RTX 3090/4090 tier at a lower price. Runs on ROCm — Linux is the more tested path for local AI runtimes.
NVIDIA · Blackwell · 32GB VRAM
RTX 5090 32GB
The current flagship consumer card at 1792 GB/s bandwidth. Comfortable 32B at Q8; 70B at Q4 only as a stretch with offload. Pushes past the 24GB ceiling.
NVIDIA · Workstation · 48GB VRAM (NVLink pooled)
Dual RTX 3090 48GB
Two RTX 3090s pooled over NVLink — the value path into the 48GB workstation tier for 70B at Q4. Requires a tensor-parallel runtime, not beginner-friendly.
Apple Silicon · M-Series Max / Ultra · 64GB Unified Memory
Apple Silicon 64GB Unified Memory
14B at FP16, 32B at Q8, 70B at Q4. The only consumer path to 70B in a single memory pool. Metal acceleration via Ollama and LM Studio — no CUDA runtimes, macOS only.
Apple Silicon · Mac Mini · 24GB Unified
Mac Mini M4 24GB
The smallest, quietest entry into Apple Silicon local AI. Comfortable 7B–8B, selective 13B at Q4.
Apple Silicon · Mac Studio Max · 128GB Unified
Mac Studio M4 Max 128GB
70B at near-FP16 quality (Q8) in a single unified memory pool — beyond what the 64GB tier can hold.
Apple Silicon · Mac Studio Ultra · 192GB+ Unified
Mac Studio Ultra 192GB+
Scales up to 512GB unified memory. Full FP16 70B and experimental frontier-scale local models beyond what any single NVIDIA GPU can hold.
CPU-only · No GPU required
CPU-Only Local AI
Running local LLMs without a GPU — which models are practical on CPU, how fast to expect, RAM requirements, and when a GPU is worth adding.
VRAM tiers
VRAM tier reference guides
Each tier guide maps a memory class to the model sizes, quantization levels, and GPU options that fit reliably within it — useful when you know your VRAM budget but not which specific card to buy.
VRAM tier · 8GB · Entry / existing-owner
8GB VRAM Tier
What can 8GB run? 1B–4B at full precision and 7B/8B at Q4, with limited context headroom. An entry point for existing owners; new buyers should start at 12GB+.
VRAM tier · 10GB · Existing-owner
10GB VRAM Tier
What can 10GB run? Comfortable 7B/8B, tight 13B–14B Q4 with limited context headroom. Useful for existing owners; new buyers should start at 12GB+.
VRAM tier · 12GB
12GB VRAM Tier
What can 12GB of GPU VRAM run? Model fit table, GPU options at this tier, and when 24GB is worth the upgrade.
VRAM tier · 16GB
16GB VRAM Tier
7B at FP16 and 13B–14B at Q8 — the two meaningful quality upgrades the 16GB tier delivers over 12GB cards.
VRAM tier · 24GB
24GB VRAM Tier
What can 24GB run? 7B FP16, 30B Q4, and a direct comparison to Apple Silicon unified memory at the same model ceiling.
VRAM tier · 32GB · High-headroom consumer
32GB VRAM Tier
Comfortable 30B–32B at Q4 with real context, concurrency, and higher-quant headroom over 24GB. 70B stays a stretch, not a conventional full-GPU fit.
VRAM tier · 48GB
48GB VRAM Tier
Workstation-class local AI — 70B at Q4 and 30B at Q8. Three hardware paths: RTX A6000, dual RTX 3090 NVLink, Apple Silicon 64GB.
Not sure where to start?
Use the compatibility checker to estimate what your current GPU, VRAM, and system RAM can reliably execute before buying new hardware, renting cloud GPUs, or migrating your local AI stack.