Hardware index · 24 guides

Local AI Hardware Guide: Match Your Budget to Real Model Limits

Choosing hardware for local AI is not about buying the biggest GPU on a spec sheet. It is about matching your budget to physical memory limits, bandwidth, model size, and the type of work you actually want to run. A small private chatbot, a coding assistant, a document RAG system, and a 70B research model all place different pressure on VRAM, system RAM, memory bandwidth, and runtime support.

17Card guides
7VRAM tier guides

VRAM is the binding constraint

Every guide is organized around what fits in memory, not spec-sheet marketing.

NVIDIA and Apple Silicon covered

CUDA-tier cards and unified-memory Macs are compared on the same practical basis.

Check before you buy

Use the compatibility checker to confirm fit before purchasing or renting hardware.

Hardware tiers

Hardware tier quick reference

VRAM is the binding constraint. Every row below reflects the model sizes that fit reliably in that memory class — bandwidth determines how fast they run, not whether they fit.

Memory tierWhat fitsPractical ceilingSpeed rangePosition
8GB1B–4B any quant · 7B/8B Q413B–14B not a comfortable full-GPU workload; buying new should start at 12GB+Comfortable 7B/8B Q4RTX 4060 / RX 6600 / Arc A750 class — entry / existing-card tier
10GB7B Q4/Q5 · some 7B Q813B–14B Q4 is tight; buying new should start at 12GB+Comfortable 7B useRTX 3080 10GB class — existing-card tier, not a 2026 buying target
12GB7B Q8 · 14B Q414B Q415–40 t/sPractical entry point
16GB7B FP16 · 14B Q814B Q820–50 t/sQuality upgrade tier
24GB7B FP16 · 14B Q8 · 30B Q430B Q435–80 t/sConsumer sweet spot
32GB30B–32B Q4 with headroom · some higher quant70B is a stretch (offload / sub-4-bit), not a full-GPU fitHigh-headroom over 24GBRTX 5090 class — high-headroom consumer, below 48GB workstation
48GB30B Q8 · 70B Q470B Q410–35 t/sWorkstation class
64GB unified (Apple)14B FP16 · 32B Q8 · 70B Q470B Q410–18 t/s on 70BmacOS / Metal only
CPU-only1B–7B at Q47B Q4 practical1–12 t/sNo GPU required

GPU and Apple Silicon

Hardware guides by card

NVIDIA GPUs offer the broadest runtime support for local AI — CUDA acceleration in Ollama, LM Studio, vLLM, llama.cpp, and TensorRT-LLM. Apple Silicon replaces discrete VRAM with a unified memory pool, enabling 70B models on a consumer Mac at the cost of lower tokens-per-second and macOS-only runtime support. CPU-only setups require no GPU but carry strict model size and speed limits.

NVIDIA · Entry / existing-owner tier · 8GB VRAM

RTX 4060 8GB

Compact models at full precision and many 7B/8B Q4 workloads, with limited context headroom. A capable entry card if you own one — not the preferred new local-AI purchase; start at 12GB+.

NVIDIA · Existing-owner tier · 10GB VRAM

RTX 3080 10GB

Comfortable 7B and 8B use on fast Ampere bandwidth. 13B–14B Q4 fits only with tight context headroom. A capable card if you own one — but new buyers should start at 12GB+.

NVIDIA · Budget · 12GB VRAM

RTX 3060 12GB

7B at Q8, 14B at Q4. Budget entry into the 12GB CUDA tier. Same model ceiling as higher-end 12GB cards, lower bandwidth.

NVIDIA · Entry · 16GB VRAM

RTX 4060 Ti 16GB / RTX 5060 Ti 16GB

The budget way into the 16GB tier — 7B at FP16, 13B–14B at Q8 with modest headroom. Lower bandwidth than the 4070 Ti Super or 5080 at the same VRAM.

NVIDIA · Prosumer · 12GB VRAM

RTX 4070 Ti 12GB

7B and 8B at Q8, 14B at Q4. Fastest Ada Lovelace 12GB option. 40% more tokens-per-second than the RTX 3060 at the same model.

NVIDIA · Prosumer · 16GB VRAM

RTX 4080 / 4080 Super 16GB

7B at FP16, 13B and 14B at Q8. The 16GB tier upgrade — full precision on 7B and near-lossless quality on 13B.

NVIDIA · Blackwell · 16GB VRAM

RTX 5080 16GB

The fastest current 16GB card at 960 GB/s. Comfortable 14B at Q8, tight on 24B at Q4 — the same practical ceiling as other 16GB cards, delivered faster.

NVIDIA · Enthusiast · 24GB VRAM · Used market

RTX 3090 24GB

7B at FP16, 14B at Q8, 30B at Q4. Same model ceiling as the RTX 4090 at lower cost. Last consumer card with NVLink support.

NVIDIA · Enthusiast · 24GB VRAM

RTX 4090 24GB

7B at FP16, 14B at Q8, 30B at Q4. Fastest consumer GPU for local AI — 1008 GB/s bandwidth, best tokens-per-second below workstation class.

AMD · RDNA 3 · 24GB VRAM

AMD RX 7900 XTX 24GB

Same 24GB model ceiling as the RTX 3090/4090 tier at a lower price. Runs on ROCm — Linux is the more tested path for local AI runtimes.

NVIDIA · Blackwell · 32GB VRAM

RTX 5090 32GB

The current flagship consumer card at 1792 GB/s bandwidth. Comfortable 32B at Q8; 70B at Q4 only as a stretch with offload. Pushes past the 24GB ceiling.

NVIDIA · Workstation · 48GB VRAM (NVLink pooled)

Dual RTX 3090 48GB

Two RTX 3090s pooled over NVLink — the value path into the 48GB workstation tier for 70B at Q4. Requires a tensor-parallel runtime, not beginner-friendly.

Apple Silicon · M-Series Max / Ultra · 64GB Unified Memory

Apple Silicon 64GB Unified Memory

14B at FP16, 32B at Q8, 70B at Q4. The only consumer path to 70B in a single memory pool. Metal acceleration via Ollama and LM Studio — no CUDA runtimes, macOS only.

Apple Silicon · Mac Mini · 24GB Unified

Mac Mini M4 24GB

The smallest, quietest entry into Apple Silicon local AI. Comfortable 7B–8B, selective 13B at Q4.

Apple Silicon · Mac Studio Max · 128GB Unified

Mac Studio M4 Max 128GB

70B at near-FP16 quality (Q8) in a single unified memory pool — beyond what the 64GB tier can hold.

Apple Silicon · Mac Studio Ultra · 192GB+ Unified

Mac Studio Ultra 192GB+

Scales up to 512GB unified memory. Full FP16 70B and experimental frontier-scale local models beyond what any single NVIDIA GPU can hold.

CPU-only · No GPU required

CPU-Only Local AI

Running local LLMs without a GPU — which models are practical on CPU, how fast to expect, RAM requirements, and when a GPU is worth adding.

VRAM tiers

VRAM tier reference guides

Each tier guide maps a memory class to the model sizes, quantization levels, and GPU options that fit reliably within it — useful when you know your VRAM budget but not which specific card to buy.

Not sure where to start?

Use the compatibility checker to estimate what your current GPU, VRAM, and system RAM can reliably execute before buying new hardware, renting cloud GPUs, or migrating your local AI stack.