Local LLM hardware tool

Free — no signupUpdated July 2026

Find local AI models that should realistically fit your PC.

Enter your CPU, GPU, Graphics Card Memory (VRAM), System RAM, and target workflow. OpenSourcesAI will suggest good, better, and best local model profiles to try, plus a practical app/runtime path for testing them on your machine.

Optional helper

Auto-detect my device basics

We ask your browser for approximate memory, CPU thread count, and graphics renderer info. This is not a full PC scan — review and correct any field before running the checker.

Local LLM Hardware Compatibility Matrix

Select your graphics card memory (VRAM) and system RAM. The grading engine matches your hardware against each model’s minimum requirements. We require at least 2 GB of VRAM headroom above the model’s minimum before rating a fit as Comfortable.

Your hardware
Presets are optional shortcuts covering CPU-only machines, NVIDIA gaming cards, Apple Silicon, and multi-GPU workstations (the 2× and 4× GPU rigs at the end of the list). Selecting one fills VRAM, RAM, and GPU name instantly — adjust any field to fine-tune. Card not listed? Skip the presets and set VRAM and RAM directly below.

Scores use practical VRAM fit rules — not live benchmarks. Confirm specs against your GPU manufacturer before long inference sessions.

Results are ranked for: Chat and general assistant · Balanced quality and speed — change the two fields above to match your goal.

Set your GPU VRAM or system RAM above to run the check — at least one of the two is needed. Not sure? Use auto-detect or pick a preset.

This checker uses practical rules, not live performance tests. Treat the result as a starting point before testing a compressed model locally — lower-memory PCs may show only one realistic local match. See exactly how results are computed →

Next steps

Quick definitions

Show definitions

VRAM

Graphics-card memory

VRAM is the dedicated memory on your GPU. Common values are 4 GB, 8 GB, 10 GB, 12 GB, 16 GB, or 24 GB.

Quantized model

A compressed local model

A 4-bit quantized model uses less memory, usually with a small quality tradeoff, so it is easier to run on a normal PC.

Runtime

The app that runs the model

Ollama and LM Studio are beginner-friendly apps for testing local models. llama.cpp is a technical engine used by many local model tools.

RAG

Answers from your files

RAG means the model checks your documents or notes first, then answers using those sources.

Open-weight

Public model weights

Open-weight means the model weights are public, but the usage terms may still have rules.

How to read your checker results

Show how to read results

This tool is for anyone who wants to run AI models on their own PC but is not sure what their hardware can handle. Enter your GPU, VRAM, and RAM, and it returns model profiles sorted into good, better, and best matches for your machine, plus a runtime path to try. Read “good” as the safe, responsive pick you can rely on every day; “better” and “best” push toward larger models that still fit but may run slower or leave less room for long conversations. If you only see one match, that is normal on lower-memory PCs — it means the checker is steering you toward the size that will actually feel usable rather than one that technically loads and then crawls.