Best list · Updated August 2026
Best Small Models for Consumer GPUs
Find small open models for consumer GPUs and local AI workflows, including Phi, Gemma, Mistral, Qwen, and embedding/reranking models.
Editorial review
AI tools, model releases, pricing, licenses, and platform terms can change quickly. Verify the official source before production or commercial use.
Disclosure: OpenSourcesAI may earn a commission when you sign up through partner links. Our listings remain editorial unless specifically labeled as sponsored.
Who this page is for
This page is for local AI users who want a responsive model on hardware they already own. Start with the job, not the parameter count: chat, coding, retrieval, and reranking have different quality and memory needs. Compare quantized memory use, context overhead, response speed, and task accuracy before buying hardware or stretching to the largest model that can barely load.
Selection criteria
- A quantization that fits with enough memory left for context, cache, and the runtime.
- Task quality measured on the prompts, documents, or code the model will actually handle.
- Interactive speed and stability under repeated requests rather than a successful load alone.
- Runtime and model-format support verified for the operating system and available accelerator.
- A clear reason to move larger when a smaller model fails the defined evaluation.
Top picks
- Phi-4 Mini
- Mistral Small 3.1
- Gemma 3 27B
- Multilingual E5 Large
- BGE Reranker v2 M3
Grouped recommendations
Best tiny/edge candidate
Phi-4 Mini
Best medium chat candidates
Mistral Small 3.1, Gemma 3 27B
Best retrieval helpers
Multilingual E5 Large, BGE Reranker v2 M3
How to choose
For consumer GPUs, start smaller than you think. Reliable speed beats a huge model that barely fits.
Related links
FAQ
Is the largest model that fits my GPU always the best choice?
No. A model that leaves almost no memory for context or cache can be slow or unstable, while a smaller model may finish more iterations and fit a longer working context. Compare both on the same tasks and include response time in the result.
Why does model memory change with context length?
The loaded weights are only part of memory use. The runtime also needs working memory and a key-value cache that grows with context and request settings. Leave headroom rather than treating the weight estimate as the entire requirement.
Do embedding and reranking models need the same GPU tier as chat models?
Usually they serve a different role and can be much smaller than the answer-generating model. Evaluate retrieval quality and latency separately instead of choosing them from a chat-model size ladder.
Related resources
Continue comparing tools, models, stacks, and guides related to this category.
Sources
Sponsorship note
Built an AI tool or open-source project? Submit it for review or sponsor a featured placement on OpenSourcesAI.
Sponsor or submit