INFRASTRUCTURE NOTE: Running unquantized weights requires a high-VRAM server or multi-GPU rig. For local workstation setups, select an architecture profile to isolate low-overhead distillation filters or Q4_K_M GGUF quantization paths.

Models · Source-aware through June 2026

Open-weight and open-source AI models

Search and filter model families and featured models by task, hardware, workflow, source status, license posture, and practical deployment fit. OpenSourcesAI separates open-source software projects from open-weight model releases so builders can check the exact model card before production or commercial use.

14Featured models
15Model families

Openness labeled clearly

Open-weight and open-source status is separated so you can verify the exact license before production use.

Hardware fit first

Every model surfaces VRAM and runtime needs alongside the description.

Decision-focused filters

Narrow by task, capability, or openness before comparing model cards.

Click a family hub to browse variants, local notes, comparison tables, source links, and related tools.

Start here

Choose models by fit, not hype

The fastest path is hardware fit first, then task fit, then a small test before committing to a stack.

Showing 14 featured models and 15 family hubs.

Browse by model type

Filter by openness and capability

License note: model openness varies by checkpoint. Treat directory labels as a starting point and verify the official model card, license file, provider terms, redistribution rights, hosted-use limits, and derivative-use rules before commercial deployment.

Featured 2026

Featured model profiles

Reasoning

DeepSeek-V4-Pro

Open weightsMIT

Frontier open-weight DeepSeek model positioned for reasoning-heavy coding, agent, and long-context work.

Best for: Teams comparing frontier-style open-weight reasoning and coding models against hosted closed models.

ParametersMoE / variesContextUp to 1M in DeepSeek docs; verify exact releaseHardwareServer-classRuntimevLLM, SGLang, Transformers, hosted providers

DeepSeek · DeepSeek

Verified/updated 2026

Review model →

Chat

Qwen3 235B A22B

Open weightsApache 2.0

Flagship open-weight Qwen3 MoE model often chosen for serious reasoning, coding, multilingual work, and agent experiments.

Best for: Builders testing frontier-style open-weight reasoning and coding in hosted or multi-GPU environments.

Parameters235BContext40,960 tokensHardwareServer-classRuntimevLLM, SGLang, Transformers, hosted providers

Alibaba Qwen · Qwen

Verified/updated 2026

Review model →

Agents

Kimi K2.6

Open weights where releasedModified MIT License

Current Kimi family model for agentic coding, tool-use behavior, and long-context workflow evaluation.

Best for: Builders testing agentic coding, tool calling, long-context planning, and workflow automation.

ParametersMoE / variesContext262,144 tokens, confirmed from the model's published configHardwareServer-classRuntimevLLM, SGLang, hosted providers

Moonshot AI · Kimi

Verified/updated 2026

Review model →

Code

GLM-5.2

Open weightsMIT

GLM-5.2 is a 753B MoE model requiring multi-GPU or distributed inference for full deployment.

Best for: GLM-5.2 is Z.ai's large mixture-of-experts model, aimed at coding, agentic tool-use, and very long context work rather than single-GPU home setups.

Parameters753BContextCheck model cardHardwareGLM-5.2 is a 753B MoE model requiring multi-GPU or distributed inference for full deployment. Ollama publishes a cloud-hosted entry only (`glm-5.2:cloud`, re-checked 2026-08-24) and exposes no downloadable local artifact, so this catalog records no local tag. Below data-center scale, the unsloth/GLM-5.2-GGUF 1-bit UD-IQ1 variants (~217–228 GB) via llama.cpp or Unsloth Studio with KTransformers are the only local execution path, at reduced fidelity. Not consumer-local hardware friendly.RuntimeOllama / LM Studio / vLLM varies

Z.ai (Zhipu AI) · GLM

Check exact model card

Review model →

Reasoning

MiMo-V2.5-Pro

Open weights where releasedMIT

Xiaomi MiMo frontier model for reasoning, coding, and long-context AI application testing.

Best for: Teams comparing newer Chinese open-weight frontier models for reasoning, code, and long-context tasks.

ParametersMoE / variesContextUp to 1M reported for Pro-class releases; verify exact model cardHardwareServer-classRuntimeHosted providers, vLLM where supported, Transformers

Xiaomi · MiMo

Verified/updated 2026

Review model →

Chat

Gemma 4

Open weightsApache 2.0

Google Gemma family entry for open-weight testing across efficient local, app, and multimodal workflows.

Best for: Developers evaluating Google-backed open-weight models for efficient local apps, hosted prototypes, and multimodal workflows where supported.

ParametersCheck model cardContextCheck current Gemma model cardHardwareVaries by sizeRuntimeOllama, LM Studio, llama.cpp, Transformers, vLLM

Google · Gemma

Verified/updated 2026

Review model →

Reasoning

MiniMax M3

Open weightsMiniMax Community License

MiniMax's current flagship: a 427B-parameter sparse mixture-of-experts model (MiniMaxM3SparseForConditionalGeneration) accepting both image and text input, with a 1,048,576-token (~1M) context window.

Best for: Teams evaluating frontier-scale open-weight reasoning/coding models with genuine multimodal (image) input and a ~1M-token context window, willing to work within a named commercial license rather than a standard OSI one.

ParametersMoE / variesContext1,048,576 tokens (~1M), confirmed from the model's published configHardwareServer-class (multi-GPU)RuntimevLLM, SGLang, or a hosted provider

MiniMax · MiniMax

Verified/updated 2026

Review model →

Multimodal

Llama 4 Scout

Open weightsLlama 4 Community License

Open-weight Llama 4 model positioned for multimodal and long-context workflows.

Best for: Teams evaluating Llama-family models for multimodal assistant, long-context, and application workflows.

ParametersCheck model cardContext10,485,760 tokensHardware~63 GB at Q4_K_M (109B total parameters, MoE) — multi-GPU classRuntimevLLM, Transformers, hosted providers where supported

Meta · Llama

Verified/updated 2026

Review model →

Multimodal

Llama 4 Maverick

Open weightsLlama 4 Community License

Open-weight Llama 4 model positioned for multimodal, reasoning, and general assistant workflows.

Best for: Builders comparing current Llama-family models for assistant, multimodal, and reasoning-oriented workflows.

ParametersCheck model cardContext1,048,576 tokens (~1M), confirmed from the model's published configHardware~243 GB at Q4_K_M (401.6B total parameters, 128-expert MoE) — multi-GPU or server classRuntimevLLM, Transformers, hosted providers where supported

Meta · Llama

Verified/updated 2026

Review model →

Chat

gpt-oss-120b

Open weightsApache 2.0

Larger gpt-oss model for users evaluating high-capacity open-weight deployments.

Best for: Teams evaluating high-capacity local, self-hosted, or developer-controlled model deployments.

Parameters120bContext131,072 tokens, confirmed from the model's published configHardwareServer-classRuntimevLLM, Transformers, hosted providers where supported

OpenAI · gpt-oss

Verified/updated 2026

Review model →

Chat

gpt-oss-20b

Open weightsApache 2.0

Smaller gpt-oss model for local and more accessible open-weight deployments.

Best for: Builders evaluating more accessible open-weight deployments for local apps, prototypes, and controlled workflows.

Parameters20bContext128,000 tokensHardware~12.5 GB at Q4_K_M (20B parameters)RuntimeOllama or LM Studio where supported, vLLM, Transformers

OpenAI · gpt-oss

Verified/updated 2026

Review model →

Audio

Whisper Large V3

Open source weights and codeApache 2.0

Open speech recognition model commonly used for transcription and multilingual audio workflows.

Best for: Builders adding local transcription, podcast processing, meeting notes, or audio translation.

ParametersCheck model cardContext30-second audio windows · 448-token cap per windowHardware~3.1 GB in fp16 (1.5B parameters)RuntimeTransformers, faster-whisper, whisper.cpp

OpenAI · Whisper

Verified/updated 2026

Review model →

Embedding

bge-m3

Open weightsMIT

Multilingual BGE embedding model that supports dense retrieval, lexical retrieval, and multi-vector retrieval in one checkpoint.

Best for: RAG teams that want one embedding model for multilingual search, long documents, and retrieval experiments beyond plain dense vectors.

ParametersCheck model cardContext8,192 tokensHardware~2.27 GB in fp32 (568M parameters)RuntimeOllama or LM Studio where supported, llama.cpp, Transformers, vLLM

BAAI · BGE

Verified/updated 2026

Review model →

Embedding

e5-mistral-7b-instruct

Open weightsMIT

Instruction-tuned E5 embedding model built on a Mistral 7B backbone for text embedding and retrieval tasks.

Best for: Teams comparing larger instruction-tuned embedding models for retrieval quality when smaller E5 checkpoints are not enough.

Parameters7bContext32,768 tokensHardware~14.2 GB in fp16 (7.1B parameters) — an LLM-scale embedding modelRuntimeOllama or LM Studio where supported, llama.cpp, Transformers, vLLM

Microsoft / intfloat · E5

Verified/updated 2026

Review model →

Browse by family

Model family hubs

BAAI

BGE

EmbeddingRerankingRAG

BGE models cover embeddings, reranking, retrieval, semantic search, and vector database workflows for RAG builders.

BAAI · Family hub

Family hub · Source-aware 2026

Family hub →

DeepSeek

DeepSeek

ReasoningCodeDistilled

DeepSeek models cover reasoning, coding, distilled local variants, and open-weight assistant workflows.

DeepSeek · Family hub

Family hub · Source-aware 2026

Family hub →

Microsoft / intfloat

E5

EmbeddingRAGMultilingual

E5 models are widely used for multilingual embeddings, semantic search, retrieval, low-overhead indexing, and RAG pipelines.

Microsoft / intfloat · Family hub

Family hub · Source-aware 2026

Family hub →

Google

Gemma

ChatLocalEfficient

Gemma models are useful for efficient local, app, and multimodal workflows, with small-to-mid-size variants that are practical for developers.

Google · Family hub

Family hub · Source-aware 2026

Family hub →

Z.ai

GLM

AgentsCodeReasoning

GLM models from Z.ai are used for agentic engineering, tool use, coding, and reasoning workflows.

Z.ai · Family hub

Family hub · Source-aware 2026

Family hub →

OpenAI

gpt-oss

Open weightsLocal AIReasoning

OpenAI gpt-oss models are open-weight reasoning models for local, self-hosted, and developer-controlled AI workflows.

OpenAI · Family hub

Family hub · Source-aware 2026

Family hub →

Moonshot AI

Kimi

AgentsCodeReasoning

Kimi models from Moonshot AI focus on agentic coding, tool use, long-context reasoning, and workflow automation.

Moonshot AI · Family hub

Family hub · Source-aware 2026

Family hub →

Meta

Llama

ChatMultimodalLocal

Meta Llama models are widely supported open-weight options for local AI stacks, multimodal workflows, assistant prototypes, and guardrail experiments.

Meta · Family hub

Family hub · Source-aware 2026

Family hub →

Xiaomi

MiMo

ReasoningCodeAgents

Xiaomi's MiMo family is a newer open-weight model family for reasoning, coding, long-context, and agent experiments.

Xiaomi · Family hub

Family hub · Source-aware 2026

Family hub →

MiniMax

MiniMax

ReasoningAgentsCode

MiniMax models are evaluated for long-context reasoning, coding, tool use, agents, and productivity workflows.

MiniMax · Family hub

Family hub · Source-aware 2026

Family hub →

Mistral AI

Mistral

ChatCodeVision

Mistral models cover multilingual workflows, coding, MoE architectures, vision-language experiments, and efficient local or hosted deployments.

Mistral AI · Family hub

Family hub · Source-aware 2026

Family hub →

Meta

Muse

Open weightsLocal AIAgents

Muse is the model family from Meta Superintelligence Labs: Muse Glimmer is the open-weight, locally runnable member, distilled from Muse Spark, the hosted flagship.

Meta · Family hub

Family hub · Source-aware 2026

Family hub →

Microsoft

Phi

EdgeSmallVision

Microsoft Phi models focus on small language models, edge deployment, reasoning efficiency, and low-resource local experiments.

Microsoft · Family hub

Family hub · Source-aware 2026

Family hub →

Alibaba Qwen

Qwen

ChatCodeReasoning

Qwen models are strong choices for multilingual chat, coding, math, vision-language, and local developer workflows.

Alibaba Qwen · Family hub

Family hub · Source-aware 2026

Family hub →

OpenAI

Whisper

AudioTranscriptionSpeech recognition

Whisper models are used for ASR, transcription, subtitles, podcast processing, meeting notes, and multilingual audio.

OpenAI · Family hub

Family hub · Source-aware 2026

Family hub →

Next step

Turn a model choice into a working stack.

After you shortlist a model, test prompts in the Playground and use stack recipes to connect it to a chat UI, RAG workflow, coding assistant, or serving layer.

Local model VRAM reference

Consumer-runnable models by parameter tier

All models run locally via Ollama or llama.cpp. VRAM values are empirical GGUF measurements, not formula estimates. Enter your GPU VRAM below to highlight which quantization fits your hardware.

≤5B

Compact

Fits in 4 GB VRAM at Q4. Fast responses on entry GPU or CPU.
ModelParamsContextQ4_K_MQ8_0FP16License
1.78B
128K
Q41.2 GB
Q82 GB
FP163.6 GB
MIT
VibeThinker-3B
ReasoningmathCoding
3B
64K
Q42.4 GB
Q83.8 GB
FP16
MIT
Phi-3 Mini 128K
ChatCodingedgeSummarisation
3.8B
128K
Q44 GB
Q86 GB
FP168 GB
MIT
Phi-4 Mini
ChatReasoningmath
3.84B
128K
Q42.6 GB
Q84.2 GB
FP167.5 GB
MIT
Qwen3 4B
ChatCodingSummarisation
4B
32K
Q42.6 GB
Q84.5 GB
FP168 GB
Apache 2.0
Gemma 3 4B
ChatSummarisationRAG
4B
128K
Q42.5 GB
Q84.3 GB
FP168 GB
Gemma Terms of Use
6–20B

Standard

Sweet spot for 8–16 GB VRAM. Best quality-to-cost ratio.
ModelParamsContextQ4_K_MQ8_0FP16License
6.7B
16K
Q44.4 GB
Q87.4 GB
FP1613.5 GB
DeepSeek License
Qwen2.5 7B
ChatCodingReasoningmultilingual
7B
128K
Q46 GB
Q89 GB
FP1616 GB
Apache 2.0
Qwen2.5-Coder 7B
CodingReasoningChatAgents
7B
128K
Q46 GB
Q89 GB
FP1616 GB
Apache 2.0
Mathstral 7B
mathReasoning
7.25B
32K
Q44.7 GB
Q88 GB
FP1614 GB
Apache 2.0
Phi-3 Small
ChatReasoning
7.4B
8K
Q44.8 GB
Q88.1 GB
FP1614.5 GB
MIT
Llama 3 8B
ChatCodingRAGSummarisation
8B
8K
Q45.6 GB
Q89.5 GB
FP1616 GB
Llama 3 Community License
Llama 3.1 8B Instruct
ChatCodingRAGAgentsSummarisation
8B
128K
Q45.6 GB
Q89.5 GB
FP1616 GB
Llama 3.1 Community License
Qwen3 8B
CodingChatRAGAgentsReasoning
8B
32K
Q45.3 GB
Q88.9 GB
FP1616 GB
Apache 2.0
DeepSeek-R1 8B (0528 Qwen3)
ReasoningCodingChat
8B
32K
Q45 GB
Q88.8 GB
FP1616 GB
MIT
Ministral 8B
ChatSummarisation
8.02B
128K
Q45.2 GB
Q88.8 GB
FP1616 GB
Mistral Research License (non-commercial)
8.03B
128K
Q45.2 GB
Q88.8 GB
FP1616 GB
MIT
8.54B
8K
Q45.6 GB
Q89.4 GB
FP1617 GB
Gemma Terms of Use
Qwythos-9B (Claude-Mythos-5 1M)
ReasoningCodingtool-uselong-contextChat
9B
1024K
Q46.5 GB
Q811 GB
FP16
Apache-2.0
Gemma 2 9B Instruct
ChatSummarisationRAG
9.24B
8K
Q46 GB
Q810 GB
FP1618 GB
Gemma Terms of Use
Qwen3.8 9B Distill
ReasoningmathCodingtool-uselong-context
9.7B
256K
Q45.4 GB
Q8
FP16
Apache 2.0
Qwen3.5 9B
ChatReasoninglong-contextmultilingual
9.7B
256K
Q46.6 GB
Q811 GB
FP1619 GB
Apache 2.0
Ornith 1.5 9B
CodingAgentsChatlong-context
9.7B
256K
Q45.6 GB
Q8
FP16
MIT
Gemma 3 12B
ChatRAGSummarisationCoding
12B
128K
Q47.6 GB
Q813 GB
FP1624 GB
Gemma Terms of Use
Gemma 4 12B
ChatReasoning
12B
256K
Q47.6 GB
Q812.8 GB
FP1624 GB
Apache 2.0
Gemma 4 12B (QAT)
ChatCodingSummarisationReasoning
12B
256K
Q49.5 GB
Q8
FP16
Apache 2.0
Phi-3 Medium 128K Instruct
ChatCodingSummarisationRAG
14B
128K
Q48.6 GB
Q815 GB
FP1628 GB
MIT
Phi-4
CodingReasoningChatSummarisation
14B
16K
Q49.1 GB
Q815 GB
FP1628 GB
MIT
Qwen3 14B
CodingChatRAGAgentsReasoning
14B
32K
Q49.3 GB
Q815.5 GB
FP1628 GB
Apache 2.0
DeepSeek-R1 Distill 14B
ReasoningCodingAgents
14B
32K
Q49 GB
Q815.5 GB
FP1628 GB
MIT
Qwen2.5 14B
ChatCodingReasoningmultilingualRAG
14B
128K
Q49 GB
Q816 GB
FP1628 GB
Apache 2.0
Qwen2.5-Coder 14B
CodingReasoningAgents
14B
128K
Q49 GB
Q815.7 GB
FP1629.6 GB
Apache 2.0
15.7B
128K
Q410 GB
Q817 GB
FP1631 GB
DeepSeek License
GPT-OSS 20B
ReasoningCodingAgents
20B
125K
Q412.5 GB
Q8
FP16
Apache 2.0
21B+

Large

Needs 20+ GB VRAM at Q4. Workstation or multi-GPU hardware.
ModelParamsContextQ4_K_MQ8_0FP16License
22.2B
32K
Q414 GB
Q824 GB
FP1643 GB
Mistral Non-Production License
Devstral
CodingAgents
23.6B
128K
Q414.3 GB
Q825 GB
FP1647 GB
Apache 2.0
Mistral Small 3.1
ChatCodingRAGReasoningAgentsSummarisation
24B
128K
Q414.4 GB
Q825.5 GB
FP1648 GB
Apache 2.0
Gemma 4 26B (A4B)
ChatReasoninglong-contextmultilingual
26B
256K
Q418 GB
Q8
FP16
Apache 2.0
Gemma 3 27B
ChatCodingRAGSummarisationReasoning
27B
128K
Q417 GB
Q829 GB
FP1656 GB
Gemma Terms of Use
Gemma 2 27B Instruct
ChatReasoningSummarisation
27.2B
8K
Q417 GB
Q829 GB
FP1653 GB
Gemma Terms of Use
Qwen3.6 27B
ChatReasoninglong-contextmultilingual
27.8B
256K
Q417 GB
Q8
FP16
Apache 2.0
Qwen3.8 27B
ChatReasoninglong-contextmultilingual
27.8B
256K
Q418 GB
Q830 GB
FP1656 GB
Apache 2.0
Muse Glimmer 30B
Agentstool-useCodingReasoninglong-contextChat
29.8B
128K
Q418 GB
Q831 GB
FP1657 GB
Apache 2.0
North Mini Code 1.0
CodingAgentstool-useReasoninglong-context
30.5B
488K
Q418 GB
Q8
FP16
Apache 2.0
Qwen3 30B (A3B)
ChatReasoninglong-contextAgents
30.5B
256K
Q419 GB
Q8
FP16
Apache 2.0
Qwen3-Coder 30B (A3B)
CodingAgentstool-uselong-context
30.5B
256K
Q419 GB
Q8
FP16
Apache 2.0
Gemma 4 31B
ChatReasoninglong-contextmultilingual
31B
256K
Q420 GB
Q8
FP16
Apache 2.0
Qwen3 32B
CodingReasoningAgentsChatRAG
32B
32K
Q420 GB
Q834 GB
FP1664 GB
Apache 2.0
DeepSeek-R1 Distill Qwen 32B
ReasoningCodingAgentsChat
32B
32K
Q420 GB
Q834 GB
FP1664 GB
MIT
Qwen2.5 32B
ChatCodingReasoningmultilingualAgents
32B
128K
Q420 GB
Q834 GB
FP1666 GB
Apache 2.0
Qwen2.5-Coder 32B
CodingReasoningAgentsChat
32B
128K
Q420 GB
Q834 GB
FP1665 GB
Apache 2.0
33B
16K
Q420.5 GB
Q835 GB
FP1664 GB
DeepSeek License
Laguna XS 2.1
CodingAgentstool-useReasoninglong-context
33.4B
256K
Q420 GB
Q8
FP16
OpenMDW-1.1
Qwen3.6 35B (A3B)
ChatReasoninglong-contextmultilingual
36B
256K
Q424 GB
Q8
FP16
Apache 2.0
Ornith 1.5 35B (A3B)
CodingAgentsChatlong-context
36B
256K
Q422 GB
Q8
FP16
MIT
Phi-3.5 MoE
ChatReasoning
41.9B
128K
Q425.5 GB
Q844.5 GB
FP1680 GB
MIT
Mixtral 8x7B Instruct
ChatCodingReasoning
46.7B
32K
Q428.5 GB
Q850 GB
FP1691 GB
Apache 2.0
Llama 3 70B
ChatCodingRAGAgentsSummarisationReasoning
70B
8K
Q441.5 GB
Q874 GB
FP16140 GB
Llama 3 Community License
Llama 3.3 70B
ChatCodingRAGAgentsReasoningSummarisation
70B
128K
Q441 GB
Q872 GB
FP16142 GB
Llama 3.3 Community License
70.6B
128K
Q443 GB
Q875 GB
FP16
MIT
Llama 3.1 70B Instruct
ChatReasoningSummarisation
70.6B
128K
Q443 GB
Q875 GB
FP16
Llama 3.1 Community License
Qwen2.5 72B
ChatCodingReasoningmultilingualAgents
72B
128K
Q445 GB
Q880 GB
FP16144 GB
Qwen License
Qwen2.5 Math 72B
mathReasoning
72.7B
4K
Q444 GB
Q877 GB
FP16
Qwen License
Llama 4 Scout
ChatCodingRAGAgentsReasoning
109B
10240K
Q463 GB
Q8
FP16
Llama 4 Community License
Mistral Large 2
ChatReasoningCoding
123B
128K
Q475 GB
Q8131 GB
FP16
Mistral Research License (non-commercial)
Mistral 8x22B Instruct
ChatReasoningCoding
141B
64K
Q486 GB
Q8150 GB
FP16
Apache 2.0
Qwen3-235B-A22B
CodingReasoningRAGlong-context
235B
125K
Q4140 GB
Q8236 GB
FP16
Apache 2.0
Qwen3 235B A22B Thinking
ReasoningAgentslong-context
235B
256K
Q4143 GB
Q8250 GB
FP16
Apache 2.0
GLM-4.5
ChatCodingReasoning
355B
125K
Q4216 GB
Q8381 GB
FP16
MIT
405B
128K
Q4245 GB
Q8430 GB
FP16
Llama 3.1 Community License
DeepSeek V3
ChatReasoningCoding
671B
128K
Q4407 GB
Q8712 GB
FP16
DeepSeek License Agreement (model weights); MIT (code)
GLM-5.2
CodingAgentsReasoninglong-contexttool-use
753B
1024K
Q4466 GB
Q8801 GB
FP16
MIT
DeepSeek-V4-Pro
CodingReasoningAgentslong-context
1600B
977K
Q4900 GB
Q81600 GB
FP16
MIT

Want a ranked shortlist for your specific hardware? The compatibility checker scores every model against your VRAM and RAM using the same engine that powers this table.

OpenSourcesAI tracks open-weight, open-source, local-capable, and hosted AI models. Use the filters to separate model families by openness, deployment fit, and primary workflow.