Best list · Updated August 2026
Best Open-Weight Models for Coding
Compare open-weight coding model candidates for Continue, Aider, Cline, Kilo Code, and other coding assistant workflows.
Editorial review
AI tools, model releases, pricing, licenses, and platform terms can change quickly. Verify the official source before production or commercial use.
Who this page is for
This page is for developers choosing a model for code completion, repo-aware chat, patch generation, or agent-style coding. Treat the shortlist as an evaluation set, not a universal ranking: the right model depends on your languages, repository size, serving hardware, acceptable latency, and whether the workflow needs reliable tool use rather than plausible-looking code alone.
Selection criteria
- Published model weights and licensing terms that can be reviewed before deployment.
- Useful fit for completion, code editing, repository questions, or tool-driven coding tasks.
- Context capacity and prompt behavior suited to the files and diffs in your evaluation set.
- A realistic local or controlled-server path for the model size you intend to run.
- Quality measured on real repository tasks, tests, review comments, and accepted diffs.
Top picks
- Qwen3 Coder
- DeepSeek Coder V2
- DeepSeek R1
- Kimi K2
- GLM-4.5
Grouped recommendations
Best coding families to test
Qwen3 Coder, DeepSeek Coder V2, DeepSeek R1
Best for agentic coding tests
Kimi K2, GLM-4.5
Best local-friendly path
Smaller Qwen, DeepSeek, Mistral, or Phi variants
How to choose
Run a repo-specific eval. Public coding benchmarks change quickly and may not match your codebase.
Related links
FAQ
Which coding model should I run locally first?
Start with the smallest credible coding model that fits comfortably in your available memory, then compare it with one stronger candidate on the same repository tasks. A model that responds quickly enough for repeated review often produces a better working loop than a larger model that barely fits.
Do coding benchmarks identify the best model for my repository?
No single benchmark captures your languages, build system, tests, coding conventions, or preferred assistant. Use public benchmarks to form a shortlist, then grade the generated patches, test results, unnecessary edits, and review effort on tasks from your own codebase.
Does every open-weight coding model work with every coding assistant?
No. Check the assistant provider format, serving API, context requirements, and tool-calling expectations before choosing a model. A model can be strong at code generation yet still be a poor fit for an agent that depends on structured tool calls.
Related resources
Continue comparing tools, models, stacks, and guides related to this category.
Sources
Sponsorship note
Built an AI tool or open-source project? Submit it for review or sponsor a featured placement on OpenSourcesAI.
Sponsor or submit