Best open source coding LLM: 3 picks I’d actually deploy
Qwen leads three deployable open-weight models, but benchmark swings, repeated runs and local hardware decide…
The short version
- Qwen3.6-35B-A3B is the default deployment pick, but coding performance belongs to a model-harness pairing.
- The same Qwen checkpoint scored 65% with Pi and 55% with OpenCode on SWE-Bench Verified.
- Repeated repository trials, end-to-end verification and hardware-aware serving matter more than a single leaderboard score.
The same Qwen checkpoint solved 65% of SWE-Bench Verified with the Pi harness and 55% with OpenCode. A ten-percentage-point swing from changing the software around the model should make every “best open source coding LLM” ranking slightly embarrassing.
My default pick is Qwen3.6-35B-A3B. Kimi K2.7 Code has the more interesting post-training story, while MiMo-V2.6-Distill-Qwen-9B is the compact option I’d test on modest hardware. I’m using “open source” because that’s what humans type into Google; “open-weight” is more accurate for these models.
A beautiful demo with one successful run proves very little. The model chooses actions and writes code, but the harness controls the files, tools, feedback loop and retry policy. Leave that context out and the ranking becomes fantasy football for GPUs.
My three best open source coding LLM picks
In a real coding run, the harness reads the repository, chooses what enters context and gives the model access to file editing, shell commands and Language Server Protocol diagnostics. The model decides what to inspect or change. The harness executes that action, captures the result and feeds it into the next turn. Tests then tell the pair whether the patch fixed the target behavior or broke something nearby. Every retry expands the history and burns more inference time. File selection, tool permissions and test policy can change the result without touching the model weights. I rank pairings I can deploy, because naked weights have never fixed one of my production bugs.

1. Qwen3.6-35B-A3B: the default
Coding Agent Bench tested the Red Hat NVFP4 checkpoint on the full SWE-Bench Verified set. Pi passed 65% of tasks, while OpenCode passed 55% with the same model. The Ansible slice tells the same story: Pi passed 48%, versus 38% for OpenCode with the identical checkpoint. That spread tells me the checkpoint is capable and Pi currently extracts more from it on these benchmarks. I’d start with Pi, reproduce the result on my repositories and keep OpenCode in the bake-off. Nobody has independently shown that these results survive repeated runs, different serving stacks and ordinary production work, and we still lack a clean cost comparison under each buyer’s context length and concurrency. Qwen gets first place in my bake-off; the mitre can stay in Rome.
2. Kimi K2.7 Code: the post-training bet
Kimi earns second place because its reinforcement-learning setup rewarded the fraction of hidden checks passed for the requested fix, then assigned zero reward if existing behavior broke. That discourages the classic agent move of fixing checkout by quietly setting the warehouse on fire. After one epoch of reported RL post-training, its Terminal-Bench pass rate rose from 67% to 82%. The evaluation included harnesses absent from training, so the gain transferred beyond the exact training scaffold. There is a catch: Changdae Oh and his co-authors found that post-training can improve single-shot accuracy while reducing the variety of solutions discovered across repeated attempts. If my workflow generates several candidates, that “sharpening tax” matters. Kimi is the model I’d choose when I can afford a proper repeated evaluation instead of one triumphant screenshot in Slack.
3. MiMo-V2.6-Distill-Qwen-9B: the compact pick
In the vendor evaluation, MiMo’s released supervised checkpoint scored 45% on SWE Pro, compared with 32% for Qwen3.5-9B under the same conditions. On AutomationBench, it reached 30% against 5% for that baseline. Vendor numbers earn a trial; they do not earn my credit card unsupervised. The compact GGUF is why MiMo makes this list: its Q4_K_M file is about 5.8 GB, compared with roughly 18 GB for the listed BF16 weights. Runtime memory still includes the KV cache and any multimodal projector, so download size is only the opening bill. I’d use MiMo for cheap internal experiments, constrained agents and laptop development. Community users have also reported trouble with its supplied llama.cpp template for reasoning and tool calls. That evidence is anecdotal, but it’s enough for me to verify the parser path before blaming the model.

Why the harness can flip the ranking
A leaderboard score belongs to a model-harness pair. I know that sounds annoyingly pedantic until the same weights lose ten percentage points because somebody swapped the scaffold.
Pi and OpenCode can expose different repository context, package tool output differently and make different decisions about retries. Those choices alter the evidence available to the model at every turn. A missed file early in the run can send the agent down a dead end, while timely diagnostics can reveal the correct symbol before it rewrites half the codebase. The harness also decides when tests run and which failure details return to the model. Sampling adds variation even when the visible setup stays fixed. Small differences accumulate across a long tool loop, so the final patches can diverge wildly from the same initial issue. That is how changing the scaffold can move a score as much as changing the model.
The strongest objection is obvious: use enough benchmark tasks and the noise should average out, revealing the best underlying model. Michael Hardy, Ruhana Azam and their co-authors examined agent benchmarks and found rankings were more reliable for fixed model-scaffold systems than for underlying models. Scaffold choice could change the model ranking itself. More tasks help, but they do not magically separate a model from the machinery controlling what it sees and does.
Repeated runs make this even messier. Eduardo Ariño de la Rubia and Szilard Pafka ran 584 coding-agent trials across six agents and six open-weight endpoints. Their fixed pairings varied enough that differences within one pairing exceeded the observed gaps among some pairings. The study covered one machine-learning task, so its scope is narrow. It still ruins a lazy one-run bake-off.
A related study found that a smaller model beat its larger same-family counterpart in 28% of individual runs, even though the larger model had the better pooled average. Its average gain was smaller than one run-to-run standard deviation. I have made purchasing decisions from thinner evidence while holding bad airport coffee. It did not end magnificamente.
For an internal evaluation, I’d lock the harness version and serving configuration, then run each pairing repeatedly on repository tasks my team actually recognizes. I’d record accepted patches alongside rejected attempts, total time to green tests and every rule violation. That still would not tell me which pairing performs best for somebody else’s repositories, languages, permissions, hardware and latency target. It would not prove the benchmark gains transfer to next year’s work either. At least we’d be ignorant professionally.
Clausebench noted:
A later date is more time to do the same amount of work, not less work.
Local serving changes what “fast” means
Model cards become much less romantic once the hardware starts wheezing. On our M3 Max with 128 GB of memory, gpt-oss:20b generated about 74 tokens per second. On the same machine, gpt-oss:120b managed about 51. Both MXFP4 models stayed fully resident in unified memory.
Prompt processing created a wider gap. The smaller model processed about 756 tokens per second, versus 215 for the larger one. Time to first token was roughly 3.7 seconds compared with 5.8 seconds. We measured both on August 25, 2026, using the same setup.
Those prompt numbers matter because a coding agent keeps rebuilding a growing conversation. It starts with an issue, reads files and appends shell or test output. The next request includes much of that history. Without prefix reuse, the server processes shared context again before generating the next action. Longer sessions therefore spend increasing time on prefill, while concurrent agents compete for KV-cache memory. A model with impressive generation speed can still feel miserable across a repository session. I benchmark elapsed time from issue to verified patch, because tokens per second cannot merge a pull request.
Caching can transform that loop. In NVIDIA’s MLPerf Edge Agentic submission, 96% of prompt tokens came from hot cache; without reuse, each turn would have processed the shared history again. The optimized Qwen workload finished 6.4 times faster than the llama.cpp reference, which took more than two and a half hours. That result came from one Jetson setup with a specific TensorRT stack, so extrapolating the speedup to another machine would be silly.
Meanwhile, ComfyUI had effectively annexed my RTX 5060 Ti for self-hosted image generation. With the card’s memory occupied, Ollama had only 150 MB of VRAM available and the language model ran entirely on the CPU. Apparently the GPU believes in work-life balance.
One machine will not suit every local AI workload. Image generation and coding agents fight over the same scarce memory in very different ways. I’d schedule them separately or buy dedicated hardware before pretending the collision disappears through positive thinking.
A green test is where verification starts
The winning patch still has to survive production-shaped tests. SWE-Serve reported a 69% pass rate when end-to-end tests were excluded, then 46% when those tests counted. Roughly one-third of patches that looked correct under narrower checks failed the serving path.
My deployment loop would generate several patches, reject anything that violates policy and then run hidden target checks alongside existing-behavior tests. The sandbox would restrict file access, credentials and network destinations independently of whatever the model says. A separate verifier would inspect surviving candidates, because generators are extremely talented at grading their own homework. End-to-end tests would exercise deployed behavior under production-shaped conditions. I’d retain failed attempts in the evaluation record, since deleting them manufactures reliability. Green tests remain evidence bounded by the suite’s coverage. Formal verification can catch counterexamples that tests miss, although synthesizing faithful specifications is still difficult.
EDPB Deputy Chair Jelena Virant Burnik said:
The new EDPB guidelines are a major step in further aligning how Data Protection Authorities decide whether an administrative fine should be imposed, either on its own or alongside other corrective measures. The GDPR significantly increased the corrective powers of DPAs, with fines serving as an important instrument for effective enforcement. The guidelines reaffirm our commitment to providing greater clarity and ensuring the consistent application of the GDPR across Europe.
Tool authorization deserves the same skepticism. In one MCP evaluation, checking permissions only inside the tool body exposed forbidden tools in 21% of attempts. Permission-aware visibility paired with invocation enforcement produced zero exposures in the reported trials. Hiding a dangerous tool from the menu does little when a scripted client can still order it from the kitchen.
The “local” label provides no magic shield either. Researchers found a saved-conversation-state authorization flaw that succeeded in every controlled attempt against one consumer local-LLM interface. That finding belongs to the specific serving software, not every local deployment. The lesson travels further: prompts can remain vulnerable in memory, wrappers and shared caches after the model file lands on your desk.
Serious coding leaderboards will eventually publish the model file, quantization, serving runtime, harness version and a distribution across repeated runs. Anything still selling one naked score will look like a restaurant proudly reviewing the flour.
Frequently asked questions
What is the best open source coding LLM to deploy?
Qwen3.6-35B-A3B is the strongest default pick among the three models examined. On SWE-Bench Verified, the same Red Hat NVFP4 checkpoint passed 65% of tasks with Pi and 55% with OpenCode, demonstrating that deployment choice must include both the model and its coding harness.
How much does the coding harness affect LLM performance?
A coding harness controls repository context, file editing, shell commands, diagnostics, retries and test feedback. Those choices change the evidence available at every turn, so identical model weights can produce different patches and scores. In the cited Qwen test, switching harnesses changed the SWE-Bench Verified result by ten percentage points.
Can coding LLMs and self-hosted image generation AI share one GPU?
Coding LLMs and self-hosted image generation AI can share a machine, but they compete for scarce GPU memory. When ComfyUI occupied the RTX 5060 Ti, Ollama had only 150 MB of VRAM available and ran the language model entirely on the CPU. Separate scheduling or dedicated hardware avoids that collision.
Sources
- Qwen3.8-Omni: Towards Native Omni-Modal Agents
- Use open weight models as your AI coding agent with Amazon Bedrock
- Contrastive Language Models
- Stanford and Nvidia's open CLM-8B caches reusable agent actions and runs up to 9x faster than Jev in tests
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
- Nvidia's OpenShell controls what AI agents can access, even when they ignore instructions