Local coding

Fit table (registry × catalog)

Model Family SWE-Verified Aider LiveCodeBench HumanEval Terminal-Bench Q4 ~GB API 100K+20K Coding focus
Qwen: Qwen3.6 27B

Qwen HF card (vendor scaffold). SWE/LiveCode are agentic/vendor numbers—treat as upper-bound orientation.

qwen3.6 77.2% 83.9% 59.3% 17.1 $0.13 Yes
Qwen: Qwen3.5-27B

From Qwen3.6-27B HF comparison table (vendor). Same scaffold caveats as 3.6.

qwen3.5 75% 80.7% 41.6% 16.5 $0.05 Yes
Qwen: Qwen3 Coder 30B A3B Instruct

SWE ~51.6% vendor OpenHands (aggregators sometimes cite ~60%). HumanEval 93% (EvalPlus/vLLM report). LiveCodeBench ~40.3% aggregator.

qwen3-coder 51.6% 40.3% 93% 15.2% 18 $0.01 Yes
Qwen2.5 Coder 32B Instruct

Qwen blog / public tables. Aider 73.7 = classic edit (vendor); Aider polyglot whole ≈16.4% on official YAML—different harness.

qwen2.5-coder 50% 73.7% 66% 92.7% 18.5 $0.09 Yes
Mistral: Codestral 2508

Aggregator SWE/HE/LCB; Aider polyglot 11.1% is Codestral 25.01 (closest public row to 2508).

codestral 40% 11.1% 37.9% 90% 13 $0.05 Yes
Qwen: Qwen3 32B

SWE ~30% aggregator; Aider polyglot 40%; LiveCodeBench v5 thinking ~65.7%; HumanEval+ ~79% aggregator.

qwen3 30% 40% 65.7% 79% 18.5 $0.01 General
Meta: Llama 3.3 70B Instruct

HE from Meta; SWE/LCB aggregators; Aider 59.4% = classic code-edit leaderboard (not polyglot).

llama3.3 22% 59.4% 26% 88.4% 40 $0.02 General
Google: Gemma 3 27B

Aider polyglot 4.9%; LiveCodeBench ~14% aggregator; HumanEval 87.8% aggregator (Google PT card HE is lower ~48.8—IT figures vary).

gemma3 4.9% 14% 87.8% 16 $0.02 General
DeepSeek: R1 Distill Llama 70B

DeepSeek distill eval table LiveCodeBench 57.5. Parent R1 SWE ~49% is NOT copied onto the distill row.

deepseek-r1-distill 57.5% 40 $0.10 Yes
IBM: Granite 4.1 8B

IBM Granite 4.1 instruct HumanEval 87.2. No curated public SWE-Verified/Aider/LCB for this desk SKU yet.

granite 87.2% 5 $0.007 General
Meta: Llama 3.1 8B Instruct

Meta model card HumanEval 72.6. Entry-memory general instruct—not SWE-proven.

llama3.1 72.6% 5 $0.007 General
Mistral: Ministral 3 14B 2512

Instruct-2512 LiveCodeBench ~35% (aggregator). Reasoning-2512 sibling reports LCB 64.6%—do not confuse checkpoints.

ministral 35.1% 4.5% 9 $0.02 General
Qwen: Qwen3 8B

LiveCodeBench ~20% (non-reasoning aggregator). HumanEval ~71% public eval reports. Thinking mode scores higher—row matches general instruct.

qwen3 20.2% 71% 2.3% 5.2 $0.02 General
Qwen: Qwen3 14B

LiveCodeBench ~59.3% (Thinking variant comparisons); HumanEval ~85% aggregator. Not a dedicated coder.

qwen3 59.3% 85% 9 $0.02 General

HumanEval alone is not enough for agentic coding. Prefer SWE-Verified + Aider when present; null cells mean we refused to invent a number.

Scores are hand-curated references (vendor blogs / HF cards / Aider YAML / aggregators), not AI Hippo harness re-runs. SWE-Verified swings with agent scaffold. Prefer SWE + Aider over HumanEval alone for agentic work.

API or rack—not a desk download

These may look like “coding models” on aggregators but are not consumer-card downloads.

← Local coding hub

Assumptions and sources

Hardware bands are classes with USD reference ranges (not street quotes). VRAM estimates use curated GGUF footprints plus size-class KV factors. Coding scores are hand-curated SWE/Aider/LiveCode/HumanEval references—not harness re-runs. API job cost uses OpenRouter primary listing when present.

Hand-curated from vendor cards, blogs, and public aggregators (incl. Aider polyglot YAML). Not AI Hippo harness re-runs. Columns: SWE-Verified, Aider, LiveCodeBench, HumanEval, Terminal-Bench. SWE/Terminal swing with agent scaffold. Aider prefers polyglot % when available; classic edit noted in scoreNote.