本地编码
能下载的编码模型
仅 allowlist。按编码基准(先 SWE-Verified)排序——站内综合榜不是编码分。
适配表(registry × catalog)
| 模型 | 家族 | SWE-Verified | Aider | LiveCodeBench | HumanEval | Terminal-Bench | Q4 ~GB | API 100K+20K | 编码向 |
|---|---|---|---|---|---|---|---|---|---|
| Qwen: Qwen3.6 27B Qwen HF card (vendor scaffold). SWE/LiveCode are agentic/vendor numbers—treat as upper-bound orientation. | qwen3.6 | 77.2% | — | 83.9% | — | 59.3% | 17.1 | $0.13 | 是 |
| Qwen: Qwen3.5-27B From Qwen3.6-27B HF comparison table (vendor). Same scaffold caveats as 3.6. | qwen3.5 | 75% | — | 80.7% | — | 41.6% | 16.5 | $0.05 | 是 |
| Qwen: Qwen3 Coder 30B A3B Instruct SWE ~51.6% vendor OpenHands (aggregators sometimes cite ~60%). HumanEval 93% (EvalPlus/vLLM report). LiveCodeBench ~40.3% aggregator. | qwen3-coder | 51.6% | — | 40.3% | 93% | 15.2% | 18 | $0.01 | 是 |
| Qwen2.5 Coder 32B Instruct Qwen blog / public tables. Aider 73.7 = classic edit (vendor); Aider polyglot whole ≈16.4% on official YAML—different harness. | qwen2.5-coder | 50% | 73.7% | 66% | 92.7% | — | 18.5 | $0.09 | 是 |
| Mistral: Codestral 2508 Aggregator SWE/HE/LCB; Aider polyglot 11.1% is Codestral 25.01 (closest public row to 2508). | codestral | 40% | 11.1% | 37.9% | 90% | — | 13 | $0.05 | 是 |
| Qwen: Qwen3 32B SWE ~30% aggregator; Aider polyglot 40%; LiveCodeBench v5 thinking ~65.7%; HumanEval+ ~79% aggregator. | qwen3 | 30% | 40% | 65.7% | 79% | — | 18.5 | $0.01 | 通用 |
| Meta: Llama 3.3 70B Instruct HE from Meta; SWE/LCB aggregators; Aider 59.4% = classic code-edit leaderboard (not polyglot). | llama3.3 | 22% | 59.4% | 26% | 88.4% | — | 40 | $0.02 | 通用 |
| Google: Gemma 3 27B Aider polyglot 4.9%; LiveCodeBench ~14% aggregator; HumanEval 87.8% aggregator (Google PT card HE is lower ~48.8—IT figures vary). | gemma3 | — | 4.9% | 14% | 87.8% | — | 16 | $0.02 | 通用 |
| DeepSeek: R1 Distill Llama 70B DeepSeek distill eval table LiveCodeBench 57.5. Parent R1 SWE ~49% is NOT copied onto the distill row. | deepseek-r1-distill | — | — | 57.5% | — | — | 40 | $0.10 | 是 |
| IBM: Granite 4.1 8B IBM Granite 4.1 instruct HumanEval 87.2. No curated public SWE-Verified/Aider/LCB for this desk SKU yet. | granite | — | — | — | 87.2% | — | 5 | $0.007 | 通用 |
| Meta: Llama 3.1 8B Instruct Meta model card HumanEval 72.6. Entry-memory general instruct—not SWE-proven. | llama3.1 | — | — | — | 72.6% | — | 5 | $0.007 | 通用 |
| Mistral: Ministral 3 14B 2512 Instruct-2512 LiveCodeBench ~35% (aggregator). Reasoning-2512 sibling reports LCB 64.6%—do not confuse checkpoints. | ministral | — | — | 35.1% | — | 4.5% | 9 | $0.02 | 通用 |
| Qwen: Qwen3 8B LiveCodeBench ~20% (non-reasoning aggregator). HumanEval ~71% public eval reports. Thinking mode scores higher—row matches general instruct. | qwen3 | — | — | 20.2% | 71% | 2.3% | 5.2 | $0.02 | 通用 |
| Qwen: Qwen3 14B LiveCodeBench ~59.3% (Thinking variant comparisons); HumanEval ~85% aggregator. Not a dedicated coder. | qwen3 | — | — | 59.3% | 85% | — | 9 | $0.02 | 通用 |
单看 HumanEval 不足以支撑智能体编码。有 SWE-Verified / Aider 时优先看它们;空单元格表示我们拒绝编造数字。
分数为手写参考(厂商博文 / HF 卡 / Aider YAML / 聚合榜),非 AI Hippo harness 复现。SWE-Verified 随脚手架波动大。智能体场景优先看 SWE + Aider,不要只看 HumanEval。
API 或机柜——不是书桌下载
聚合器上可能看起来像「编码模型」,但不是消费卡能下的体积。
- Qwen: Qwen3 Coder 480B A35B — ~480B MoE coder — 付 API 或上机柜。 $0.05
- DeepSeek: DeepSeek V4 Flash 0423 — ~284B MoE — API 便宜,不是 24GB 下载。 $0.01
- DeepSeek: DeepSeek V4 Flash 0731 — ~284B MoE — API 便宜,不是 24GB 下载。 $0.02
假设与来源
硬件档是「类」+ 美元参考区间(非街价报价)。显存估算用 curated GGUF 体积 + 按规模的 KV 系数。编码分为手写 SWE/Aider/LiveCode/HumanEval 参考——非 harness 复现。API 任务成本取 OpenRouter 主 listing(若有)。
Hand-curated from vendor cards, blogs, and public aggregators (incl. Aider polyglot YAML). Not AI Hippo harness re-runs. Columns: SWE-Verified, Aider, LiveCodeBench, HumanEval, Terminal-Bench. SWE/Terminal swing with agent scaffold. Aider prefers polyglot % when available; classic edit noted in scoreNote.