Qwen3.8-27B: a 24GB consumer GPU against Opus 4.6’s coding table
Apache-2.0 dense 27.78B. A third-party Q4 (~17GB) fits a 24GB consumer GPU. OpenRouter primary listing $0.45/$3.20. Qwen’s launch table (vendor-reported) has SWE-bench Pro 61.7 vs Opus 4.6 Max’s imported 53.4; Terminal-Bench 73.0 vs 78.2. Same 100K+20K job: local ~$0, hosted ~$0.11, Opus $1.00. Not an independent benchmark.
Enlarge image What it is, and why this table should startle you
Qwen3.8-27B is the Apache-2.0 dense multimodal checkpoint published 2026-08-14: 27,781,427,952 parameters on Hugging Face, text/image/video in, text out, native 262K. It is not Qwen3-8B, and it is not API-only Qwen3.8-Max.
The shock is the floor, not “another 27B.” A third-party Q4_K_M file is about 17.1GB — it fits a 24GB consumer GPU. You can boot a coding agent on a desk tonight. The peer on the other side of that sentence is Claude Opus 4.6: OpenRouter primary listing $5/$25, weights not downloadable.
The four bars are Qwen’s own launch table. 3.8-27B leads the in-table Opus 4.6 Max on SWE-Pro, LiveCodeBench, and OSWorld; it trails by 5.2 on Terminal-Bench. That is a marketing harness, not an independent reproduction. It is still the comparison this page exists to nail down.
Spec sheet, side by side
Put the numbers you can check on the table first. All four columns are on-site catalog plus OpenRouter primary listings (`qwen/qwen3.8-27b` is now $0.45/$3.20).
| 3.8-27B | 3.6-27B | 3.8-Max | Opus 4.6 | |
|---|---|---|---|---|
| Released | 2026-08-14 (HF) | catalog first_seen 2026-06-22 | catalog first_seen 2026-08-04 | in catalog |
| Params | 27.78B dense | ~27.8B dense | flagship 3.8 API | frontier API |
| Weights | Apache-2.0 download | Apache-2.0 + API | API only | API only |
| Context | 262K native / 1M YaRN | 262K | Best in this comparison 1M | Best in this comparison 1M |
| Modality | text+image+video→text | text+image+video→text | text+image+video→text | multimodal API |
| OpenRouter $/1M | $0.45 / $3.20 | Best in this comparison $0.289 / $2.40 | $2.00 / $6.00 | $5.00 / $25.00 |
The published decoder for 3.8-27B and 3.6-27B is essentially the same shape (64-layer hybrid attention, hidden 5120). The coding jump is a training story, not a new skeleton. Max and Opus are a different product line: hosted 1M, billed per token, you never own the weights. 3.8-27B has both doors: downloadable local weights, and a hosted API on the same checkpoint.
Same job: local ~$0, hosted ~$0.11, Opus $1
Enlarge image Hold the job fixed at 100K in + 20K out and recompute from on-site OpenRouter primary listings. This is not a capability score. It is who you pay for the same token work.
| Model | In $/1M | Out $/1M | 100K+20K | Role |
|---|---|---|---|---|
| 3.8-27B local | — | — | Best in this comparison ~$0 (electricity) | 24GB Q4 desk |
| Qwen3 Coder 30B A3B | Best in this comparison $0.07 | Best in this comparison $0.27 | $0.012 | cheap code API |
| Qwen3 Coder 480B | $0.30 | $1.00 | $0.050 | hosted code flagship |
| Qwen3.6-27B API | $0.289 | $2.40 | $0.077 | same-size predecessor |
| Qwen3.8-27B API | $0.45 | $3.20 | $0.109 | same weights, hosted |
| Qwen3.8-Max | $2.00 | $6.00 | $0.32 | same-gen cloud flagship |
| Claude Opus 4.6 | $5.00 | $25.00 | $1.00 | frontier coding API |
| GPT-5.6 Sol | $5.00 | $30.00 | $1.10 | OpenAI coding flagship |
Read it this way: Opus 4.6 and GPT-5.6 Sol bill about $1 for that job. Same-generation Max is about $0.32. Hosted 3.8-27B is about $0.11 (a bit above predecessor 3.6-27B at $0.077). Hosted Coder flagship is about $0.05. The same weights on your own 24GB card have a token sticker of zero — you pay electricity and depreciation. DeepSeek V4 Flash 0731 is the other extreme: 284B MoE at $0.09/$0.18, cheap as an API, not a consumer-card download.
Coding: versus Opus 4.6, not just the family
Qwen’s launch table puts 3.8-27B, predecessor 3.6-27B, hosted 3.7-Plus, and Claude Opus 4.6 Max on one sheet. The copy below is that sheet. Attribution: vendor-reported; several rows used a Claude Code harness; the Opus SWE-bench Pro cell is an imported official score (starred), not a same-harness rerun; QwenSWEBench is in-house.
| Benchmark | 3.8-27B | 3.6-27B | 3.7-Plus | Opus 4.6 Max |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 73.0 | 63.4 | 64.0 | Best in this comparison 78.2 |
| SWE-bench Pro* | Best in this comparison 61.7 | 53.5 | 57.6 | 53.4 |
| LiveCodeBench v6 | Best in this comparison 90.3 | 83.9 | 89.6 | 88.8 |
| OSWorld-Verified | Best in this comparison 84.3 | 63.9 | 73.3 | 72.7 |
| QwenSWEBench | Best in this comparison 79.0 | 49.3 | 59.2 | 63.8 |
| DeepSWE 1.1 | Best in this comparison 42.2 | 13.3 | 14.2 | — |
| NL2Repo-Bench | 42.3 | 36.2 | 41.1 | Best in this comparison 47.6 |
| GPQA Diamond | 89.2 | 87.8 | 90.3 | Best in this comparison 91.3 |
How to read it against the first rank: versus Opus 4.6, this 27B leads or sits slightly ahead on SWE-Pro, LiveCodeBench, OSWorld, and QwenSWEBench; it still loses Terminal-Bench, NL2Repo, and GPQA. Versus 3.6-27B the agent jump is a break (DeepSWE 13.3→42.2, QwenSWEBench 49.3→79.0). Versus the Plus API, the dense 27B holds most coding rows — and Plus is billed per token while the 27B can leave the building.
Coder 480B and GPT-5.6 Sol are not on this harness. Do not turn 61.7 into “it beat every frontier model.” The decision that survives: screenshots, private repos, vendor-table caveats accepted → 3.8-27B on a 24GB card is the first time this size looks near-frontier. SLA, hosted 1M, no GPU to nurse → keep paying Opus / Sol / Max.
Beginner hardware: tonight, not someday
Enlarge image Calling it “can run” is too soft. The accurate sentence: a 24GB consumer GPU (4090-class) can load a third-party 4-bit build tonight and act as a coding agent at 8K–32K. Opus 4.6 cannot be “loaded” — you are not given the weights. Coder 480B and V4 Flash are not this chassis.
| Rig | Fits | Context to start | What you get | What you give up |
|---|---|---|---|---|
| 24GB consumer GPU | Q4_K_M ~17.1GB* | 8K–32K | Coding agent tonight | Opus never leaves the API |
| 32–48GB workstation | FP8 30.9GB | 32K–64K | Less quant damage | Still no download for Opus |
| 80GB-class card | BF16 55.6GB | native 262K is tight | Need KV budget (~16GiB @262K) | API still simpler |
| No local GPU | n/a | 262K–1M hosted | Pay OpenRouter | 3.8-27B $0.45/$3.20; Max $2/$6; Opus $5/$25 |
Do not weld two claims together: “fits in 17GB” and “supports 262K.” The full-attention KV lower bound at 262K BF16 is about 16GiB (the 16 full-attention layers only). On 24GB, a full window dies of VRAM before it dies of skill. Beginner protocol: pin 8K–32K, leave 1M YaRN off, rerun the same failing repository-repair tasks to see what quantization cost you. Comfort is FP8 on 48GB. Native BF16 at full 262K is an 80GB-class plan.
That is the low-floor shock: a first-line API is $5/$25 and never downloads. This 27B sits on a gaming GPU. Quantization still hurts reasoning, vision, and tool format — the shock is that you can start, not that it already is Opus.
On-site board: a proxy, not this coding table
This site’s unified composite rewards long context and low unit price, not SWE. In this snapshot 3.8-27B (now listed at $0.45/$3.20), 3.6-27B, 3.8-Max, and Opus 4.6 do not rise because they “code well” — 262K at a mid sticker does not beat a 1M Flash slot.
| On-site unified slot | Model | Why it ranks |
|---|---|---|
| #21 | Qwen3.7-Flash | 1M ctx + $0.03/$0.13 |
| #93 | Qwen3.5-Flash | 1M + cheap |
| #97 | Qwen3 Coder Flash | 1M code API, still cheap |
| #100 | Qwen3.7-Plus | 1M, $0.32/$1.28 |
| not on board | 3.8-27B / Opus 4.6 | listing ≠ 1M-cheap Flash slot / frontier sticker |
HF Hot still places official Qwen3.6-27B at sort_index=45 with about 6,895,121 downloads. 3.8-27B is not in that hot snapshot yet. The composite board will not climb because of SWE 61.7. None of this is a capability benchmark.
When to run the 27B, when to keep paying Opus
Choose 3.8-27B when you have a 24GB card, the repo must not leave the building, you need screenshots or a terminal, and you can live with a vendor table instead of an independent board. Treat it as a coding upgrade over 3.6-27B and measure the Q4 on the same failing tests. Without a card, take OpenRouter `qwen/qwen3.8-27b` ($0.45/$3.20, ~$0.11 for the same job) — an order of magnitude cheaper than Opus, still pricier than small Coder APIs.
Keep paying Opus 4.6 / GPT-5.6 Sol when you need SLA, hosted 1M, no GPU to nurse, and you will not swallow imported-score / in-house-suite risk. Pay Max when you want the same-generation Qwen flagship API cheaper than Opus ($2/$6 vs $5/$25). Pay Coder 480B when you want text-only code and a ready 262K API without downloading 27B.
Do not call hosted 3.8-27B “the cheapest 27B API” — same-size predecessor 3.6-27B and Coder 30B are cheaper. Do not call the 24GB on-ramp a full 262K. Decide with the Token page, compare, and the on-site model page.