Qwen3.8-27B: a 24GB consumer GPU against Opus 4.6’s coding table

Apache-2.0 dense 27.78B. A third-party Q4 (~17GB) fits a 24GB consumer GPU. OpenRouter primary listing $0.45/$3.20. Qwen’s launch table (vendor-reported) has SWE-bench Pro 61.7 vs Opus 4.6 Max’s imported 53.4; Terminal-Bench 73.0 vs 78.2. Same 100K+20K job: local ~$0, hosted ~$0.11, Opus $1.00. Not an independent benchmark.

Qwen launch table: 3.8-27B vs 3.6-27B, 3.7-Plus, and Opus 4.6 Max Enlarge image
Vendor-reported. The Opus SWE-Pro cell is an imported official score, not a same-harness rerun.

What it is, and why this table should startle you

Qwen3.8-27B is the Apache-2.0 dense multimodal checkpoint published 2026-08-14: 27,781,427,952 parameters on Hugging Face, text/image/video in, text out, native 262K. It is not Qwen3-8B, and it is not API-only Qwen3.8-Max.

The shock is the floor, not “another 27B.” A third-party Q4_K_M file is about 17.1GB — it fits a 24GB consumer GPU. You can boot a coding agent on a desk tonight. The peer on the other side of that sentence is Claude Opus 4.6: OpenRouter primary listing $5/$25, weights not downloadable.

The four bars are Qwen’s own launch table. 3.8-27B leads the in-table Opus 4.6 Max on SWE-Pro, LiveCodeBench, and OSWorld; it trails by 5.2 on Terminal-Bench. That is a marketing harness, not an independent reproduction. It is still the comparison this page exists to nail down.

Spec sheet, side by side

Put the numbers you can check on the table first. All four columns are on-site catalog plus OpenRouter primary listings (`qwen/qwen3.8-27b` is now $0.45/$3.20).

3.8-27B 3.6-27B 3.8-Max Opus 4.6
Released 2026-08-14 (HF) catalog first_seen 2026-06-22 catalog first_seen 2026-08-04 in catalog
Params 27.78B dense ~27.8B dense flagship 3.8 API frontier API
Weights Apache-2.0 download Apache-2.0 + API API only API only
Context 262K native / 1M YaRN 262K Best in this comparison 1M Best in this comparison 1M
Modality text+image+video→text text+image+video→text text+image+video→text multimodal API
OpenRouter $/1M $0.45 / $3.20 Best in this comparison $0.289 / $2.40 $2.00 / $6.00 $5.00 / $25.00

The published decoder for 3.8-27B and 3.6-27B is essentially the same shape (64-layer hybrid attention, hidden 5120). The coding jump is a training story, not a new skeleton. Max and Opus are a different product line: hosted 1M, billed per token, you never own the weights. 3.8-27B has both doors: downloadable local weights, and a hosted API on the same checkpoint.

Same job: local ~$0, hosted ~$0.11, Opus $1

USD cost of 100K input + 20K output Enlarge image
OpenRouter primary listings. Local 3.8-27B is electricity only; hosted same weights ~$0.11.

Hold the job fixed at 100K in + 20K out and recompute from on-site OpenRouter primary listings. This is not a capability score. It is who you pay for the same token work.

Model In $/1M Out $/1M 100K+20K Role
3.8-27B local Best in this comparison ~$0 (electricity) 24GB Q4 desk
Qwen3 Coder 30B A3B Best in this comparison $0.07 Best in this comparison $0.27 $0.012 cheap code API
Qwen3 Coder 480B $0.30 $1.00 $0.050 hosted code flagship
Qwen3.6-27B API $0.289 $2.40 $0.077 same-size predecessor
Qwen3.8-27B API $0.45 $3.20 $0.109 same weights, hosted
Qwen3.8-Max $2.00 $6.00 $0.32 same-gen cloud flagship
Claude Opus 4.6 $5.00 $25.00 $1.00 frontier coding API
GPT-5.6 Sol $5.00 $30.00 $1.10 OpenAI coding flagship

Read it this way: Opus 4.6 and GPT-5.6 Sol bill about $1 for that job. Same-generation Max is about $0.32. Hosted 3.8-27B is about $0.11 (a bit above predecessor 3.6-27B at $0.077). Hosted Coder flagship is about $0.05. The same weights on your own 24GB card have a token sticker of zero — you pay electricity and depreciation. DeepSeek V4 Flash 0731 is the other extreme: 284B MoE at $0.09/$0.18, cheap as an API, not a consumer-card download.

Coding: versus Opus 4.6, not just the family

Qwen’s launch table puts 3.8-27B, predecessor 3.6-27B, hosted 3.7-Plus, and Claude Opus 4.6 Max on one sheet. The copy below is that sheet. Attribution: vendor-reported; several rows used a Claude Code harness; the Opus SWE-bench Pro cell is an imported official score (starred), not a same-harness rerun; QwenSWEBench is in-house.

Benchmark 3.8-27B 3.6-27B 3.7-Plus Opus 4.6 Max
Terminal-Bench 2.1 73.0 63.4 64.0 Best in this comparison 78.2
SWE-bench Pro* Best in this comparison 61.7 53.5 57.6 53.4
LiveCodeBench v6 Best in this comparison 90.3 83.9 89.6 88.8
OSWorld-Verified Best in this comparison 84.3 63.9 73.3 72.7
QwenSWEBench Best in this comparison 79.0 49.3 59.2 63.8
DeepSWE 1.1 Best in this comparison 42.2 13.3 14.2
NL2Repo-Bench 42.3 36.2 41.1 Best in this comparison 47.6
GPQA Diamond 89.2 87.8 90.3 Best in this comparison 91.3

How to read it against the first rank: versus Opus 4.6, this 27B leads or sits slightly ahead on SWE-Pro, LiveCodeBench, OSWorld, and QwenSWEBench; it still loses Terminal-Bench, NL2Repo, and GPQA. Versus 3.6-27B the agent jump is a break (DeepSWE 13.3→42.2, QwenSWEBench 49.3→79.0). Versus the Plus API, the dense 27B holds most coding rows — and Plus is billed per token while the 27B can leave the building.

Coder 480B and GPT-5.6 Sol are not on this harness. Do not turn 61.7 into “it beat every frontier model.” The decision that survives: screenshots, private repos, vendor-table caveats accepted → 3.8-27B on a 24GB card is the first time this size looks near-frontier. SLA, hosted 1M, no GPU to nurse → keep paying Opus / Sol / Max.

Beginner hardware: tonight, not someday

Weight footprint versus a 24GB consumer GPU Enlarge image
Weights only. KV is extra. Q4 is third-party Unsloth. V4 Flash ~149GB is a public serving footprint.

Calling it “can run” is too soft. The accurate sentence: a 24GB consumer GPU (4090-class) can load a third-party 4-bit build tonight and act as a coding agent at 8K–32K. Opus 4.6 cannot be “loaded” — you are not given the weights. Coder 480B and V4 Flash are not this chassis.

Rig Fits Context to start What you get What you give up
24GB consumer GPU Q4_K_M ~17.1GB* 8K–32K Coding agent tonight Opus never leaves the API
32–48GB workstation FP8 30.9GB 32K–64K Less quant damage Still no download for Opus
80GB-class card BF16 55.6GB native 262K is tight Need KV budget (~16GiB @262K) API still simpler
No local GPU n/a 262K–1M hosted Pay OpenRouter 3.8-27B $0.45/$3.20; Max $2/$6; Opus $5/$25

Do not weld two claims together: “fits in 17GB” and “supports 262K.” The full-attention KV lower bound at 262K BF16 is about 16GiB (the 16 full-attention layers only). On 24GB, a full window dies of VRAM before it dies of skill. Beginner protocol: pin 8K–32K, leave 1M YaRN off, rerun the same failing repository-repair tasks to see what quantization cost you. Comfort is FP8 on 48GB. Native BF16 at full 262K is an 80GB-class plan.

That is the low-floor shock: a first-line API is $5/$25 and never downloads. This 27B sits on a gaming GPU. Quantization still hurts reasoning, vision, and tool format — the shock is that you can start, not that it already is Opus.

On-site board: a proxy, not this coding table

This site’s unified composite rewards long context and low unit price, not SWE. In this snapshot 3.8-27B (now listed at $0.45/$3.20), 3.6-27B, 3.8-Max, and Opus 4.6 do not rise because they “code well” — 262K at a mid sticker does not beat a 1M Flash slot.

On-site unified slot Model Why it ranks
#21 Qwen3.7-Flash 1M ctx + $0.03/$0.13
#93 Qwen3.5-Flash 1M + cheap
#97 Qwen3 Coder Flash 1M code API, still cheap
#100 Qwen3.7-Plus 1M, $0.32/$1.28
not on board 3.8-27B / Opus 4.6 listing ≠ 1M-cheap Flash slot / frontier sticker

HF Hot still places official Qwen3.6-27B at sort_index=45 with about 6,895,121 downloads. 3.8-27B is not in that hot snapshot yet. The composite board will not climb because of SWE 61.7. None of this is a capability benchmark.

When to run the 27B, when to keep paying Opus

Choose 3.8-27B when you have a 24GB card, the repo must not leave the building, you need screenshots or a terminal, and you can live with a vendor table instead of an independent board. Treat it as a coding upgrade over 3.6-27B and measure the Q4 on the same failing tests. Without a card, take OpenRouter `qwen/qwen3.8-27b` ($0.45/$3.20, ~$0.11 for the same job) — an order of magnitude cheaper than Opus, still pricier than small Coder APIs.

Keep paying Opus 4.6 / GPT-5.6 Sol when you need SLA, hosted 1M, no GPU to nurse, and you will not swallow imported-score / in-house-suite risk. Pay Max when you want the same-generation Qwen flagship API cheaper than Opus ($2/$6 vs $5/$25). Pay Coder 480B when you want text-only code and a ready 262K API without downloading 27B.

Do not call hosted 3.8-27B “the cheapest 27B API” — same-size predecessor 3.6-27B and Coder 30B are cheaper. Do not call the 24GB on-ramp a full 262K. Decide with the Token page, compare, and the on-site model page.

Insights