Qwen model series deep dive: a selection guide from Qwen2.5 to Qwen3.7
49 Qwen SKUs span Max, Plus, Flash, Coder, and VL lines—the flagship Qwen3.7 Max is $1.25/$3.75, Flash is $0.065/$0.26, and the free Coder ranks #6 on the unified board.
What the Qwen series is
Qwen (Tongyi Qianwen) is Alibaba’s open and hosted LLM brand; in our catalog the vendor_id is qwen with 49 OpenRouter-mirrored SKUs today. It is not one model but a scenario-tiered product matrix: Max for hard reasoning, Plus/Flash for everyday long context, Coder for programming and agent workflows, VL for image-text multimodal inputs. Unlike vendors with one or two API endpoints, Qwen spans a full price gradient from $0 free to $1.25/$3.75 flagship under the same vendor—pick the task type first, then compare SKUs within it.
Product line map
By generation and scenario, the 49 SKUs group into six main lines.
Qwen2.5 (prior generation): Qwen2.5 72B Instruct (131K ctx), Qwen2.5 Coder 32B, Qwen2.5 VL 72B—for mature, lower-cost legacy workflows.
Qwen3 base: 8B/14B/32B/235B MoE tiers; Qwen3 32B at $0.08/$0.28 is the value anchor for lightweight APIs.
Qwen3.5/3.6/3.7 iterations: Qwen3.5-Flash (1M ctx, multimodal), Qwen3.6 Plus/Flash, Qwen3.7 Plus/Max form the current core; Qwen3.7 Max is the 1M flagship, Qwen3 Max the prior 262K flagship.
Max/Plus/Flash tiers: Max is priciest (Qwen3.7 Max $1.25/$3.75), Plus mid-range ($0.32/$1.28), Flash cheapest ($0.065/$0.26).
Coder line: Qwen3 Coder 480B MoE (1M ctx, $0.22/$1.80) and free qwen3-coder:free ($0, board #6); plus Coder Flash/Plus/Next variants.
VL multimodal: Qwen3 VL 235B/32B/8B; Qwen3 VL 235B Instruct is $0.20/$0.88, 262K ctx, accepts image input.
Spec sheet and pricing at a glance
All figures below are OpenRouter primary listing unit prices (USD/1M tokens, input/output), verifiable on the site token page.
Qwen3.7 Max: 1M context, $1.25/$3.75—the priciest flagship today.
Qwen3 Max: 262K context, $0.78/$3.90—prior flagship, shorter window but still expensive output.
Qwen3.7 Plus: 1M context, $0.32/$1.28—the balanced tier for long documents and agents.
Qwen3.5-Flash: 1M context, $0.065/$0.26—the series’ lowest unit price, for high-volume light tasks.
Qwen3 Coder 480B: 1M context, $0.22/$1.80—MoE code flagship; free qwen3-coder:free is also 1M ctx at $0/$0.
Qwen3 32B: 131K context, $0.08/$0.28—lightweight API anchor; also listed on DeepInfra, Groq, and others.
Qwen3 VL 235B Instruct: 262K context, $0.20/$0.88—multimodal image-text input.
Blended price reference (100K input + 20K output): Flash ≈ $0.01, Plus/Coder ≈ $0.06, Max ≈ $0.20—a gap of up to 20×.
On-site board ranks and open-source traction
On the AI Hippo unified composite board, 12 Qwen SKUs are ranked. The highest is qwen3-coder:free (#6), followed by paid qwen3-coder (#74), Qwen3.5-Flash (#88), and the Plus/Flash cluster between #90–#98.
Key caveat: the composite score is a context-length-over-price proxy that rewards long context and low unit price—it is not a capability benchmark like MMLU or HumanEval. A high rank means strong spec-to-price ratio, not necessarily the strongest model.
On the open-source side, Hugging Face hot-list data shows Qwen3.6-27B at ~4.96M downloads, Qwen3.6-27B-MTP-GGUF at ~2.9M, and Qwen3.6-35B-A3B variants at ~2.51M—signaling sustained demand for local deployment. Hosted API and local GGUF are separate paths: the former bills per token with no ops; the latter has zero token cost but requires your own hardware.
The above are spec and cost signals, not capability benchmarks.
Cost in practice and selection matrix
Two typical scenarios with dollar costs (OpenRouter primary listing prices).
Scenario A: 100K input + 20K output (medium chat/document summary) Qwen3.7 Max ≈ $0.20 | Qwen3.7 Plus ≈ $0.06 | Qwen3.5-Flash ≈ $0.01 | Qwen3 Coder ≈ $0.06 | qwen3-coder:free = $0
Scenario B: 500K input + 5K output (long-context retrieval/whole-repo scan) Qwen3.7 Max ≈ $0.64 | Qwen3.7 Plus ≈ $0.17 | Qwen3.5-Flash ≈ $0.03 | Qwen3 Coder ≈ $0.12
Selection matrix: Flagship reasoning and complex agents → Qwen3.7 Max (1M ctx, priciest but longest window) Everyday long documents and general agents → Qwen3.7 Plus or Qwen3.6 Plus (1M ctx, $0.32/$1.28 tier) High-volume light tasks → Qwen3.5-Flash ($0.065/$0.26, lowest unit price in the series) Programming and code agents → Qwen3 Coder 480B or free tier (1M ctx, MoE architecture) Image-text multimodal → Qwen3 VL 235B Instruct ($0.20/$0.88, 262K ctx) Local deployment/zero token cost → HF open weights (Qwen3.6-27B, ~4.96M downloads)
The above is a spec and cost comparison, not a capability benchmark.
When to pick which SKU
When picking Qwen, start with three questions: how much context do you need? What share of tokens is output? Is multimodal or local deployment mandatory?
If the answer is "1M context + highest-quality reasoning + cost-insensitive," choose Qwen3.7 Max. If "1M context + everyday agent + cost-sensitive," Qwen3.7 Plus or Flash fits better—Flash costs 1/20 of Max in a 100K+20K scenario.
For programming: prototype on free qwen3-coder:free (unified board #6, $0), then scale to paid Coder 480B or Coder Plus in production. Image-text multimodal tasks go to VL 235B; legacy Qwen2.5 workflows can stay on 72B Instruct ($0.36/$0.40, 131K ctx) to avoid migration cost.
Skip Qwen when: you need the lowest latency and Qwen3 32B's 131K window is not enough—compare other vendors on Groq/DeepInfra; for enterprise compliance with official SLA, connect directly to Alibaba Cloud Model Studio, not only the OpenRouter mirror.
The site already has a qwen-qwen3-coder-free-review single-model deep dive and a vendor page listing all 49 SKUs—start from the vendor page, drill into a specific model, and validate side-by-side on compare.
Sources
Evidence and actions
- Time window: catalog snapshot
- Observation count: 12
- Source type: catalog_and_pricing
- Browse all 49 Qwen SKUs on /en/model/vendor/qwen/
- Compare provider quotes on /en/token/
- Read the free Coder deep dive on /en/insights/qwen-qwen3-coder-free-review/