DeepSeek V4.1 Flash: CED 552B, vendor agent scores, and the V4.1 Pro that is not listed
Catalog first seen 2026-09-10. OpenRouter primary $0.15/$0.60; Fireworks $0.22/$0.66, Together $0.30/$1.20. HF card: CED 552B, 8B prefill / 16B decode, 890 B/token global KV, MIT. Vendor table: Terminal-Bench 2.1 90.6, DeepSWE 74.2 (max effort). No V4.1 Pro listing in this snapshot. Composite #100 (context-per-dollar proxy, not capability). 0731 is still about 3× cheaper on the same job.
Enlarge image What it is, and why there is no V4.1 Pro
DeepSeek V4.1 Flash (OpenRouter deepseek/deepseek-v4.1-flash) entered this site’s catalog on 2026-09-10. The official HF card describes the first Causal Encoder-Decoder (CED): a 40-layer Transformer (20-layer causal encoder + 20-layer decoder), 552B backbone parameters, 8B activated per token in prefill and 16B in decode. Native image+text in, text out. Context 1,048,576. MIT weights. HF also prints “Model size 763B params” — that figure includes sparse Engram memory and is not the same sentence as the 552B backbone.
This snapshot’s core_model table has no deepseek/deepseek-v4.1-pro and no V4.1 Pro alias. The comparable flagship SKUs are still last-generation V4 Pro 0813 (GA) and V4 Pro 0423. Using Flash’s vendor table as a stand-in for an unlisted Pro is a marketing jump. This page does not take it.
Versus 0731: the skeleton changed (CED vs 284B/13B sparse MoE), vision is native, persistent KV is claimed at about 1/4, and the sticker moved from $0.06/$0.12 to $0.15/$0.60. 0731’s catalog context is now 1,310,720; V4.1 stays at 1,048,576. Do not treat the word “Flash” as one checkpoint.
Spec sheet the catalog can check
All four columns are on-site catalog plus OpenRouter primary listings. HF Hot and the composite board are this snapshot, not a capability score.
| V4.1 Flash | V4 Flash 0731 | V4 Pro 0813 | V4 Pro 0423 | |
|---|---|---|---|---|
| Catalog first_seen | 2026-09-10 | 2026-08-01 | 2026-08-18 | 2026-06-22 |
| OpenRouter id | deepseek/deepseek-v4.1-flash | …-v4-flash-0731 | …-v4-pro-0813 | …-v4-pro |
| Architecture | CED MoE 552B | sparse MoE 284B | MoE (GA Pro) | MoE 1.6T |
| Activated | 8B prefill / 16B decode | ~13B | Pro-scale (card: 49B on 0423) | Best in this comparison 49B |
| Context | 1,048,576 | Best in this comparison 1,310,720 | 1,048,576 | 1,048,576 |
| Modality | text+image→text | text→text | text→text | text→text |
| Weights | MIT on HF | family weights | API listing | API listing |
| OR primary $/1M | $0.15 / $0.60 | Best in this comparison $0.06 / $0.12 | $0.98 / $2.95 | $1.60 / $3.20 |
| Fireworks / Together | $0.22/$0.66 · $0.30/$1.20 | — | — | — |
| HF Hot (this snap) | #1 · +66% · ~326K DL | not this row | not this row | not this row |
| Unified composite | #100 (ctx/$ proxy) | #3 | not top-cheap | not top-cheap |
Read it this way: V4.1 is a new skeleton, native vision, and a dearer Flash sticker. 0731 remains the cheapest long-context DeepSeek listing. The two Pro columns are last-generation flagships, not “the V4.1 Pro.” Fireworks and Together are the same weights at other cash registers.
Vendor agent table: leads TB2.1, still trails Opus on TB4
The sheet below is copied from the HF card’s “Comparison with frontier models (Max reasoning effort).” Attribution on the card: reasoning_effort=100, temperature=1.0, top_p=0.95; code-agent rows use DeepSeek Harness Minimal at a 1M window; DeepSWE uses mini-SWE; some visual-agent rows use a Claude Code harness. † is the text-only HLE subset. This site did not rerun it.
| Benchmark | V4.1 Flash | V4 Flash | V4 Pro | Opus 5.0 |
|---|---|---|---|---|
| GPQA Diamond | 90.9 | 89.9 | 92.4 | Best in this comparison 93.4 |
| Terminal-Bench 2.1 | Best in this comparison 90.6 | 82.7 | 87.9 | 89.1 |
| Terminal-Bench 3.0 | 30.0 | 7.6 | 11.8 | Best in this comparison 43.3 |
| Terminal-Bench 4.0 | 31.2 | 7.0 | 12.4 | Best in this comparison 51.8 |
| DeepSWE v1.1 | Best in this comparison 74.2 | 54.4 | 62.7 | 74.0 |
| NL2Repo-Bench | 64.0 | 54.2 | 61.5 | Best in this comparison 75.3 |
| AutomationBench | Best in this comparison 54.8 | 37.7 | 43.2 | 50.3 |
| HLE (text† / tools) | 36.8 (39.1†) / 63.9 | 37.8† / 51.5 | 42.7† / 60.0 | 56.3 / 63.6 |
What survives: versus family V4 Flash the agent jump is a break (DeepSWE 54.4→74.2, TB3 7.6→30.0). Versus last-gen V4 Pro most agent rows also rise. Versus Opus 5.0, TB2.1 and DeepSWE sit slightly ahead; AutomationBench and HLE-with-tools are close; TB3 / TB4 / NL2Repo still lose clearly. Do not turn 90.6 into “it beat every frontier model.” The card’s own scaffold table drops DeepSWE from 74.2 (mini-SWE) to 69.8 (Claude Code).
Same job: 0731 ~$0.008, V4.1 ~$0.027, Pro 0813 ~$0.16
Enlarge image Hold the job at 100K in + 20K out. This is not a capability score. It is who you pay for the same tokens.
| Listing | In $/1M | Out $/1M | 100K+20K | Note |
|---|---|---|---|---|
| V4 Flash 0731 OR | Best in this comparison $0.06 | Best in this comparison $0.12 | Best in this comparison $0.008 | cheapest Flash SKU |
| V4.1 Flash OR | $0.15 | $0.60 | $0.027 | primary listing |
| V4.1 Flash Fireworks | $0.22 | $0.66 | $0.035 | same model, other host |
| V4.1 Flash Together | $0.30 | $1.20 | $0.054 | same model, other host |
| V4 Pro 0813 OR | $0.98 | $2.95 | $0.157 | last-gen Pro GA |
| V4 Pro 0423 OR | $1.60 | $3.20 | $0.224 | last-gen Pro 0423 |
V4.1’s OpenRouter primary is no longer the cheapest Flash. 0731 bills about $0.008 for the same job — roughly one third. Together charges about $0.054 for the same weights, about 2× the primary listing. Pro 0813 is about $0.16 and 0423 about $0.22: last-generation flagships, not a missing V4.1 Pro.
KV: 890 B/token on the card, about 1/4 of V4 Flash
Enlarge image The card puts the pitch in the title: KV cache compression. Decoder global KV is projected from the last encoder hidden states rather than stored per decoder layer. CSA2 shares main KV and indexer state across layers. Main KV is cached in FP4 (E2M1). Result: 890 bytes/token. At 1M context that is about 890MB of persistent KV — still not “run 552B on a laptop,” only less SSD traffic on the serving side.
The card also claims SWA Bounded Replay cuts persistent SWA KV to about 1/8 of V4-Flash, and a 437× reduction versus DeepSeek-V1. The last comparison spans the whole product history. This page nails 890 B/token and the ~4× versus V4 Flash. There is no public third-party KV dump to audit it.
On-site #100: a proxy, not this agent table
The unified composite rewards long context and low unit price. In this snapshot V4.1 Flash sits at #100 (sort_index=99) because its blended sticker is about $0.375/1M — it loses to 0731 at $0.09 with a 1.31M window. 0731 is #3; ~deepseek/deepseek-v4-flash-latest is #2.
HF Hot places official DeepSeek-V4.1-Flash at sort_index=0 with +66% 24h growth and about 325,712 downloads, flagged new_listing. Heat is not the composite. The composite is not TB 90.6.
None of this is a capability benchmark.
When to take V4.1 Flash, when to stay on 0731 / last-gen Pro
Take V4.1 Flash when you need native vision, the vendor agent band (especially TB2.1 / DeepSWE), can pay $0.15/$0.60, and will live with a marketing harness. OpenRouter is the primary listing; Fireworks and Together are the other registers. MIT weights download — 552B MoE is not a 24GB consumer-card story.
Stay on 0731 when you want text-only and the cheapest 1M+ window at $0.06/$0.12. Do not pay ~3× for a “.1.”
Last-gen Pro 0813 / 0423: there is still no V4.1 Pro in the catalog. If your comparison class is “larger-activation DeepSeek flagship,” those two SKUs are what you can click today, at roughly an order of magnitude more. The next article writes itself when a Pro listing exists. This one will not invent it.
Decide with the Token page, compare, and the on-site model page.