DeepSeek V4.1 Flash: CED 552B, vendor agent scores, and the V4.1 Pro that is not listed

Catalog first seen 2026-09-10. OpenRouter primary $0.15/$0.60; Fireworks $0.22/$0.66, Together $0.30/$1.20. HF card: CED 552B, 8B prefill / 16B decode, 890 B/token global KV, MIT. Vendor table: Terminal-Bench 2.1 90.6, DeepSWE 74.2 (max effort). No V4.1 Pro listing in this snapshot. Composite #100 (context-per-dollar proxy, not capability). 0731 is still about 3× cheaper on the same job.

Vendor agent table: V4.1 Flash vs V4 Flash, V4 Pro, and Opus 5.0 Enlarge image
DeepSeek HF card, max reasoning_effort=100. DeepSWE uses mini-SWE; Terminal-Bench uses DeepSeek Harness Minimal. Not an independent same-harness rerun.

What it is, and why there is no V4.1 Pro

DeepSeek V4.1 Flash (OpenRouter deepseek/deepseek-v4.1-flash) entered this site’s catalog on 2026-09-10. The official HF card describes the first Causal Encoder-Decoder (CED): a 40-layer Transformer (20-layer causal encoder + 20-layer decoder), 552B backbone parameters, 8B activated per token in prefill and 16B in decode. Native image+text in, text out. Context 1,048,576. MIT weights. HF also prints “Model size 763B params” — that figure includes sparse Engram memory and is not the same sentence as the 552B backbone.

This snapshot’s core_model table has no deepseek/deepseek-v4.1-pro and no V4.1 Pro alias. The comparable flagship SKUs are still last-generation V4 Pro 0813 (GA) and V4 Pro 0423. Using Flash’s vendor table as a stand-in for an unlisted Pro is a marketing jump. This page does not take it.

Versus 0731: the skeleton changed (CED vs 284B/13B sparse MoE), vision is native, persistent KV is claimed at about 1/4, and the sticker moved from $0.06/$0.12 to $0.15/$0.60. 0731’s catalog context is now 1,310,720; V4.1 stays at 1,048,576. Do not treat the word “Flash” as one checkpoint.

Spec sheet the catalog can check

All four columns are on-site catalog plus OpenRouter primary listings. HF Hot and the composite board are this snapshot, not a capability score.

V4.1 Flash V4 Flash 0731 V4 Pro 0813 V4 Pro 0423
Catalog first_seen 2026-09-10 2026-08-01 2026-08-18 2026-06-22
OpenRouter id deepseek/deepseek-v4.1-flash …-v4-flash-0731 …-v4-pro-0813 …-v4-pro
Architecture CED MoE 552B sparse MoE 284B MoE (GA Pro) MoE 1.6T
Activated 8B prefill / 16B decode ~13B Pro-scale (card: 49B on 0423) Best in this comparison 49B
Context 1,048,576 Best in this comparison 1,310,720 1,048,576 1,048,576
Modality text+image→text text→text text→text text→text
Weights MIT on HF family weights API listing API listing
OR primary $/1M $0.15 / $0.60 Best in this comparison $0.06 / $0.12 $0.98 / $2.95 $1.60 / $3.20
Fireworks / Together $0.22/$0.66 · $0.30/$1.20
HF Hot (this snap) #1 · +66% · ~326K DL not this row not this row not this row
Unified composite #100 (ctx/$ proxy) #3 not top-cheap not top-cheap

Read it this way: V4.1 is a new skeleton, native vision, and a dearer Flash sticker. 0731 remains the cheapest long-context DeepSeek listing. The two Pro columns are last-generation flagships, not “the V4.1 Pro.” Fireworks and Together are the same weights at other cash registers.

Vendor agent table: leads TB2.1, still trails Opus on TB4

The sheet below is copied from the HF card’s “Comparison with frontier models (Max reasoning effort).” Attribution on the card: reasoning_effort=100, temperature=1.0, top_p=0.95; code-agent rows use DeepSeek Harness Minimal at a 1M window; DeepSWE uses mini-SWE; some visual-agent rows use a Claude Code harness. † is the text-only HLE subset. This site did not rerun it.

Benchmark V4.1 Flash V4 Flash V4 Pro Opus 5.0
GPQA Diamond 90.9 89.9 92.4 Best in this comparison 93.4
Terminal-Bench 2.1 Best in this comparison 90.6 82.7 87.9 89.1
Terminal-Bench 3.0 30.0 7.6 11.8 Best in this comparison 43.3
Terminal-Bench 4.0 31.2 7.0 12.4 Best in this comparison 51.8
DeepSWE v1.1 Best in this comparison 74.2 54.4 62.7 74.0
NL2Repo-Bench 64.0 54.2 61.5 Best in this comparison 75.3
AutomationBench Best in this comparison 54.8 37.7 43.2 50.3
HLE (text† / tools) 36.8 (39.1†) / 63.9 37.8† / 51.5 42.7† / 60.0 56.3 / 63.6

What survives: versus family V4 Flash the agent jump is a break (DeepSWE 54.4→74.2, TB3 7.6→30.0). Versus last-gen V4 Pro most agent rows also rise. Versus Opus 5.0, TB2.1 and DeepSWE sit slightly ahead; AutomationBench and HLE-with-tools are close; TB3 / TB4 / NL2Repo still lose clearly. Do not turn 90.6 into “it beat every frontier model.” The card’s own scaffold table drops DeepSWE from 74.2 (mini-SWE) to 69.8 (Claude Code).

Same job: 0731 ~$0.008, V4.1 ~$0.027, Pro 0813 ~$0.16

USD cost of 100K input + 20K output Enlarge image
On-site OpenRouter primary listings. 0731 is still ~3× cheaper. Together is ~$0.054 for the same job.

Hold the job at 100K in + 20K out. This is not a capability score. It is who you pay for the same tokens.

Listing In $/1M Out $/1M 100K+20K Note
V4 Flash 0731 OR Best in this comparison $0.06 Best in this comparison $0.12 Best in this comparison $0.008 cheapest Flash SKU
V4.1 Flash OR $0.15 $0.60 $0.027 primary listing
V4.1 Flash Fireworks $0.22 $0.66 $0.035 same model, other host
V4.1 Flash Together $0.30 $1.20 $0.054 same model, other host
V4 Pro 0813 OR $0.98 $2.95 $0.157 last-gen Pro GA
V4 Pro 0423 OR $1.60 $3.20 $0.224 last-gen Pro 0423

V4.1’s OpenRouter primary is no longer the cheapest Flash. 0731 bills about $0.008 for the same job — roughly one third. Together charges about $0.054 for the same weights, about 2× the primary listing. Pro 0813 is about $0.16 and 0423 about $0.22: last-generation flagships, not a missing V4.1 Pro.

KV: 890 B/token on the card, about 1/4 of V4 Flash

Global KV bytes per token (vendor claim) Enlarge image
HF card: CSA2 + FP4 main KV = 890 B/token, about 4× smaller than V4 Flash. Not an independent serving dump.

The card puts the pitch in the title: KV cache compression. Decoder global KV is projected from the last encoder hidden states rather than stored per decoder layer. CSA2 shares main KV and indexer state across layers. Main KV is cached in FP4 (E2M1). Result: 890 bytes/token. At 1M context that is about 890MB of persistent KV — still not “run 552B on a laptop,” only less SSD traffic on the serving side.

The card also claims SWA Bounded Replay cuts persistent SWA KV to about 1/8 of V4-Flash, and a 437× reduction versus DeepSeek-V1. The last comparison spans the whole product history. This page nails 890 B/token and the ~4× versus V4 Flash. There is no public third-party KV dump to audit it.

On-site #100: a proxy, not this agent table

The unified composite rewards long context and low unit price. In this snapshot V4.1 Flash sits at #100 (sort_index=99) because its blended sticker is about $0.375/1M — it loses to 0731 at $0.09 with a 1.31M window. 0731 is #3; ~deepseek/deepseek-v4-flash-latest is #2.

HF Hot places official DeepSeek-V4.1-Flash at sort_index=0 with +66% 24h growth and about 325,712 downloads, flagged new_listing. Heat is not the composite. The composite is not TB 90.6.

None of this is a capability benchmark.

When to take V4.1 Flash, when to stay on 0731 / last-gen Pro

Take V4.1 Flash when you need native vision, the vendor agent band (especially TB2.1 / DeepSWE), can pay $0.15/$0.60, and will live with a marketing harness. OpenRouter is the primary listing; Fireworks and Together are the other registers. MIT weights download — 552B MoE is not a 24GB consumer-card story.

Stay on 0731 when you want text-only and the cheapest 1M+ window at $0.06/$0.12. Do not pay ~3× for a “.1.”

Last-gen Pro 0813 / 0423: there is still no V4.1 Pro in the catalog. If your comparison class is “larger-activation DeepSeek flagship,” those two SKUs are what you can click today, at roughly an order of magnitude more. The next article writes itself when a Pro listing exists. This one will not invent it.

Decide with the Token page, compare, and the on-site model page.

Insights