DeepSeek V4 Flash 0731 deep dive: 1M-context MoE at $0.09/$0.18 on OpenRouter

Sparse MoE at 284B/13B active, 1M-token context, reasoning and tools, priced $0.09/$0.18 on the OpenRouter primary listing. HF Hot shows +64% growth; catalog first seen 2026-08-01. AI Hippo composite rank #9 (a context-per-dollar proxy, not a capability score).

DeepSeek V4 Flash — high-speed sparse MoE concept Enlarge image
Flash: sparse activation aimed at long-context throughput

What it is

DeepSeek V4 Flash 0731 (`deepseek/deepseek-v4-flash-0731`) is the efficiency SKU on DeepSeek’s V4 line. Catalog copy describes a sparse mixture-of-experts model with 284B total parameters and about 13B activated, aimed at coding, reasoning, and agent workflows, with a 1M-token context window. Versus sibling V4 Pro (1.6T total / 49B activated), Flash trades activation scale for throughput and a much lower sticker.

Site `ops_model_first_seen` records the 0731 revision on 2026-08-01; the earlier Flash 0423 cut and V4 Pro both first appear on 2026-06-22. Catalog text calls 0731 a re-post-trained revision—treat it as a pin-dated family SKU, not a rebranded model line.

On the community side, this snapshot’s HF Hot list places DeepSeek-V4-Flash-0731 immediately after Kimi-K3 (sort_index=1) with +64% 24h growth and about 236K HF downloads; an unsloth GGUF derivative also charts. Momentum is real, but downloads are not a production readiness certificate—decide on specs and multi-host quotes.

Spec sheet at a glance

Sparse MoE: 284B total, ~13B activated Enlarge image
Most experts idle; a few paths light up

By the numbers (`core_model`): context length 1,048,576; modality `text->text` (text in and text out—not multimodal). The API surface exposes 20 parameters, including frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, and top_p.

Versus Flash 0423: same 1M window and text→text stack, with a nearly aligned parameter list (both expose reasoning_effort and top_a). Versus V4 Pro: same 1M context, a slightly narrower parameter list (no top_a), and a much larger activation footprint positioned as the flagship MoE. Knowledge cutoff for 0731 is undisclosed in this snapshot. Vendor: DeepSeek.

Pricing and where to run it

V4 family price-ladder sketch Enlarge image
0731 at the bottom rung: cheapest in-family tier

Flash 0731: OpenRouter primary listing (is_primary_listing=1) at $0.09 input and $0.18 output per 1M tokens. Fireworks and Together quote the same model at $0.14/$0.28 (model_bound_verified)—do not assume the $0.09 sticker everywhere. Fireworks also exposes cached input at $0.028 per 1M.

Family peers on OpenRouter primary: Flash 0423 at $0.14/$0.28; V4 Pro at $0.435/$0.87; earlier V3.2 about $0.269/$0.40. Versus 0423, 0731 is about 36% cheaper on input (and the same ratio on output)—the revision carries a price signal as well as a training signal.

DeepInfra / deepseek-direct rows in this snapshot mainly attach to 0423 and Pro; verified multi-host coverage for 0731 centers on OpenRouter + Fireworks + Together. Re-check the Token page before production.

Inside the family: Flash 0423 → Flash 0731 → V4 Pro

DeepSeek V4 ladder in this snapshot (OpenRouter primary listing, 1M context aligned): Flash 0731 blends to about $0.135 per 1M; Flash 0423 about $0.21; V4 Pro about $0.65. Activation footprint: Flash line ~13B versus Pro ~49B; total parameters 284B versus 1.6T.

On AI Hippo’s composite board (a context-per-dollar proxy, not a capability benchmark): Flash 0731 sits around #9; the `~deepseek/deepseek-v4-flash-latest` alias around #8; Flash 0423 around #72; V4 Pro around #75. The proxy rewards cheap long context—exactly 0731’s profile—so it outranks the far pricier Pro.

Rule of thumb: default long-context batch and agent drafts → 0731; must pin an older route or eval baseline → 0423; want the larger-activation catalog flagship and will pay roughly 4.8× the Flash blend → V4 Pro.

How it compares across vendors

Blended cost (OpenRouter primary listing mean): Flash 0731 about $0.135 per 1M (1.0M ctx, text→text). Peers: Llama 4 Scout about $0.20 (1.3M, composite #2), GPT-5.6 Luna Pro about $0.35, Qwen3.7 Flash about $0.08, MiniMax M3 about $0.75, GLM-5.2 about $1.59, Kimi K3 about $9, Grok 4.5 about $4, Claude Sonnet 5 about $6.

0731 lands in the ultra-cheap long-context band alongside Scout / Luna / Qwen Flash, with stronger HF-hot momentum and a clearer DeepSeek family story. Caveat: 0731 is text-only—pick a multimodal SKU when you need image or file inputs. These are spec and cost comparisons, not capability benchmarks.

Cost in practice: two scenarios

Modeled with this snapshot’s OpenRouter primary listing (excluding extra reasoning tokens).

Scenario A—agent coding turn (100K input + 20K output): Flash 0731 about $0.013; Flash 0423 about $0.020; Llama 4 Scout about $0.016; Luna Pro about $0.022; V4 Pro about $0.061; MiniMax M3 about $0.054; Kimi K3 about $0.60; Sonnet 5 about $0.40.

Scenario B—long-document pass (500K input + 5K output): 0731 about $0.046; 0423 about $0.071; Scout about $0.052; Luna about $0.053; Pro about $0.222; K3 about $1.58.

Takeaway: 0731 pushes million-token agent drafts into cent territory. Versus K3/Sonnet the gap is an order of magnitude; versus Scout it is still meaningfully cheaper. Output is only 2× input—output-heavy loops hurt less than on 5×-output tiers. These are spec and cost comparisons, not capability benchmarks.

When to choose it—and when not to

Choose Flash 0731 when you need 1M-context coding, reasoning, or agent drafts on a razor budget and text→text is enough; when you want the hottest DeepSeek efficiency SKU in this HF snapshot; and when you can start on OpenRouter’s $0.09/$0.18 listing. Pin the dated 0731 model ID—do not silently float on a latest alias.

Skip or look elsewhere when you need image/file inputs (pick multimodal); when you want the larger-activation flagship story and will pay for it (V4 Pro); when an eval harness is locked to 0423; or when Fireworks/Together is the stabler path and you accept $0.14/$0.28—route by host quote, not by the OpenRouter sticker alone.

Decide with AI Hippo’s model, Token, compare, and Trends pages; the composite board is a context-per-dollar proxy, not a capability benchmark.

Deployment checklist

Before you integrate: route explicitly to `deepseek/deepseek-v4-flash-0731`, kept distinct from 0423, Pro, and `~deepseek/deepseek-v4-flash-latest`; confirm the harness supports reasoning, reasoning_effort, tools, and structured_outputs; budget with your real input/output ratio and branch by host (OpenRouter $0.09/$0.18 versus Fireworks/Together $0.14/$0.28).

Operationally: monitor tokens per turn including reasoning; load-test truncation and latency on long-context paths; re-check the Token page and Trends (HF Hot). Treat SQLite snapshot prices as source of truth—verify again before you ship.

Insights