DeepSeek’s change log for September 10, 2026 reads like a quiet product update and a strategic announcement at the same time. The product is DeepSeek-V4.1-Flash, “the smallest model in our new architecture family, with native multimodal visual understanding.” The strategy is the rest of the page: V4-Flash and V4-Flash-Vision-Exp are retired, deepseek-v4-pro is being routed to V4.1 Flash at V4.1 Flash prices starting September 14, and the new architecture is “designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models.” This isn’t a model release. It’s the foundation model for the next family — and the cheapest member is the one that shipped first.
For an AI engineer running multi-agent pipelines, three things from the announcement matter before anything else: the asymmetric 8B/16B active-parameter split, the 890-byte KV cache footprint, and the fact that the model that just surpassed V4-Pro is being billed at the Flash price.
The architecture, in one diagram
V4.1 Flash is a 552B-parameter multimodal MoE with 1 shared expert plus 384 routed experts per MoE layer, activating 6 routed experts per token. The Causal Encoder–Decoder (CED) split is the new part: an asymmetric active-parameter budget — 8B for input, 16B for output. That isn’t two MoEs glued together; it’s one backbone with separate capacity allocated to the prefill pass (where context understanding happens) and the decode pass (where token-by-token reasoning happens).
The intuition behind asymmetric active counts is that prefill and decode have different compute profiles. Prefill is dominated by long-context attention against the prompt; decode is dominated by reasoning depth per generated token. Allocating fewer active parameters to prefill and more to decode means the same total parameter budget can deliver lower first-token latency without compromising reasoning quality — at the cost of a more complex training recipe (the encoder and decoder halves need their own load-balancing, and the routing has to be aware of which pass it’s serving).
The third orthogonal subsystem is Engram conditional memory: 196B parameters, sparsely accessed via token-based lookup. This is a learned associative memory that’s consulted during inference, distinct from the routed experts. It’s why the headline total (552B) is so much higher than the active budget (8B/16B): the routed experts are the model’s reasoning, and Engram is its lookup-table memory.
The fourth subsystem, DSpark speculative decoding, sits at decode time: a semi-autoregressive draft model generates candidate continuations, and a confidence-scheduled verifier decides which to keep. DSpark is what makes the 16B-active decode pass practical to serve — the draft model absorbs most of the per-token work, and the verifier only engages on uncertain continuations.
The fifth, Mega-mHC, is the kernel. “Mega-mHC” is DeepSeek’s manifold-constrained Hopfield network kernel; the HF model card mentions it as the implementation that makes the Engram lookup tractable on modern accelerators. We don’t have the full math in the public materials yet, but the practical consequence is that token-keyed lookups over a 196B-parameter table don’t blow the latency budget.
The training recipe behind all of this is 45T tokens from a multimodal corpus, sparse attention trained at 64K sequence length, and context extended to 1M tokens at 34T tokens. The 64K-then-1M two-stage context extension is the same progressive-context recipe DeepSeek has used since V3; what’s new is that it’s being done in a multimodal corpus from scratch, not adapted from a text-only base.
What the agent benchmarks actually show
The change log lists 17 benchmark numbers. I’ll pull out the agent-relevant ones because those are the cells that matter for anyone running tool-using pipelines:
| Benchmark | V4.1 Flash | Notes |
|---|---|---|
| Terminal-Bench 2.1 | 90.6 | agent harness scoring |
| Terminal-Bench 3.0 | 30.0 | harder version |
| Terminal-Bench 4.0 | 31.2 | even harder version |
| DeepSWE v1.1 | 74.2 | SWE-bench-style code agent |
| ProgramBench | 20.3 | competitive programming |
| NL2Repo-Bench | 65.4 | repo-from-spec |
| SEC-Bench Pro | 62.8 | long-horizon financial tasks |
| ExploitGym | 15.3 | cybersecurity agent |
| HLE (w/tools) | 63.9 | with tools |
| Agents’ Last Exam | 31.8 | hardest agent exam |
| Automation-Bench | 54.8 | browser/desktop automation |
| Chartography (w/tools) | 78.9 | chart understanding w/ tools |
| BabyVision (w/tools) | 89.6 | vision agent w/ tools |
| ZeroBench-main (w/tools) | 49.0 | zero-shot hard w/ tools |
The Terminal-Bench progression is the most interesting cell to sit with. V4.1 Flash scores 90.6 on Terminal-Bench 2.1, then drops to 30.0 on Terminal-Bench 3.0 and 31.2 on Terminal-Bench 4.0. Terminal-Bench is a series where each generation is deliberately harder than the last — version 4.0 is supposed to be near the frontier ceiling. The 90→30→31 collapse tells you exactly where V4.1 Flash sits relative to the frontier: it’s a strong V2-tier agent that doesn’t yet climb to V3-tier difficulty.
That isn’t a criticism; it’s a positioning statement. V4.1 Flash is being billed as a high-throughput Flash-tier model that outperforms its own flagship V4-Pro on cost-adjusted metrics. The benchmark table is honest about what the model can and can’t do. The “98% of Astra at 1.4% the cost” claim that’s been circulating on LinkedIn and Reddit is from a single third-party benchmark (OpenDesign Arena) that I can’t independently verify against DeepSeek’s published numbers — I’d treat that specific framing as marketing rather than a reproducible result until DeepSeek or Artificial Analysis publish it. The verifiable version is closer to: V4.1 Flash sits within a couple of points of GPT-6 Astra on most agent benchmarks, at roughly 7-10% of the input cost and 20% of the output cost.
The KV cache is the actual product
The line that mattered most to me was this one from the change log: “V4.1-Flash’s KV cache needs just 1/4 the HBM and 1/8 the SSD storage of the previous generation.” The HF model card puts a number on it: 890 bytes per token in the global KV cache footprint.
For serving 1M-context requests, the KV cache is the binding constraint. At 890 bytes per token, a full 1M-token conversation fits in about 890 MB — small enough that a single 80GB H100 can hold dozens of concurrent long-context requests, and the residual KV spillover to SSD is cheap. The previous generation (V4-Flash) was roughly 4× that in HBM and 8× that in SSD; the new compression is what makes the 1M context window practical to serve at Flash prices.
The compression is documented in a separate paper at the HF model card: “DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression.” I haven’t read the full paper yet, but the headline is that the team has found ways to push the per-token KV footprint below what most quantization recipes manage (Q4 KV cache typically lands around 4-8 bits per element × 2 elements per token × hidden-dim-of-K-and-V; 890 bytes implies something like 1.5-2 bits per element effective, which is below standard 2-bit KV quantization and well below FP8 KV).
There’s a side observation worth pausing on: the V4-Flash 0731 update (released July 31) was widely covered for its 10-point jump on Artificial Analysis’s Intelligence Index, but the deeper engineering story there was the cache compression — that release is what made 1M-context viable, and V4.1 Flash is the architecture-level refinement of the same compression bet. The 0731 checkpoint and the V4.1 release are the same research line at different maturity levels.
The pricing math
Off-peak API pricing for V4.1 Flash: $0.22 per million input tokens (cache miss), $0.66 per million output tokens. Peak pricing is 2× off-peak. Cache-hit pricing is dramatically cheaper — DeepSeek’s change log notes that cache-hit charges “often account for a large share of agent costs,” and the new cache compression lets them price cache hits aggressively because the marginal cost of serving a cached token is now near-zero.
For comparison, the published Flash pricing before the V4.1 release was $0.14/M input cache-miss and $0.28/M output (from the eesel.ai and Spheron Network pricing breakdowns). V4.1 Flash is priced higher per token than V4-Flash was — the off-peak input rate went from $0.14 to $0.22, output from $0.28 to $0.66. But the cache-hit economics, the KV cache compression, and the asymmetric active-parameter split mean that agent-shaped workloads (long context, high cache hit rate, lots of decoding) get cheaper in practice even though the per-token list price went up.
For a senior engineer choosing between V4.1 Flash and a competitor for an agent pipeline, the comparison that matters isn’t $0.22 vs $X/M input — it’s dollars per completed agent task at a given cache-hit ratio and context length. The numbers from the V4-Flash 0731 era (when it was a “kill line” on the cost-per-weighted-task chart) suggest V4.1 Flash will land similarly aggressive on that metric, but I want to actually run a real workload before claiming it.
What changes for an existing DeepSeek user
Three operational things to do this week if you’re already shipping with DeepSeek models:
-
Rename your API calls from
deepseek-v4-flashtodeepseek-flash. The legacy names still work, but they’re being routed to V4.1 Flash at Flash prices anyway, and the new name is the one that will keep working after September 14. Don’t wait for the routing change to force your hand. -
Update your billing projections. V4.1 Flash is priced higher per token than V4-Flash was. If you have a fixed monthly budget for DeepSeek inference, your dollar burn at the same traffic pattern will go up unless you’re benefiting from the cache compression. The honest comparison is dollars-per-completed-agent-task, not dollars-per-token — but if you’re not measuring that, prepare for a billing surprise.
-
Test against your existing V4-Flash traffic before the September 14 deadline. After that date, every
deepseek-v4-prorequest will be served by V4.1 Flash at V4.1 Flash prices. If you have any code path that callsdeepseek-v4-profor a specific capability (long-context reasoning, particular prompting format), validate that V4.1 Flash reproduces the behavior before the silent switch happens. DeepSeek’s claim that V4.1 Flash beats V4-Pro “across performance, cost, speed, and total time” is the kind of claim that’s true on average and not always true on your specific workload.
What doesn’t change
Two notes on what V4.1 Flash is not:
It’s not a self-hosted release in the sense that you can pull the weights and serve them today. The HF repo (deepseek-ai/DeepSeek-V4.1-Flash) is published under MIT with the model card, the inference code, and the evaluation instructions — but a 552B-parameter MoE isn’t a single-GPU deploy. EXL3 quants are starting to appear (there’s already a TP4 EXL3 build in the HF ecosystem), and the V4.1-Flash inference support is being co-developed with the open-source community per DeepSeek’s note. Realistic self-hosting timelines are weeks-to-months for the major inference engines to ship V4.1-Flash support, not days. If you want the model today, the API is the path.
It’s also not a “small model” in any sense that matters for inference cost. 552B parameters is a frontier-scale backbone; the active parameter counts are Flash-tier, but the memory footprint and the multi-GPU serving requirements are not. Anyone describing V4.1 Flash as “the small DeepSeek model” is confused about what active vs total parameters mean for serving.
Open questions I didn’t get to answer
The technical report PDF is on the HF model card but I haven’t finished reading it. Three things I want to dig into before I claim to understand the architecture:
- The exact mechanism of the Mega-mHC kernel and how it interacts with Engram lookup. The model card mentions Mega-mHC as enabling efficient Engram access, but doesn’t explain whether the lookup is hash-based, learned-key-based, or a hybrid.
- How the asymmetric encoder/decoder load balancing is enforced during training. Naive MoE load-balancing treats all tokens symmetrically; CED has to maintain separate balance losses for the encoder-side 8B budget and the decoder-side 16B budget, and I’d like to see how DeepSeek handles the case where one side is over-utilized and the other is under-utilized.
- Whether DSpark speculative decoding is a separate draft model or a learned head off the main backbone. Confidence-scheduled verification is consistent with either design; the deployment implications are very different (separate draft model = double the VRAM; learned head = integrated).
I’ll write up what I find once I’ve read the paper. For now, V4.1 Flash is a credible bet that the Flash-tier price point can absorb a frontier-class architectural redesign — and DeepSeek has put its money where its mouth is by routing its own flagship V4-Pro to the new model at the new price.
References and where to dig further
- DeepSeek change log (Sep 10, 2026): the canonical source for the benchmark table and the routing changes —
https://api-docs.deepseek.com/updates/ - DeepSeek-V4.1-Flash announcement page (asymmetric architecture, KV cache, multimodal):
https://api-docs.deepseek.com/news/news260910/ - DeepSeek-V4.1-Flash HF model card (multimodal MoE, 552B backbone, 1M context, 384 routed experts / 6 active, 45T training tokens, MIT license):
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash - Technical report PDF: “DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression”, same HF repo
- Pricing details (off-peak $0.22/M input cache-miss, $0.66/M output, peak 2×, cache-hit aggressive):
https://api-docs.deepseek.com/quick_start/pricing - DeepSeek Harness guide (V4.1 Flash ships with native DeepSeek Harness integration; the agent benchmarks above use Minimal mode):
https://api-docs.deepseek.com/guides/ - Earlier reference: V4-Flash 0731 (Jul 31, 2026) and the Artificial Analysis Intelligence Index progression that put Flash-tier models within a point of GLM-5.2 — same research line, earlier checkpoint
The V4.1 architecture is going to take a few weeks to fully land in the open-source inference stack. If you’re already on DeepSeek, the September 14 routing change is the date that matters. If you’re not, this is the release to actually read the technical report for — the asymmetric active-parameter split is a real architectural idea, not just a marketing claim, and the KV cache compression is what makes the 1M context window serveable at Flash prices.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.