runaiicloud
ModelsAppsGPUsServerless GPUPricingDocsConnectCompareEnterprisePlayground
Log inGet started
home/blog/890 bytes per token: how DeepSeek V4.1-Flash's KV cache works
2026-10-09·runaii engineering·13 min read

890 bytes per token: how DeepSeek V4.1-Flash's KV cache works

DeepSeek V4.1-Flash stores its KV cache in about 890 bytes per token — roughly 3.9x less than the previous generation. We tear down the published architecture behind that number: the causal encoder-decoder split, cross-layer KV sharing, FP4 quantization, and the sparse indexer that makes it all retrievable.

deepseekkv-cachearchitectureresearch
890 bytes per token: how DeepSeek V4.1-Flash's KV cache works — illustrated summary card

tl;dr

The single most important number in the DeepSeek V4.1-Flash tech report isn't a benchmark — it's a memory footprint. The model keeps its global key-value cache at roughly 890 bytes per token, against approximately 3,514 bytes per token for the previous V4-Flash generation, according to the report and the community architecture teardowns that followed it. That ~3.9x shrink is why a 1M-token context is even servable, why DeepSeek can price cached prompt reads at $0.006/1M tokens (and half that off-peak), and why the September-2026 "is this the cheapest frontier-class API ever" discourse happened at all. This post is a plain-language teardown of the four published mechanisms behind the number — the causal encoder-decoder split, cross-layer KV sharing, FP4-quantized cache trained in with the model, and the hierarchical sparse indexer that keeps retrieval accurate — plus the engineering tradeoffs each one buys and pays. Everything here comes from the public tech report and public teardowns; links throughout.

Bar chart comparing KV cache bytes per token: DeepSeek V4-Flash about 3,514 bytes versus V4.1-Flash about 890 bytes — a 3.9x reduction

runaii.cloud

Start free — $5 in credits, no application

Per-token serverless inference on open models behind one OpenAI-compatible URL. Cached prefixes bill for almost nothing.

Create an account →See pricing

Why cache size is the whole cost story

Every autoregressive transformer pays the same tax: to generate the next token, it must attend over the entire context, and to do that fast, the key-value projections of every past token must sit in memory. That state — the KV cache — grows linearly with context and with model depth. For a big dense model it's tens of kilobytes per token; even efficient variants have historically cost multiple kilobytes. The consequences cascade through everything a user feels:

  • Context ceilings. If a token costs 3.5KB, a 1M-token context needs ~3.5GB of state per sequence — before you've stored weights, activations, or anyone else's request. Serving many long sessions at that rate is economically impossible, which is why million-token contexts were, for most of the industry, a spec-sheet fantasy.
  • Batch sizes. GPU memory given to cache is memory not given to batching. Smaller caches mean more concurrent sequences per card, which means better throughput and lower per-token cost for everyone.
  • Cache-hit pricing. Prompt caching — reusing the KV state of a repeated prefix across requests — is only generous when the stored state is cheap. You cannot bill $0.003/1M for cached reads if the cache eats premium memory like the old generations did.

So when DeepSeek's report lands on ~890 bytes per token of global cache state, it's not an incremental optimization. It's the enabling fact behind the model's entire economic pitch. Here is how they got there, mechanism by mechanism, as published.

Mechanism 1: the causal encoder-decoder split

V4.1-Flash is a 552B-parameter mixture-of-experts model with an unusual spine: a YOCO-style causal encoder-decoder architecture — roughly 20 layers acting as a causal encoder and 20 as a decoder, sharing information through a compact global representation. The community teardown (linked below) describes the encoder building a shared global KV state that the decoder layers attend to, rather than every layer maintaining its own full copy.

Two things fall out of that design. First, the KV state that actually needs to be kept is the shared global one — the per-layer duplication that inflates cache size in conventional stacks is structurally reduced from the start. Second, the report indicates prefill compute for a fresh context runs at roughly half of full-attention cost at long ranges (about 51% at 131K tokens in the report's framing) — the encoder does the expensive global pass once, and later layers reference its output instead of redoing the work.

For agent workloads — where a huge, stable prefix (system prompt, repo context, tool definitions) is re-sent constantly — this is the architectural soul of the "cheap agent" story: the expensive part of the context is computed once into a compact state, and everything downstream reuses it.

Mechanism 2: cross-layer KV sharing — only 4 source layers

The second published trick compounds the first. In a conventional transformer, each of the model's layers keeps its own key-value projections; cache size scales with layer count. V4.1-Flash instead uses cross-layer KV sharing with only four KV source layers: per the teardown, just a handful of layers produce the canonical key-value states, and the other layers consume shared/derived versions.

The memory math is direct. If 4 of the decoder layers carry the stored KV burden rather than dozens, the per-token cache drops by nearly an order of magnitude before quantization even enters the picture. The quality risk — different layers usually want different attention states — is the obvious objection, and it's why the model was trained with this structure from the start rather than retrofitted: the sharing pattern is learned, not imposed post-hoc. The report's benchmark section is the answer to "what does quality cost": SWE-bench Verified-class coding results in the high-50s and strong long-context retrieval scores, which is to say, the sharing is lossy in principle but trained-away in practice on the benchmarks that matter.

Mechanism 3: FP4 KV via quantization-aware training

Third: the KV values themselves are stored in FP4 (MXFP4) precision, achieved through quantization-aware training — the model learned to keep its attention math accurate while the cache lives in 4-bit floating point.

This matters more than it sounds. Post-training quantization of KV caches is a known trick, but it degrades long-context retrieval — the tiny per-token errors compound over a million positions. Training with the quantized cache in the loop means the model's weights adapted to the precision loss before deployment, which is the difference between "quantized cache as a lossy compression hack" and "quantized cache as the model's native memory format." FP4 versus, say, 16-bit state is another ~4x on the footprint — the single largest multiplier in the whole stack, stacked on top of the layer-sharing reduction.

The industry context makes this the least surprising of the four mechanisms: MXFP4 expert weights and FP8/FP4 cache paths are where the entire open-weights ecosystem is heading in 2026, and DeepSeek's contribution is demonstrating it at flagship scale with a QAT process that defends the quality.

Mechanism 4: the hierarchical sparse indexer

Compact storage is worthless if you can't find what you stored — a 1M-token context with byte-cheap KV still needs attention to pick the relevant few thousand tokens per generation step. V4.1-Flash's answer, per the report and teardown, is a hierarchical sparse indexer: a learned, coarse-to-fine retrieval structure over the context with a candidate pool of roughly 16,384 positions that each generation step attends over.

Think of it as a two-stage search: the indexer first narrows the million positions down to a pool of plausible candidates, then full-precision attention runs only over that pool. The compute saving at long context is dramatic — attention cost scales with the candidate pool, not the full context — and it pairs naturally with the shared cache: the indexer's keys and the FP4 main cache are separate structures, so retrieval precision doesn't inherit the main cache's quantization.

The tradeoff is the most intellectually interesting one: sparse attention over a candidate pool is an approximation of full attention, and adversarial needle-in-haystack patterns (the answer sitting just outside the retrieved pool) are its known weak spot. The report's long-context evals argue the approximation holds on realistic distributions; long-context benchmark culture argues the adversarial cases exist. Both are true. For coding agents — where relevance is lumpy (the file you're editing, the error string you're chasing) rather than adversarially diffuse — the sparse bet is about as favorable as the approximation game gets.

The bonus layer: Engram conditional memory

The teardown also describes an Engram conditional-memory module — a reported ~196B-parameter auxiliary component that conditionally stores and retrieves session state, separate from the KV path. In plain terms: a learned memory that decides what's worth keeping beyond the attention window. It is the least-documented piece of the architecture (the tech report is the primary source; third-party details are thin), so we flag it honestly as "reported, structure not fully public" rather than pretending to a teardown we don't have. Its direction is clear though: the industry is converging on explicit memory management as the next lever after raw cache compression — keeping the right tokens matters as much as shrinking them.

What the numbers mean on an API bill

Architecture is upstream of pricing, and DeepSeek's official API makes the linkage explicit. As of October 2026, deepseek-flash bills $0.30/1M fresh input, $0.006/1M cached reads, $1.20/1M output at peak — and everything halves off-peak (UTC 01:00–04:00 and 06:00–10:00 on weekdays), taking cache hits to $0.003/1M. Those cache prices are only sustainable because the stored state is ~890 bytes/token: generosity in the price column is frugality in the memory column. For workload math at agent scale — sessions, heavy days, and how this stacks against other providers — we ran the full comparison in Best LLM API for coding agents (2026).

Worth reading alongside it: the September freak-out post that put this model at the center of the cost discourse got several things right and several things loud — we fact-checked it here.

If you want to try the model behind an OpenAI-compatible endpoint with transparent cache accounting, runaii.cloud serves DeepSeek V4.1-Flash at $0.22/1M fresh input, $0.022/1M cached, $0.66/1M output, with the cache split visible in every response — pricing here.

What's lost: the honest tradeoff ledger

Every mechanism above trades something. The ledger as published:

  • Sparse retrieval is approximate. Candidate-pool attention can miss adversarial placements, and no indexer is perfect. Realistic workloads look fine; needle-in-haystack edge cases are the known failure mode.
  • Shared KV is a straitjacket on layer specialization. QAT clawed the benchmarks back, but the architecture forecloses some per-layer attention diversity that dense models get for free.
  • FP4 is precision ceiling. Retrieval-sensitive tasks (exact-match over huge contexts, some retrieval-augmented setups) are where 4-bit state would show stress first if anywhere.
  • Complexity moves to the serving stack. A model with a causal encoder split, four KV source layers, a sparse indexer, and a conditional memory module asks much more of whoever serves it. That's invisible to API users but is why "we serve the official weights" is a meaningful claim in the serving market.

None of these are gotchas — they're the visible costs of the invisible subsidy. The 890 bytes is real, and so is the bill of goods it comes with.

The math, end to end

It's worth composing the four mechanisms into one back-of-envelope, because the composition — not any single trick — is what produces ~890 bytes. Take a conventional decoder of this class as the baseline: call it a few dozen layers, each storing key and value projections per token at 16-bit or 8-bit precision — the territory of multiple kilobytes per token, in line with the ~3,514 B/token the teardown attributes to V4-Flash. Now stack V4.1-Flash's reductions:

  • Cross-layer sharing replaces "every layer stores its own KV" with "four source layers store canonical state" — if the baseline's footprint scaled with N layers and four carry the burden, that's roughly an N/4-style cut on the stored component before anything else.
  • The causal encoder-decoder split means the long-term global state the decoder attends to is the encoder's compact shared representation — the expensive full-history KV isn't duplicated across the decoder stack at all.
  • FP4 storage quarters the bytes of whatever remains versus 16-bit (with QAT making the model tolerate it).
  • The sparse indexer keeps its own retrieval keys — a comparatively small, bounded structure over positions — while the main cache stays FP4.

Multiply a layer-sharing cut by a precision cut, subtract what the encoder split never stores, and a few-kilobyte baseline collapsing to under a kilobyte stops being mysterious. The design isn't one clever idea; it's four mutually-reinforcing choices, each of which makes the next more valuable — shared state is worth more when it's also 4-bit, and sparse retrieval is worth more when the state it indexes is compact. That compounding is why retrofitting any single trick onto an existing model buys so much less than training with all four from scratch.

Where this sits in the 2026 research landscape

V4.1-Flash did not appear in a vacuum, and framing it alongside the public research makes both clearer. Three threads worth knowing:

  • Eviction policy research. A much-discussed September study, "LRU is harder to beat than the KV-cache papers suggest", argues plain least-recently-used eviction remains surprisingly competitive with learned eviction schemes on agentic workloads — a reminder that what to drop is still an open question even as how much each token costs shrinks.
  • Compression-to-the-limit research. October papers push the same frontier V4.1-Flash industrialized: "A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention" and "REMORY: Learning Residual Memory for Context Compaction" both attack the cache from the "decide what's worth keeping" angle — the same direction as V4.1's Engram module, published as research rather than shipped as product.
  • Serving-engine research. Token-level routing (TokenRouter, arXiv:2610.12242) and sparse-first engines (SparseEngine, arXiv:2609.39068) attack the same economics from the serving side: cheaper decisions about which model and which attention path each token needs, rather than cheaper storage per token.

The read: 2026's inference-cost curve is being pushed from three directions at once — model architecture (DeepSeek's cache design), memory policy (the eviction/compaction literature), and serving policy (routing and sparse engines). V4.1-Flash is the strongest public proof that the first direction can land in a production API at flagship scale; the other two are, for now, mostly papers — which is the standard lag between a good idea and a bill you can check.

Sources and freshness

  • DeepSeek-V4.1-Flash tech report: DeepSeek_V41_Tech_Report.pdf on Hugging Face (release Sept 10, 2026).
  • Community architecture teardown with the per-generation cache-byte analysis: zartbot — DSV41Flash architecture deep-dive (Sept 17, 2026).
  • DeepSeek API pricing (peak/off-peak, cache-hit rates): api-docs.deepseek.com, verified 2026-10-09.
  • DeepSeek V4.1-Flash release notes and model family transitions: api-docs.deepseek.com/news.
  • Adjacent research context on KV compression and eviction: see the papers linked in our fact-check companion post.

Byte-per-token figures are from the sources above as of their publication dates; where the report and teardowns frame numbers differently (per-layer vs global state), we follow the teardown's global-cache framing and note the ambiguity.

Changelog

  • 2026-10-09: initial publication, built from the V4.1 tech report, the zartbot teardown, and DeepSeek's public pricing pages (all verified this date). Will update if the tech report is revised or a V4.1-Pro variant lands with a different cache design.

FAQ

▸How much KV cache does DeepSeek V4.1-Flash use per token?

Roughly 890 bytes of global KV state per token, versus about 3,514 bytes per token for the previous V4-Flash generation — a ~3.9x reduction, per the V4.1 tech report and the community architecture teardowns. For scale, that puts a 200K-token context at roughly 178MB of cache state instead of ~703MB.

▸Why does a smaller KV cache make the API cheaper?

The KV cache lives in GPU memory, so its size determines how many concurrent requests fit per card, how big contexts can get, and whether prompt caching is affordable to offer. A cache ~3.9x smaller lets a serving fleet store more sessions, support the 1M-token window, and bill cached reads at $0.006/1M or less — the frugality shows up directly in the price column.

▸What is cross-layer KV sharing?

A design where only a few layers (four, in V4.1-Flash's case, per the teardown) compute and store the canonical key-value states, while other layers reuse or derive from them, instead of every layer keeping its own copy. It cuts cache size nearly proportionally to the sharing ratio; the quality cost is handled by training the model with the sharing built in rather than retrofitting it.

▸Does FP4 KV cache hurt quality?

Not measurably on the reported benchmarks — because DeepSeek used quantization-aware training, so the model adapted to 4-bit cache during training rather than having it imposed afterward. Post-training KV quantization is the riskier version of this trick; QAT is the version that defends the benchmark sheet. Retrieval-stress cases over million-token contexts remain the place to watch.

▸Is the sparse indexer the same as prompt caching?

No — they're complementary. Prompt caching reuses KV state for repeated prefixes across requests (a billing and reuse mechanism). The sparse indexer decides which parts of the current context each generation step attends over (a compute mechanism). V4.1-Flash uses both: cheap storage makes caching generous, and sparse retrieval keeps long-context attention affordable.

▸Can other models adopt these techniques?

Individually, yes — YOCO-style encoder splits, cross-layer sharing, QAT-quantized cache, and sparse retrieval are all active research areas with public papers (we link several in the companion post). The hard part is adopting them together in one trained model: the tricks interact, and the QAT process has to absorb all of the constraints simultaneously. That integration, at flagship scale, is the actual V4.1 achievement.

Glossary

KV cache
the stored key-value projections of every past token that attention needs to generate the next token; grows linearly with context length and dominates serving memory at long ranges.
YOCO (You Only Cache Once)
an architecture family where early layers build a shared global KV state that later layers attend to, instead of every layer keeping its own — the backbone of V4.1-Flash's causal encoder-decoder split.
Cross-layer KV sharing
storing canonical KV states in only a few source layers while other layers consume shared or derived states, cutting cache size by roughly the sharing ratio.
Quantization-aware training (QAT)
training a model with quantized values (here, FP4/MXFP4 cache) in the loop so weights adapt to the precision loss before deployment, instead of losing accuracy to post-hoc quantization.
MXFP4
a 4-bit floating-point block format (microscaling) used for weights and, in V4.1-Flash, the main KV cache.
Sparse indexer
a learned coarse-to-fine retrieval structure that narrows a long context to a small candidate pool (about 16,384 positions in V4.1-Flash) for each attention step, making million-token attention tractable.
Prefill
the compute pass that processes a fresh prompt into KV state before generation; agent workloads are prefill-heavy, which is why cache pricing and prefill efficiency dominate their bills.

runaii.cloud

Start free — $5 in credits, no application

Per-token serverless inference on open models behind one OpenAI-compatible URL. Cached prefixes bill for almost nothing.

Create an account →See pricing
Keep reading
The DeepSeek V4.1-Flash freak-out, fact-checked2026-10-09Best LLM API for coding agents (2026): priced for real agent workloads2026-10-09

On this page

Why cache size is the whole cost storyMechanism 1: the causal encoder-decoder splitMechanism 2: cross-layer KV sharing — only 4 source layersMechanism 3: FP4 KV via quantization-aware trainingMechanism 4: the hierarchical sparse indexerThe bonus layer: Engram conditional memoryWhat the numbers mean on an API billWhat's lost: the honest tradeoff ledgerThe math, end to endWhere this sits in the 2026 research landscapeSources and freshnessChangelog
runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryAppsGPUsServerless GPUPricingToken-Max coding plansCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy