DeepSeek's new V4.1-Flash model slashes KV cache to a quarter of its predecessor

DeepSeek v4.1 Flash

DeepSeek's new V4.1-Flash model slashes KV cache to a quarter of its predecessor

DeepSeek has announced DeepSeek-V4.1-Flash, the smallest model in its new architecture family, featuring native visual understanding and an asymmetric Causal Encoder–Decoder design. With 552B total parameters but only 8B active for input and 16B for output, it delivers faster inference and higher throughput. The KV cache requires just one-quarter the HBM and one-eighth the SSD storage of the previous generation, cutting agent costs. It's now live on the DeepSeek API, with off-peak pricing at half the peak rate.

Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.
  1. kouteiheika

    It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers.

    [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...

    [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

  2. rao-v

    As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.

    I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.

    They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.

  3. k9294

    I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the chat completion API obsolete because the cost of context transfer over network is going to dominate the task total cost?

    Here's the same token usage priced at different rates: a real long-running coding task, medium codebase, 447 turns.

    Input 1,026,957

    Output 164,667

    Cache read 36,554,368

    GPT-6-astra

    Type Rate Cost Share

    Input 10.000 10.270 19%

    Output 50.000 8.233 15%

    Cache 1.000 36.554 66%

    Total 55.057 100%

    DeepSeek v4.1 Flash, $0.003 cache hit

    Type Rate Cost Share

    Input 0.300 0.308 50%

    Output 1.200 0.198 32%

    Cache 0.003 0.110 18%

    Total 0.615 100%

    DeepSeek v4.1 Flash, $0.006 cache hit

    Type Rate Cost Share

    Input 0.300 0.308 42%

    Output 1.200 0.198 27%

    Cache 0.006 0.219 30%

    Total 0.725 100%

    Hypothetical: same DeepSeek input/output rates, but cache priced so it accounts for 66% of the bill.

    Type Rate Cost Share

    Input 0.300 0.308 21%

    Output 1.200 0.198 13%

    Cache 0.027 0.982 66%

    Total 1.487 100%

    This cache it improvement makes the model x2-x2.5 more efficient on a long horizon tas […]

  4. revolvingthrow

    Already on HuggingFace: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

    The bad news is that the original v4 flash was 284B, which was large but still somewhat reasonable for running locally. This one is 552B so almost twice that, so the huge gains in benchmark scores make sense - it's not really flash anymore, imo.

    I've no idea about actual performance vs benchmaxxing, though deepseek was fairly trustworthy as far as Chinese models go. If that holds (and if it doesn't think forever, as deepseek 4 sometimes did) it's probably the newest king of the hill amongst open weights models.

    It does include vision, and they do something funky with KV cache so it's very efficient: "[...] these designs reduce the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash". I do appreciate the high focus on efficiency, but at this point we sure could use a flash-flash version.

    @edit: I couldn't make sense what the actual parameter count is, with the addition of Engram memory. To my understanding the 4.1 flash is 552B parameters you want in vram or ram, out of which ~16B is active (8B for prefill). It also includes additional 196B Engram memory which you can put on an SSD. I think.

    Assuming that's correct 256 GB memory is insufficient to even load the model at q4 - you'd be 1GB short, assuming you can fill it to 100% (so no mac). You'd also want some for kv cache of course. A 256 GB desktop with some extra VRAM from GPU could run it, but normal consumer boar […]

  5. wren6991

    That's a lot of architectural innovation for a .1 release! I guess there's precedent there: they introduced sparse attention (DSA) in V3.2.

  6. XCSme

    Seems just slightly better than last v4 release, considerably (3x) more expensive, but also faster and slightly more token efficient.

    https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-hi...

  7. impulser_

    I think it's very clear that DeepSeek is obviously the best AI lab in the world.

    Every model release seems like it packed with wonderful research and advancements.

  8. LaurensBER

    Initial impressions: this is a really strong model and the fact that they reduced prices at the same time makes it an awesome backup model to use when your primary subscription runs out and you need to bridge a few days before it resets.

    It also seems to be more willing to just do whatever you ask of it. My favourite benchmark for this is to ask it to download a rom for an old game, that I own. Legal in my juristiction but the US models (except Grok) have a tendency to refuse it.

  9. mentalgear

    https://xcancel.com/deepseek_ai/status/2097930608790167907

    Should be the link ( now that it works again! :) )

  10. cdnsteve

    Absolutely insane performance and benchmark results. It's beating Opus 5 and Sol 5.6 https://tokenstead.ai/models/deepseek-v4-1-flash

More from this day

2026-09-10