‹ BackHN Continuity

Thread

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

131 points · 10 comments · mfiguiere

  1. arikrahman · · focus · HN ↗
    I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free.
  2. smy20011 · · focus · HN ↗
    Removed
    1. sebmellen · · focus · HN ↗
      The writing feels human to me… and I call out AI slop as much as possible.
    2. girvo · · focus · HN ↗
      The website design definitely is, but I don’t know if the content is? This reads pretty human to me and is quite interesting to boot!
    3. yunfei · · focus · HN ↗
      you don't know him?
  3. vivzkestrel · · focus · HN ↗
    404 on the blog page? <a href="https:&#x2F;&#x2F;zartbot.github.io&#x2F;blog&#x2F;" rel="nofollow">https:&#x2F;&#x2F;zartbot.github.io&#x2F;blog&#x2F;
    1. sanufar · · focus · HN ↗
      Yeah, for some reason their &#x2F;blog&#x2F; is 404ing but for anyone interested in their other articles, <a href="https:&#x2F;&#x2F;github.com&#x2F;zartbot&#x2F;blog&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;zartbot&#x2F;blog&#x2F; has all of them, just sub out anything after &#x2F;blog&#x2F;.* with the folder path (e.g <a href="https:&#x2F;&#x2F;zartbot.github.io&#x2F;blog&#x2F;arch&#x2F;jalapeno&#x2F;" rel="nofollow">https:&#x2F;&#x2F;zartbot.github.io&#x2F;blog&#x2F;arch&#x2F;jalapeno&#x2F; from <a href="https:&#x2F;&#x2F;github.com&#x2F;zartbot&#x2F;blog&#x2F;tree&#x2F;main&#x2F;arch&#x2F;jalapeno" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;zartbot&#x2F;blog&#x2F;tree&#x2F;main&#x2F;arch&#x2F;jalapeno)
  4. N_Lens · · focus · HN ↗

    [dead]

  5. mmastrac · · focus · HN ↗
    I&#x27;ve been working with an automatic incremental context compactor enabled and it&#x27;s been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence.

    TBH I also ran the 400tok&#x2F;s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems

    1. nchmy · · focus · HN ↗
      Can you share a link to thr automatic compactor? I&#x27;ve been noticing that when I get to around 60% context window, the cache will simply break and suddenly I&#x27;ve paid 50x more than expected. The only solution seems to be to compact or start a new session.
      1. mmastrac · · focus · HN ↗
        Send me an email- it&#x27;s not public just yet
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.