‹ BackHN Continuity

Thread

Context Language Models

176 points · 51 comments · emersonmacro

  1. _jayhack_ · · focus · HN ↗
    Letting a model manage its own context is very bitter lesson-pilled

    Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.

    This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure

    See related: &quot;KV Cache Rules Everything Around Me&quot;: <a href="https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-around" rel="nofollow">https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-...

    1. wangii · · focus · HN ↗
      is kvcache a prematural optimization?
      1. f_devd · · focus · HN ↗
        It&#x27;s pretty much required to make it work with large contexts, so probably not
      2. gfrecvh · · focus · HN ↗
        Do you even understand what you&#x27;re suggesting? Do you understand how the data flow would look like without a kv cache?
        1. wangii · · focus · HN ↗
          Java spent decades making pointer-rich identity semantics cheaper, before increasingly recognizing that many things programmers call “objects” are actually values.

          We may spend a decade making persistent token-level KV state cheaper, before recognizing that much of what we call “context” should actually be transient, compressed, reconstructed, or represented in an entirely different state space.

          KV cache risks turning an optimization of today&#x27;s representation of history into tomorrow&#x27;s semantics of memory.

      3. wat10000 · · focus · HN ↗
        Not having it would be like a text editor that loads your entire document from disk and re-saves it every time you press a key. On a floppy drive.
        1. wangii · · focus · HN ↗

          [dead]

    2. metalspot · · focus · HN ↗
      this is done on the model side so it is invisible to the caller. all of deepseeks kv cache compression stuff is basically the same thing. it doesn&#x27;t really put the entire context into kv cache and does various inference side stuff to decide what to actually use as context.
    3. trenchgun · · focus · HN ↗
      If you read the paper, they had a solution for this.
    4. Neywiny · · focus · HN ↗
      True facts. Cline does this with auto compact and it&#x27;s so dumb. At the start of the session it reads the plan.md. during compaction, it forgets it. Same with its own output. It&#x27;ll tell me the list of steps to take. By step 4 it needs to compact. So I prompt to do step 4. It spends 15 minutes searching the codebase and comes back that it couldn&#x27;t find such a step. So dumb. So so dumb
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.