‹ BackHN Continuity

Thread

Context Language Models

176 points · 51 comments · emersonmacro

  1. _jayhack_ · · focus · HN ↗
    Letting a model manage its own context is very bitter lesson-pilled

    Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.

    This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure

    See related: &quot;KV Cache Rules Everything Around Me&quot;: <a href="https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-around" rel="nofollow">https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-...

    1. wangii · · focus · HN ↗
      is kvcache a prematural optimization?
      1. gfrecvh · · focus · HN ↗
        Do you even understand what you&#x27;re suggesting? Do you understand how the data flow would look like without a kv cache?
        1. wangii · · focus · HN ↗
          Java spent decades making pointer-rich identity semantics cheaper, before increasingly recognizing that many things programmers call “objects” are actually values.

          We may spend a decade making persistent token-level KV state cheaper, before recognizing that much of what we call “context” should actually be transient, compressed, reconstructed, or represented in an entirely different state space.

          KV cache risks turning an optimization of today&#x27;s representation of history into tomorrow&#x27;s semantics of memory.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.