‹ BackHN Continuity

Thread

Context Language Models

176 points · 51 comments · emersonmacro

  1. _jayhack_ · · focus · HN ↗
    Letting a model manage its own context is very bitter lesson-pilled

    Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.

    This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure

    See related: &quot;KV Cache Rules Everything Around Me&quot;: <a href="https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-around" rel="nofollow">https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-...

    1. wangii · · focus · HN ↗
      is kvcache a prematural optimization?
      1. f_devd · · focus · HN ↗
        It&#x27;s pretty much required to make it work with large contexts, so probably not
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.