‹ BackHN Continuity

Thread

Context Language Models

176 points · 51 comments · emersonmacro

  1. _jayhack_ · · focus · HN ↗
    Letting a model manage its own context is very bitter lesson-pilled

    Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.

    This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure

    See related: &quot;KV Cache Rules Everything Around Me&quot;: <a href="https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-around" rel="nofollow">https:&#x2F;&#x2F;www.completeskeptic.com&#x2F;p&#x2F;kv-cache-rules-everything-...

    1. wangii · · focus · HN ↗
      is kvcache a prematural optimization?
      1. wat10000 · · focus · HN ↗
        Not having it would be like a text editor that loads your entire document from disk and re-saves it every time you press a key. On a floppy drive.
        1. wangii · · focus · HN ↗

          [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.