Letting a model manage its own context is very bitter lesson-pilled
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
this is done on the model side so it is invisible to the caller. all of deepseeks kv cache compression stuff is basically the same thing. it doesn't really put the entire context into kv cache and does various inference side stuff to decide what to actually use as context.
_jayhack_ · · focus · HN ↗
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
metalspot · · focus · HN ↗