Letting a model manage its own context is very bitter lesson-pilled
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
_jayhack_ · · focus · HN ↗
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
wangii · · focus · HN ↗
wat10000 · · focus · HN ↗
wangii · · focus · HN ↗
[dead]