Letting a model manage its own context is very bitter lesson-pilled
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
True facts. Cline does this with auto compact and it's so dumb. At the start of the session it reads the plan.md. during compaction, it forgets it. Same with its own output. It'll tell me the list of steps to take. By step 4 it needs to compact. So I prompt to do step 4. It spends 15 minutes searching the codebase and comes back that it couldn't find such a step. So dumb. So so dumb
_jayhack_ · · focus · HN ↗
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
Neywiny · · focus · HN ↗