Letting a model manage its own context is very bitter lesson-pilled
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
Java spent decades making pointer-rich identity semantics cheaper, before increasingly recognizing that many things programmers call “objects” are actually values.
We may spend a decade making persistent token-level KV state cheaper, before recognizing that much of what we call “context” should actually be transient, compressed, reconstructed, or represented in an entirely different state space.
KV cache risks turning an optimization of today's representation of history into tomorrow's semantics of memory.
this is done on the model side so it is invisible to the caller. all of deepseeks kv cache compression stuff is basically the same thing. it doesn't really put the entire context into kv cache and does various inference side stuff to decide what to actually use as context.
True facts. Cline does this with auto compact and it's so dumb. At the start of the session it reads the plan.md. during compaction, it forgets it. Same with its own output. It'll tell me the list of steps to take. By step 4 it needs to compact. So I prompt to do step 4. It spends 15 minutes searching the codebase and comes back that it couldn't find such a step. So dumb. So so dumb
_jayhack_ · · focus · HN ↗
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
wangii · · focus · HN ↗
f_devd · · focus · HN ↗
gfrecvh · · focus · HN ↗
wangii · · focus · HN ↗
We may spend a decade making persistent token-level KV state cheaper, before recognizing that much of what we call “context” should actually be transient, compressed, reconstructed, or represented in an entirely different state space.
KV cache risks turning an optimization of today's representation of history into tomorrow's semantics of memory.
wat10000 · · focus · HN ↗
wangii · · focus · HN ↗
[dead]
metalspot · · focus · HN ↗
trenchgun · · focus · HN ↗
Neywiny · · focus · HN ↗