I wonder how bad the performance would be if they plain ignored the whole rotary encoding dance and just back-filled precisely the parts of the cache that changed directly. Would it break the model? Confuse the model? Or would the model internally correct for it?
Kimi K3 is an example of a modern LLM that doesn't use positional encodings. It uses NoPE'd MLA for global attention and a variant of Gated DeltaNet (KDA) for local attention. I wonder how much the KDA would degrade if you just did a bounded replay over the last few thousand tokens to recover its approximate state when assembling your context window from chunks of known MLA KV.
Bolwin · · focus · HN ↗
TeMPOraL · · focus · HN ↗
wren6991 · · focus · HN ↗