I wonder how bad the performance would be if they plain ignored the whole rotary encoding dance and just back-filled precisely the parts of the cache that changed directly. Would it break the model? Confuse the model? Or would the model internally correct for it?
The issue is this: rotary encoding currently numbers the tokens on the request level.
If you pre-process a text file, you cannot insert it into the context because even if you numbered the tokens within a file properly, it is numbered wrong in the global context of the request. So you have cached data but the request structure does not allow you to insert it.
I think one way out of this would be to split a token into its request and content components. Then you might have to recalculate the request level information but never the content level information.
Edit: Now that I think about, why even cache the RoPE'd values to begin with? If you only need the RoPE'd data during token generation, then you could just RoPE on the fly instead of baking it into the KV cache, meaning you unlock block level suffix and infix caching, not just prefix caching.
Edit 2: In case caching the RoPE'd values is necessary, it might still be possible to apply the inverse RoPE for recalculation purposes.
Edit 3: By reserving a fixed block of tokens for a summary at the front right after the system prompt, you could now update that summary for the cost of prompt processing the summary without worrying that you modified something at the front.
The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture, but inversely speaking, if all you are doing is generating a summary that is strictly meant to reflect the content of the messages 1, 2, 3, then replacing them with a less information dense version that still retains the same meaning for 4, 5 means you could implement this directly in the inference engine without modifying the model at all. This would just be a special form of sliding window attention where the prefix is updated and causes partial preprocessing.
> The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture
What does "you cannot modify without changing the architecture" mean exactly? I.e. will the runtime crash if you do it wrong, or would it merely risk confusing the LLM?
Because if it's the latter, I wonder if this isn't less of a problem in practice than it sounds.
Humans don't update their memories all at the same time, either. We often end up in inconsistent mental picture that we resolve automatically when it becomes apparent (the brief "pause to think" moments).
Nothing will crash. It might create "surprise".
Example:
Trajectory: >>What is the capital of France? Let me think<<
the KV cache tokens for "let me think" will have baked in them among other things "France", "Paris", "major city", "question"
Now you compact, and reuse the later tokens (suffix)
Compacted trajectory: >>Let me think<<all right, let's see<<
Now the compacted KV cache for "Let me think" will have baked in them "France", "Paris", and the model can be "that is weird, why am I thinking about France and Paris? there is nothing related in the previous context"
So, from this follows there must be a point between "pristine cache" and "100% stale cache" where the surprise is small enough the model can recover. Perhaps with extra hint at the end, "your context cache is no fully rebuilt after a change, except confusion".
Bit like a person who just woke up, or was suddenly distracted, and now has some stray thoughts from previous context floating around. We normally dismiss them. Maybe the models can, too, and this could achieve more flexibility with cache management (therefore much lower costs of context management), at the price of slight capability reduction after context management events?
most of this stuff will be "unconscious" the model will have a pull in a particular direction without being aware
> stray thoughts from previous context floating around. We normally dismiss them
It's not that simple, see <a href="https://en.wikipedia.org/wiki/Priming_(psychology)" rel="nofollow">https://en.wikipedia.org/wiki/Priming_(psychology).
> Priming is a concept in psychology and psycholinguistics to describe how exposure to one stimulus may influence a response to a subsequent stimulus, without conscious guidance or intention
Also, we are AGI, the model isn't, it can't (re)organize its thoughts as easily as we can.
But nobody knows, what will happen is people will experiment with this, you can run the benchmarks, if it works it will be used, if not, well...
However given it's a super-obvious thing to do, drop middle tool calls from cache and keep the rest without re-prefilling, I would guess it degrades performance quite a lot, otherwise the labs would have been doing this already.
Bolwin · · focus · HN ↗
TeMPOraL · · focus · HN ↗
dist-epoch · · focus · HN ↗
But this is the kind of thing you could ask your agent to test locally.
imtringued · · focus · HN ↗
If you pre-process a text file, you cannot insert it into the context because even if you numbered the tokens within a file properly, it is numbered wrong in the global context of the request. So you have cached data but the request structure does not allow you to insert it.
I think one way out of this would be to split a token into its request and content components. Then you might have to recalculate the request level information but never the content level information.
Edit: Now that I think about, why even cache the RoPE'd values to begin with? If you only need the RoPE'd data during token generation, then you could just RoPE on the fly instead of baking it into the KV cache, meaning you unlock block level suffix and infix caching, not just prefix caching.
Edit 2: In case caching the RoPE'd values is necessary, it might still be possible to apply the inverse RoPE for recalculation purposes.
Edit 3: By reserving a fixed block of tokens for a summary at the front right after the system prompt, you could now update that summary for the cost of prompt processing the summary without worrying that you modified something at the front.
The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture, but inversely speaking, if all you are doing is generating a summary that is strictly meant to reflect the content of the messages 1, 2, 3, then replacing them with a less information dense version that still retains the same meaning for 4, 5 means you could implement this directly in the inference engine without modifying the model at all. This would just be a special form of sliding window attention where the prefix is updated and causes partial preprocessing.
TeMPOraL · · focus · HN ↗
What does "you cannot modify without changing the architecture" mean exactly? I.e. will the runtime crash if you do it wrong, or would it merely risk confusing the LLM?
Because if it's the latter, I wonder if this isn't less of a problem in practice than it sounds.
Humans don't update their memories all at the same time, either. We often end up in inconsistent mental picture that we resolve automatically when it becomes apparent (the brief "pause to think" moments).
dist-epoch · · focus · HN ↗
Example:
Trajectory: >>What is the capital of France? Let me think<<
the KV cache tokens for "let me think" will have baked in them among other things "France", "Paris", "major city", "question"
Now you compact, and reuse the later tokens (suffix)
Compacted trajectory: >>Let me think<<all right, let's see<<
Now the compacted KV cache for "Let me think" will have baked in them "France", "Paris", and the model can be "that is weird, why am I thinking about France and Paris? there is nothing related in the previous context"
TeMPOraL · · focus · HN ↗
Bit like a person who just woke up, or was suddenly distracted, and now has some stray thoughts from previous context floating around. We normally dismiss them. Maybe the models can, too, and this could achieve more flexibility with cache management (therefore much lower costs of context management), at the price of slight capability reduction after context management events?
dist-epoch · · focus · HN ↗
most of this stuff will be "unconscious" the model will have a pull in a particular direction without being aware
> stray thoughts from previous context floating around. We normally dismiss them
It's not that simple, see <a href="https://en.wikipedia.org/wiki/Priming_(psychology)" rel="nofollow">https://en.wikipedia.org/wiki/Priming_(psychology).
> Priming is a concept in psychology and psycholinguistics to describe how exposure to one stimulus may influence a response to a subsequent stimulus, without conscious guidance or intention
Also, we are AGI, the model isn't, it can't (re)organize its thoughts as easily as we can.
But nobody knows, what will happen is people will experiment with this, you can run the benchmarks, if it works it will be used, if not, well...
However given it's a super-obvious thing to do, drop middle tool calls from cache and keep the rest without re-prefilling, I would guess it degrades performance quite a lot, otherwise the labs would have been doing this already.