> We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files.
So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?
MMUs are already emulated in engines like vLLM (paged attention).
>This allows the model to learn what is most important to maintain in context
DeepSeek's Lightning Fast Indexer already does something similar, although without context compaction. It identifies which tokens are most important to attend to, which allows the model to skip irrelevant ones. A similar idea could be used to remove unnecessary tokens from the context altogether, while somehow strengthening the representation of the important ones (increasing their attention weight, merging information from discarded tokens into them, or creating compressed summary representations)
plastic-enjoyer · · focus · HN ↗
So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?
kgeist · · focus · HN ↗
MMUs are already emulated in engines like vLLM (paged attention).
>This allows the model to learn what is most important to maintain in context
DeepSeek's Lightning Fast Indexer already does something similar, although without context compaction. It identifies which tokens are most important to attend to, which allows the model to skip irrelevant ones. A similar idea could be used to remove unnecessary tokens from the context altogether, while somehow strengthening the representation of the important ones (increasing their attention weight, merging information from discarded tokens into them, or creating compressed summary representations)