‹ BackHN Continuity

Thread

Context Language Models

176 points · 51 comments · emersonmacro

  1. plastic-enjoyer · · focus · HN ↗
    > We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files.

    So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?

    1. actionfromafar · · focus · HN ↗
      yes, and "kubernetes"
    2. kgeist · · focus · HN ↗
      >Do we have to reinvent MMUs for LLMs

      MMUs are already emulated in engines like vLLM (paged attention).

      >This allows the model to learn what is most important to maintain in context

      DeepSeek's Lightning Fast Indexer already does something similar, although without context compaction. It identifies which tokens are most important to attend to, which allows the model to skip irrelevant ones. A similar idea could be used to remove unnecessary tokens from the context altogether, while somehow strengthening the representation of the important ones (increasing their attention weight, merging information from discarded tokens into them, or creating compressed summary representations)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.