‹ BackHN Continuity

Thread

Context Language Models

176 points · 51 comments · emersonmacro

  1. Bolwin · · focus · HN ↗
    The biggest discovery might actually be that they ignored regular caching rules and kept invalid cache suffixes and it didn't hurt performance
    1. TeMPOraL · · focus · HN ↗
      I wonder how bad the performance would be if they plain ignored the whole rotary encoding dance and just back-filled precisely the parts of the cache that changed directly. Would it break the model? Confuse the model? Or would the model internally correct for it?
      1. wren6991 · · focus · HN ↗
        Kimi K3 is an example of a modern LLM that doesn't use positional encodings. It uses NoPE'd MLA for global attention and a variant of Gated DeltaNet (KDA) for local attention. I wonder how much the KDA would degrade if you just did a bounded replay over the last few thousand tokens to recover its approximate state when assembling your context window from chunks of known MLA KV.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.