Context Language Models
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Context Language Models
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
svachalek · · focus · HN ↗
Bolwin · · focus · HN ↗
TeMPOraL · · focus · HN ↗
dist-epoch · · focus · HN ↗
But this is the kind of thing you could ask your agent to test locally.
imtringued · · focus · HN ↗
If you pre-process a text file you cannot insert it into the context because even if you numbered the tokens within a file properly, it is numbered wrong in the global context of the request. So you have cached data but the request structure does not allow you to insert it.
TeMPOraL · · focus · HN ↗
What does "you cannot modify without changing the architecture" mean exactly? I.e. will the runtime crash if you do it wrong, or would it merely risk confusing the LLM?
Because if it's the latter, I wonder if this isn't less of a problem in practice than it sounds.
Humans don't update their memories all at the same time, either. We often end up in inconsistent mental picture that we resolve automatically when it becomes apparent (the brief "pause to think" moments).
dist-epoch · · focus · HN ↗
Example:
Trajectory: >>What is the capital of France? Let me think<<
the KV cache tokens for "let me think" will have baked in them among other things "France", "Paris", "major city", "question"
Now you compact, and reuse the later tokens (suffix)
Compacted trajectory: >>Let me think, all right, let's see<<
Now the compacted KV cache for "Let me think" will have baked in them "France", "Paris", and the model can be "that is weird, why am I thinking about France and Paris? there is nothing related in the previous context"
TeMPOraL · · focus · HN ↗
Bit like a person who just woke up, or was suddenly distracted, and now has some stray thoughts from previous context floating around. We normally dismiss them. Maybe the models can, too, and this could achieve more flexibility with cache management (therefore much lower costs of context management), at the price of slight capability reduction after context management events?
dist-epoch · · focus · HN ↗
most of this stuff will be "unconscious" the model will have a pull in a particular direction.
> stray thoughts from previous context floating around. We normally dismiss them
It's not that simple, see <a href="https://en.wikipedia.org/wiki/Priming_(psychology)" rel="nofollow">https://en.wikipedia.org/wiki/Priming_(psychology)
> Priming is a concept in psychology and psycholinguistics to describe how exposure to one stimulus may influence a response to a subsequent stimulus, without conscious guidance or intention
But nobody knows, what will happen is people will experiment with this, you can run the benchmarks, if it works it will be used, if not, well...
imtringued · · focus · HN ↗
wren6991 · · focus · HN ↗
visarga · · focus · HN ↗
nsingh2 · · focus · HN ↗
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
visarga · · focus · HN ↗
Do you have a public repo for your approach?
nsingh2 · · focus · HN ↗
flyinfra · · focus · HN ↗
[dead]
ijidak · · focus · HN ↗
TeMPOraL · · focus · HN ↗
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
veqq · · focus · HN ↗
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens. - <a href="https://alexalejandre.com/interviews/interview-with-claude-roux/" rel="nofollow">https://alexalejandre.com/interviews/interview-with-claude-r...
verdverm · · focus · HN ↗
bob1029 · · focus · HN ↗
Do you want your agent solving its own memory crisis, or do you want it solving the actual task? It can probably do both at the same time, but I suspect there is a non trivial cost associated with this.
A separate hypervisor agent that manages the main agent's context would be much better in my experience. You can run it on a different schedule and the main agent has to spend zero tokens thinking about it. This also makes it a lot easier to control when caches will be missed.
alightsoul · · focus · HN ↗
theroadnotbacon · · focus · HN ↗
alightsoul · · focus · HN ↗
vatsachak · · focus · HN ↗
verdverm · · focus · HN ↗
naw103 · · focus · HN ↗
[dead]
rajveerb · · focus · HN ↗
A separate hypervisor agent at least doubles the cost because (1) the underlying model needs to be the same so that we get the same degree of intelligence, and (2) it needs to have the same context + more tokens for the work it does.
flyinfra · · focus · HN ↗
[dead]
gitghxst · · focus · HN ↗
aghuang · · focus · HN ↗
plastic-enjoyer · · focus · HN ↗
So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?
actionfromafar · · focus · HN ↗
kgeist · · focus · HN ↗
MMUs are already emulated in engines like vLLM (paged attention).
gavinray · · focus · HN ↗
royal__ · · focus · HN ↗
vatsachak · · focus · HN ↗
And there will be multiple contexts like hot vs cold pages in DBs
jkhdigital · · focus · HN ↗
killerstorm · · focus · HN ↗
bob1029 · · focus · HN ↗
_jayhack_ · · focus · HN ↗
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": <a href="https://www.completeskeptic.com/p/kv-cache-rules-everything-around" rel="nofollow">https://www.completeskeptic.com/p/kv-cache-rules-everything-...
wangii · · focus · HN ↗
f_devd · · focus · HN ↗
gfrecvh · · focus · HN ↗
wangii · · focus · HN ↗
We may spend a decade making persistent token-level KV state cheaper, before recognizing that much of what we call “context” should actually be transient, compressed, reconstructed, or represented in an entirely different state space.
KV cache risks turning an optimization of today's representation of history into tomorrow's semantics of memory.
wat10000 · · focus · HN ↗
wangii · · focus · HN ↗
[dead]
metalspot · · focus · HN ↗
trenchgun · · focus · HN ↗
Neywiny · · focus · HN ↗
hashmap · · focus · HN ↗
sayamss · · focus · HN ↗
rajveerb · · focus · HN ↗
imtringued · · focus · HN ↗
No, we have DEQ at home.
DEQ at home: We just put the context into a file first and then prompt the model to edit it.
thelocalnode · · focus · HN ↗
[dead]
darkoob12 · · focus · HN ↗
It is pathetic