Can't we do this trick today with any model? Just send the file as next context. Of course you pay the price for cache misses, depending how deep you make changes, while CLM just ignores the recomputation.
One approximation of this is the experimental context management Codex has been moving towards (not released yet). Rather than relying on summary compaction, the model maintains notes as it works and as it approaches the context limit. A new session is just a fresh context with those notes attached, and a pointer back to the previous session.
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
That is similar to what I am thinking... not just edit the context as a file or string, but have a way to evict blocks and replace them with summary notes and also be able to retrieve them on demand.
No public repo but it's not too difficult to point a LLM at the general idea. Since Codex is open source, you can even take a look at how they do it. Here is what the model is instructed to do in codex (look at "guidance_message"): <a href="https://github.com/openai/codex/blob/d91294c39edb93d204926b33f21310dc968edc34/codex-rs/models-manager/models.json#L109-L112" rel="nofollow">https://github.com/openai/codex/blob/d91294c39edb93d204926b3...
Interesting. It matches my manual workflow with all harnesses (including vanilla web ChatGPT/Gemini/Claude) for the past year or so: when the session gets compacted, or (ideally) when I feel it's about to be, I just tell it to write a handover note, and start a new session.
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens.
- <a href="https://alexalejandre.com/interviews/interview-with-claude-roux/#how-do-you-leverage-llms-these-days" rel="nofollow">https://alexalejandre.com/interviews/interview-with-claude-r...
yup, I build a set of fs tools in my custom coding harness that worked like this, doing it again in another custom harness that I only expect to take one turn per request, you can still get decent caching by ordering things so most dynamic comes later
visarga · · focus · HN ↗
nsingh2 · · focus · HN ↗
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
visarga · · focus · HN ↗
Do you have a public repo for your approach?
nsingh2 · · focus · HN ↗
flyinfra · · focus · HN ↗
[dead]
ijidak · · focus · HN ↗
TeMPOraL · · focus · HN ↗
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
veqq · · focus · HN ↗
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens. - <a href="https://alexalejandre.com/interviews/interview-with-claude-roux/#how-do-you-leverage-llms-these-days" rel="nofollow">https://alexalejandre.com/interviews/interview-with-claude-r...
verdverm · · focus · HN ↗