Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
3.8 Flash is just quite good, and so is the Antigravity harness.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window.
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
I have a "wrap up the session" skill that I use when the session gets >50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.
I do something similar - as I approach the context limits I have a pre-compact flush skill that extracts anything useful from the context, updates the MEMORY.md and my Obsidian vaults (set up as a poor man's graph DB) and so forth. Once everything's been stored I run /compact to keep the general session flow intact. Recently I added a small embedder/vector search setup to the same skill, which seems promising so far.
On another environment I've been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.
This is what I've been working toward as well. It's interesting how having the agent do its own reasoning about what it thinks is the most relevant knowledge to carry forward into the next pieces of work is vastly more effective (and even fast sometimes) than whatever the mystery-meat "compaction" process is.
Mentioned by the author in a recent HN thread, I'm also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.
The idea to kill "mystery-meat compaction" and use an external handoff primitive is brilliant, but doesn't letting the agent author its own handoff tickets re-introduces the same failure mode?
You would think so, right? But as with others in the thread, I'd found directing the agent itself to prepare the handoff does deliver much better continuity.
I assume Anthropic & friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.
taylorfinley · · focus · HN ↗
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
spankalee · · focus · HN ↗
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
mapontosevenths · · focus · HN ↗
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
drusepth · · focus · HN ↗
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
walthamstow · · focus · HN ↗
KeplerBoy · · focus · HN ↗
macNchz · · focus · HN ↗
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
marcus_holmes · · focus · HN ↗
Still works better than compaction.
rnxrx · · focus · HN ↗
On another environment I've been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.
mikepurvis · · focus · HN ↗
Mentioned by the author in a recent HN thread, I'm also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.
[1]: <a href="https://ljtn.github.io/epiq/" rel="nofollow">https://ljtn.github.io/epiq/
varman11 · · focus · HN ↗
mikepurvis · · focus · HN ↗
I assume Anthropic & friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.
[deleted] · · focus · HN ↗
[deleted]
w0m · · focus · HN ↗
Same teamates also post 'Sol deleted my git repo!' or 'Sorry, ignore those 300 PR comments i was just looking!' ~once a month.