‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. spankalee · · focus · HN ↗
      3.8 Flash is just quite good, and so is the Antigravity harness.

      I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it&#x27;s own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

      1. mapontosevenths · · focus · HN ↗
        Even if agy was the best (it&#x27;s not, and is missing basic features) you wouldn&#x27;t rather have a choice?

        I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

        1. drusepth · · focus · HN ↗
          What basic features are missing from agy? I&#x27;ve been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

          I actually just cancelled Ultra also because I couldn&#x27;t subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

          1. walthamstow · · focus · HN ↗
            I have used it for little more than 6 hours or so in total but I&#x27;m pretty sure it doesn&#x27;t have compaction?
            1. KeplerBoy · · focus · HN ↗
              How else would it work? Less technical people don&#x27;t even watch their context usage.
              1. macNchz · · focus · HN ↗
                In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model&#x27;s window.

                That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I&#x27;ve used has never actually produced compelling results, to where if I see I&#x27;m getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.

                1. marcus_holmes · · focus · HN ↗
                  I have a &quot;wrap up the session&quot; skill that I use when the session gets &gt;50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.

                  Still works better than compaction.

                  1. rnxrx · · focus · HN ↗
                    I do something similar - as I approach the context limits I have a pre-compact flush skill that extracts anything useful from the context, updates the MEMORY.md and my Obsidian vaults (set up as a poor man&#x27;s graph DB) and so forth. Once everything&#x27;s been stored I run &#x2F;compact to keep the general session flow intact. Recently I added a small embedder&#x2F;vector search setup to the same skill, which seems promising so far.

                    On another environment I&#x27;ve been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.