‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. spankalee · · focus · HN ↗
      3.8 Flash is just quite good, and so is the Antigravity harness.

      I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it&#x27;s own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

      1. mapontosevenths · · focus · HN ↗
        Even if agy was the best (it&#x27;s not, and is missing basic features) you wouldn&#x27;t rather have a choice?

        I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

        1. drusepth · · focus · HN ↗
          What basic features are missing from agy? I&#x27;ve been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

          I actually just cancelled Ultra also because I couldn&#x27;t subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

          1. walthamstow · · focus · HN ↗
            I have used it for little more than 6 hours or so in total but I&#x27;m pretty sure it doesn&#x27;t have compaction?
            1. KeplerBoy · · focus · HN ↗
              How else would it work? Less technical people don&#x27;t even watch their context usage.
              1. macNchz · · focus · HN ↗
                In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model&#x27;s window.

                That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I&#x27;ve used has never actually produced compelling results, to where if I see I&#x27;m getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.

                1. Vacyyyy · · focus · HN ↗
                  Do you have experience with OAI&#x27;s, it&#x27;s been known to be to be good for a while now, going off public consensus and my experience.
                  1. macNchz · · focus · HN ↗
                    Yes, and I think it has improved some, but just this week 6 Astra lost the most key details of a project across a compaction and got confused about what we were actually trying to do. I would have preferred to stop at 85%, interactively develop a next-steps prompt and continue from there when ready, rather than seeing it compact and become 5x dumber from one turn to the next.
                2. marcus_holmes · · focus · HN ↗
                  I have a &quot;wrap up the session&quot; skill that I use when the session gets &gt;50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.

                  Still works better than compaction.

                  1. rnxrx · · focus · HN ↗
                    I do something similar - as I approach the context limits I have a pre-compact flush skill that extracts anything useful from the context, updates the MEMORY.md and my Obsidian vaults (set up as a poor man&#x27;s graph DB) and so forth. Once everything&#x27;s been stored I run &#x2F;compact to keep the general session flow intact. Recently I added a small embedder&#x2F;vector search setup to the same skill, which seems promising so far.

                    On another environment I&#x27;ve been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.

                  2. mikepurvis · · focus · HN ↗
                    This is what I&#x27;ve been working toward as well. It&#x27;s interesting how having the agent do its own reasoning about what it thinks is the most relevant knowledge to carry forward into the next pieces of work is vastly more effective (and even fast sometimes) than whatever the mystery-meat &quot;compaction&quot; process is.

                    Mentioned by the author in a recent HN thread, I&#x27;m also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.

                    [1]: <a href="https:&#x2F;&#x2F;ljtn.github.io&#x2F;epiq&#x2F;" rel="nofollow">https:&#x2F;&#x2F;ljtn.github.io&#x2F;epiq&#x2F;

                    1. varman11 · · focus · HN ↗
                      The idea to kill &quot;mystery-meat compaction&quot; and use an external handoff primitive is brilliant, but doesn&#x27;t letting the agent author its own handoff tickets re-introduces the same failure mode?
                      1. mikepurvis · · focus · HN ↗
                        You would think so, right? But as with others in the thread, I&#x27;d found directing the agent itself to prepare the handoff does deliver much better continuity.

                        I assume Anthropic &amp; friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.

                  3. [deleted] · · focus · HN ↗

                    [deleted]

                  4. w0m · · focus · HN ↗
                    i have active disagreements with teammates on the value of compaction&#x2F;months-long sessions.

                    Same teamates also post &#x27;Sol deleted my git repo!&#x27; or &#x27;Sorry, ignore those 300 PR comments i was just looking!&#x27; ~once a month.

                3. sroussey · · focus · HN ↗
                  No compelling results because summarization is really hard.
            2. honr · · focus · HN ↗
              It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
              1. tobias2014 · · focus · HN ↗
                That is really the biggest beef I have with agy over others, the forced auto compaction at the 250k token threshold (3.8-flash), while the model itself (via API) would be fine with a 1M context window. Even if the model is great, restricting context to 250k tokens (and auto compacting no matter what) limits certain applications and workflows somewhat.
                1. parasti · · focus · HN ↗
                  There is auto compaction at 250k? Having used agy for months with multi day sessions, I have never seen this.
                2. adastra22 · · focus · HN ↗
                  Wtf. Can you turn that off? On CC I have compaction completely turned off. I’d rather hit the hard out of context limit at 1M.
                3. sigseg1v · · focus · HN ↗
                  I find this discussion interesting. I&#x27;ve had huge increases in accuracy and huge reductions in token usage by capping my Claude models at 200k instead of 1M. I find 1M unusable and wasteful and feel that 200k should be the default. This also makes sense given that the whole reason people use &quot;Ralph loops&quot; is to keep the context window small for all tasks to get better results. Of course, clearing it yourself and manually managing it is better, but if I have 8 projects going in different terminal tabs I&#x27;m not watching any one of them that closely to effectively do that.

                  What are people using 1M context window for?

            3. piyh · · focus · HN ↗
              They only released auto mode in the last 2 weeks. Before that it was bypass permissions or manually approve every single tool call. Antigravity is permanently 6 months behind.

              I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.

              The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.

              For almost a year they didn&#x27;t allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It&#x27;s a little better now, but it&#x27;s still painfully behind the curve.

              1. throwuxiytayq · · focus · HN ↗
                Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
                1. miroljub · · focus · HN ↗
                  They are now infested with Indian style middle management making them an Infosys &#x2F; Cognizant &#x2F; Tata clone.

                  Do you know a single product from Infosys &#x2F; Cognizant &#x2F; Tata done right?

                  1. gitowiec · · focus · HN ↗
                    Yeah, that is a cancer, too much brown
                    1. miroljub · · focus · HN ↗
                      It&#x27;s not about the colour, but with the management and engineering culture.
                    2. speerer · · focus · HN ↗
                      What a disgusting and unworthy comment.
                2. p_l · · focus · HN ↗
                  Arguably gemini-cli was done in similar style, but claims on reasons for switching were about speed and efficiency of the internal jetski tool in comparison (antigravity toolkit wraps jetski code)
              2. kaszanka · · focus · HN ↗
                Is auto mode only in the IDE? I&#x27;m not seeing it in the CLI on my end, version 1.2.14.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.