‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. spankalee · · focus · HN ↗
      3.8 Flash is just quite good, and so is the Antigravity harness.

      I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it&#x27;s own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

      1. mapontosevenths · · focus · HN ↗
        Even if agy was the best (it&#x27;s not, and is missing basic features) you wouldn&#x27;t rather have a choice?

        I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

        1. drusepth · · focus · HN ↗
          What basic features are missing from agy? I&#x27;ve been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

          I actually just cancelled Ultra also because I couldn&#x27;t subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

          1. esafak · · focus · HN ↗
            I use a variety of models for various subagents. I don&#x27;t want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
          2. walthamstow · · focus · HN ↗
            I have used it for little more than 6 hours or so in total but I&#x27;m pretty sure it doesn&#x27;t have compaction?
            1. KeplerBoy · · focus · HN ↗
              How else would it work? Less technical people don&#x27;t even watch their context usage.
              1. macNchz · · focus · HN ↗
                In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model&#x27;s window.

                That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I&#x27;ve used has never actually produced compelling results, to where if I see I&#x27;m getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.

                1. Vacyyyy · · focus · HN ↗
                  Do you have experience with OAI&#x27;s, it&#x27;s been known to be to be good for a while now, going off public consensus and my experience.
                  1. macNchz · · focus · HN ↗
                    Yes, and I think it has improved some, but just this week 6 Astra lost the most key details of a project across a compaction and got confused about what we were actually trying to do. I would have preferred to stop at 85%, interactively develop a next-steps prompt and continue from there when ready, rather than seeing it compact and become 5x dumber from one turn to the next.
                2. marcus_holmes · · focus · HN ↗
                  I have a &quot;wrap up the session&quot; skill that I use when the session gets &gt;50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.

                  Still works better than compaction.

                  1. rnxrx · · focus · HN ↗
                    I do something similar - as I approach the context limits I have a pre-compact flush skill that extracts anything useful from the context, updates the MEMORY.md and my Obsidian vaults (set up as a poor man&#x27;s graph DB) and so forth. Once everything&#x27;s been stored I run &#x2F;compact to keep the general session flow intact. Recently I added a small embedder&#x2F;vector search setup to the same skill, which seems promising so far.

                    On another environment I&#x27;ve been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.

                  2. mikepurvis · · focus · HN ↗
                    This is what I&#x27;ve been working toward as well. It&#x27;s interesting how having the agent do its own reasoning about what it thinks is the most relevant knowledge to carry forward into the next pieces of work is vastly more effective (and even fast sometimes) than whatever the mystery-meat &quot;compaction&quot; process is.

                    Mentioned by the author in a recent HN thread, I&#x27;m also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.

                    [1]: <a href="https:&#x2F;&#x2F;ljtn.github.io&#x2F;epiq&#x2F;" rel="nofollow">https:&#x2F;&#x2F;ljtn.github.io&#x2F;epiq&#x2F;

                    1. varman11 · · focus · HN ↗
                      The idea to kill &quot;mystery-meat compaction&quot; and use an external handoff primitive is brilliant, but doesn&#x27;t letting the agent author its own handoff tickets re-introduces the same failure mode?
                      1. mikepurvis · · focus · HN ↗
                        You would think so, right? But as with others in the thread, I&#x27;d found directing the agent itself to prepare the handoff does deliver much better continuity.

                        I assume Anthropic &amp; friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.

                  3. [deleted] · · focus · HN ↗

                    [deleted]

                  4. w0m · · focus · HN ↗
                    i have active disagreements with teammates on the value of compaction&#x2F;months-long sessions.

                    Same teamates also post &#x27;Sol deleted my git repo!&#x27; or &#x27;Sorry, ignore those 300 PR comments i was just looking!&#x27; ~once a month.

                3. sroussey · · focus · HN ↗
                  No compelling results because summarization is really hard.
            2. honr · · focus · HN ↗
              It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
              1. tobias2014 · · focus · HN ↗
                That is really the biggest beef I have with agy over others, the forced auto compaction at the 250k token threshold (3.8-flash), while the model itself (via API) would be fine with a 1M context window. Even if the model is great, restricting context to 250k tokens (and auto compacting no matter what) limits certain applications and workflows somewhat.
                1. parasti · · focus · HN ↗
                  There is auto compaction at 250k? Having used agy for months with multi day sessions, I have never seen this.
                2. adastra22 · · focus · HN ↗
                  Wtf. Can you turn that off? On CC I have compaction completely turned off. I’d rather hit the hard out of context limit at 1M.
                3. sigseg1v · · focus · HN ↗
                  I find this discussion interesting. I&#x27;ve had huge increases in accuracy and huge reductions in token usage by capping my Claude models at 200k instead of 1M. I find 1M unusable and wasteful and feel that 200k should be the default. This also makes sense given that the whole reason people use &quot;Ralph loops&quot; is to keep the context window small for all tasks to get better results. Of course, clearing it yourself and manually managing it is better, but if I have 8 projects going in different terminal tabs I&#x27;m not watching any one of them that closely to effectively do that.

                  What are people using 1M context window for?

            3. piyh · · focus · HN ↗
              They only released auto mode in the last 2 weeks. Before that it was bypass permissions or manually approve every single tool call. Antigravity is permanently 6 months behind.

              I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.

              The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.

              For almost a year they didn&#x27;t allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It&#x27;s a little better now, but it&#x27;s still painfully behind the curve.

              1. throwuxiytayq · · focus · HN ↗
                Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
                1. miroljub · · focus · HN ↗
                  They are now infested with Indian style middle management making them an Infosys &#x2F; Cognizant &#x2F; Tata clone.

                  Do you know a single product from Infosys &#x2F; Cognizant &#x2F; Tata done right?

                  1. gitowiec · · focus · HN ↗
                    Yeah, that is a cancer, too much brown
                    1. miroljub · · focus · HN ↗
                      It&#x27;s not about the colour, but with the management and engineering culture.
                    2. speerer · · focus · HN ↗
                      What a disgusting and unworthy comment.
                2. p_l · · focus · HN ↗
                  Arguably gemini-cli was done in similar style, but claims on reasons for switching were about speed and efficiency of the internal jetski tool in comparison (antigravity toolkit wraps jetski code)
              2. kaszanka · · focus · HN ↗
                Is auto mode only in the IDE? I&#x27;m not seeing it in the CLI on my end, version 1.2.14.
          3. arizen · · focus · HN ↗
            Does it have &#x2F;goal feature similar to Codex?
            1. sorrybutidontha · · focus · HN ↗
              yes
          4. sarjann · · focus · HN ↗
            Auto mode?
            1. KeplerBoy · · focus · HN ↗
              It absolutely has auto mode.
              1. levelZero · · focus · HN ↗
                Via cli switch, but in process w&#x2F;o fine graining? If so please tell
              2. SomaticPirate · · focus · HN ↗
                [delayed]
              3. smartbit · · focus · HN ↗
                agy cli does not have auto mode. I&#x27;ve tried and tried and tried to work with agy cli sandbox-mode and just failed.

                  agy --dangerously-skip-permissions
                
                in my experience is the only workable solution that doesn&#x27;t ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don&#x27;t want to work with the GUI and b) it was poorly implemented as I couldn&#x27;t get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.

                gemini-cli supported &#x27;pre-write diff tabs&#x27; (y&#x2F;n) in external editors like vscode. In Claude Code I heavily use &#x27;pre-write diff tabs&#x27; for documentation and miss it sincerely in agy cli.

                IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.

                1. Kostchei · · focus · HN ↗
                  PSA, in the agy ui there is a button. It was very annoying until i set it. Slight downside- if I ask it to write a planning doc it will write the doc and then implement it without asking. but as long as you know that, no problem....
                  1. smartbit · · focus · HN ↗
                    do you mean?

                      accept-edits
                    
                    also available with shift-tab [0]. That is not related to executing commands, only to allowing agy edit files. Unless you set Turbo-mode == yolo-mode, agy prompts a zillion times.

                    I&#x27;m referencing this but can&#x27;t see change in daily work: v2.14.0 (September 15, 2026) &quot;New Permissions System&quot; [1]

                    Details: Introduced the new unified permissions system, presets (Default, Request Review, Turbo), syntax-highlighted permission requests, and restructured the settings under Global Permissions and project-level Inherit Global.

                    [0] <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;cli&#x2F;modes&#x2F;#available-modes" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;cli&#x2F;modes&#x2F;#available-modes [1] <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;changelog&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;changelog&#x2F;

                2. Phineas_here · · focus · HN ↗
                  have you had any experiences where the agent just made unintended edits to the code as you tried to make it run auto? cuz I&#x27;m always skeptical about letting it go auto but there&#x27;s not much I can do when it gets repetitive
              4. LoganDark · · focus · HN ↗
                Auto mode means that another model reviews tool calls to ensure they&#x27;re safe before they&#x27;re allowed. It&#x27;s different from bypass permissions mode which typically just doesn&#x27;t filter at all.
            2. augusto-moura · · focus · HN ↗
              No auto mode is the thing that bothers me, I don&#x27;t trust the cli blindly nor do I trust myself to read every python script it throws at me. Auto mode is an acceptable middle ground in my experience
              1. Phineas_here · · focus · HN ↗
                you could use an ai governance agent if you don&#x27;t want to manually review every script. and if you already use any which ones do you think are the most recommendable?
                1. KshitizLoharuka · · focus · HN ↗

                  [dead]

                2. KshitizLoharuka · · focus · HN ↗

                  [dead]

              2. tiborsaas · · focus · HN ↗
                Shift+TAB sets accept-edits and plan mode.
          5. dleslie · · focus · HN ↗
            Emacs integration over ACP.

            They&#x27;ve got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM

            1. p_l · · focus · HN ↗
              agent-shell works with antigravity but I haven&#x27;t tested it much yet
              1. dleslie · · focus · HN ↗
                It works but it&#x27;s a violation of the TOS to use it.

                I would rather not risk my Google account.

                1. p_l · · focus · HN ↗
                  It&#x27;s not - it uses the same interface as editor extensions like the one for VScode.

                  It does mean however that it cannot operate as flexibly as it it could with raw API, IMO, but agent-shell is essentially designed towards wrapping the official clients

                  1. dleslie · · focus · HN ↗
                    They removed the ACP command line flag that was in the old Gemini client. That&#x27;s a signal that they won&#x27;t support it.

                    While it may be technically allowed, I&#x27;m not about to risk my account. Google has proven themselves to be capricious and arbitrary when it comes to TOS enforcement, and their appeals system doesn&#x27;t meaningfully exist in practice.

                    1. p_l · · focus · HN ↗
                      I am not talking about ACP command flag - there is a separate binary that exists for integration into editors, and it&#x27;s how VScode extension works. So it&#x27;s the same usage pattern underneath
                      1. dleslie · · focus · HN ↗
                        I don&#x27;t see a download link for that:

                        <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;

                        What I see is integrations for specific editors.

                        1. p_l · · focus · HN ↗
                          Agent-shell uses ACP server published through ACP registry here [1] - those are AFAIK used by editor extensions as backend since VScode can&#x27;t exactly run Go-based extensions - Zed instructions explicitly talk about using ACP Registry too [2]

                          [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;agentclientprotocol&#x2F;registry&#x2F;blob&#x2F;main&#x2F;antigravity-acp%2Fagent.json" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;agentclientprotocol&#x2F;registry&#x2F;blob&#x2F;main&#x2F;an...

                          [2] <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;zed" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;zed

                          1. dleslie · · focus · HN ↗
                            They should really make this clear on their website.

                            But thank you, it appears there is an official and generic ACP client.

          6. mapontosevenths · · focus · HN ↗
            Sorry for the delay, I didn&#x27;t want to drop a glib half answer on you. Using agy is like going back in time. It&#x27;s better than Gemini CLI was, but that&#x27;s a really low bar.

            I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.

            1) Try to integrate agy into a workflow. It can&#x27;t do standard I&#x2F;O like: tail -200 app.log | claude -p &quot;Find the problem&quot;

            2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.

            3) Not open source so I can&#x27;t fix any of these problems.

            4) No skills. In 2026. Yikes.

            5) No persistent memory (see Claudes auto memory)

            6) No sub-agents or orchestration of any type really.

            7) Weird hard coded limits and constant API errors on everything (scaling problems?)

            8) No &#x2F;loop command

            9) &#x2F;btw is weird and ephemeral. No way to merge it back to the conversation.

            10) Unstable in general.

            11) No way to control it via API.

            I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it&#x27;s so worth it. Then if you want try to go back to agy. Don&#x27;t worry, agy won&#x27;t have changed much. It improves at a snails pace.

            1. trevorm4 · · focus · HN ↗
              It has both skills (<a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;skills&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;skills&#x2F;) and agents (<a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;subagents&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;subagents&#x2F;)
            2. thanhhaimai · · focus · HN ↗
              Opinions are my own.

              I&#x27;m not sure this list is correct. Number 4 is especially wrong, since Skills are available with the launch of Antigravity 2:

              <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;blog&#x2F;introducing-google-antigravity-2" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;blog&#x2F;introducing-google-antigravi...

            3. anyg · · focus · HN ↗
              Isn&#x27;t &#x2F;btw meant to be ephemeral?
              1. p_l · · focus · HN ↗
                and has explicit &quot;&#x2F;copy btw&quot; if you want to save it
            4. drusepth · · focus · HN ↗
              I don&#x27;t know when you last tried agy, but if you ever go back to try it again, you&#x27;ll hopefully be happy to know it does indeed support skills, sub-agents with pretty good inter-agent communication, and probably more.
          7. krisgenre · · focus · HN ↗
            &gt;I couldn&#x27;t subscribe to a YouTube Family plan while I had it active

            Luckily for me both expired yesterday and I was able to subscribe back again (first Youtube family and then Google AI plan).

          8. edg5000 · · focus · HN ↗
            Why use Codex CLI if you can use the ChatGPT app (which is a Codex GUI in all but name). Personally I weirdly got used to the terrible TUI stuff.
            1. krzyk · · focus · HN ↗
              Why use some GUI when there is a TUI?
              1. esafak · · focus · HN ↗
                Because &#x27;graphical&#x27; TUIs are pale imitations of GUIs. I don&#x27;t have a TUI fetish, despite having grown up with them.
                1. krzyk · · focus · HN ↗
                  It is not a fetish, some people like this and others like that.

                  I like CLI more than TUI, and TUI more than GUI where appropriate. For working with text TUI is better, for e.g. images GUI (GIMP).

                2. crossroadsguy · · focus · HN ↗
                  [delayed]
                  1. nsonha · · focus · HN ↗
                    You&#x27;re a software engineer living in a middle of an AI revolution and you confuse what IS with what CAN BE? GUI can be good people just didn&#x27;t do them because it took time and the foundation was shit (for the options that didn&#x27;t take time). That all changes now when a new class of software can be generated with AI. Unfortunately the people generating software and their users still base the decision on what IS (with highly technical reasoning such as &quot;infinitely better&quot;). People literally have the power to define what IS these days. I&#x27;ve only recently switched from claude and codex cli to their respective apps and those apps while not the best gui apps, already infinitely better than tui. The tui is actually hurting my main usage of them which is controlling agent programmatically. The only really point I&#x27;d give for TUI is compatibility. Running coding agent directly on android&#x2F;ios is nice sometimes.
                    1. antonvs · · focus · HN ↗
                      &gt; That all changes now when a new class of software can be generated with AI.

                      All the AI-generated UIs I&#x27;ve seen have been very derivative, certainly not eliminating any of the disadvantages of typical GUI interfaces.

                      Realistically, the whole &quot;overlapping windows&quot; GUI model, and everything that derives from that, was a metaphor geared towards people who&#x27;d never seen a computer before. It fit the increasing consumer focus of computing interfaces. It&#x27;s no wonder that technical people often prefer TUIs.

                      Maybe AI will bring real advancements in GUIs, but someone&#x27;s still going to have to make it happen.

                      1. nsonha · · focus · HN ↗
                        GUI is not just your &quot;overlapping windows&quot; straw man, it&#x27;s interactivity, plus graphical display. Yes via a taxonomy hack TUI is not GUI and there are technical hacks that bring graphics to TUIs these days, but the point remains that we need complex (but not overlapping windows, sure!) interactivity and rich display capability.

                        That is what we need, and if you&#x27;re making the argument that a terminal shell is the best place to provide them then I don&#x27;t know what to say.

                        1. antonvs · · focus · HN ↗
                          I&#x27;m pointing out that current GUIs are pretty primitive and haven&#x27;t undergone much serious thought about functional improvement since Xerox PARC in the late 1970s, and that&#x27;s why TUIs can still have an edge with technically-inclined people.

                          It&#x27;s not that GUIs are inherently worse in principle, but in practice they often are.

                          The point about overlapping windows is that that &quot;desktop&quot; model permeates the thinking about GUI design, but it&#x27;s fundamentally limiting and misguided.

                          &gt; we need complex (but not overlapping windows, sure!) interactivity and rich display capability.

                          Yep. Pity today&#x27;s GUIs can&#x27;t deliver that.

                    2. crossroadsguy · · focus · HN ↗
                      [delayed]
                      1. nsonha · · focus · HN ↗
                        This is not about vision or anything like that it&#x27;s about thinking like actual engineers and understand that what accidentally IS will always be inferior to what is designated to (can) be. TUIs will always be a hack and good by accident.
                3. agentcoops · · focus · HN ↗
                  “ 39. Re graphics: A picture is worth 10K words - but only those to describe the picture. Hardly any sets of 10K words can be adequately described with pictures.”

                  Perlis has an aphorism for this, as he does every important problem [0].

                  [0] <a href="https:&#x2F;&#x2F;www.cs.yale.edu&#x2F;homes&#x2F;perlis-alan&#x2F;quotes.html" rel="nofollow">https:&#x2F;&#x2F;www.cs.yale.edu&#x2F;homes&#x2F;perlis-alan&#x2F;quotes.html

          9. nxdmum · · focus · HN ↗
            IF you havent written your own harness - you would not understand what you can do when you&#x27;re writing your own harness . the current set of harnesses - all of them are crap tier. The only clue i can give you - it&#x27;s not in the model providers interest to have token efficiency - but when you are coding the harness yourself you can shoot for that .

            In today&#x27;s world - and idea stated stated is an idea stolen .

            1. cowl · · focus · HN ↗
              the only clue you can give... why? because it&#x27;s just words? plenty of opensource harnesses there that invalidate the conspiracy in your only clue that you can give.
            2. dmos62 · · focus · HN ↗
              What you&#x27;re saying is sus. If you have a harness that&#x27;s a tier above frontier labs&#x27; offerings, link to it.
              1. imtringued · · focus · HN ↗
                It&#x27;s not sus at all. You&#x27;re looking at it from the perspective of a general purpose system that is not adapted to your use case. Basically you prompt the model and then let the agent do everything. That&#x27;s the use case you have in mind when you think that it&#x27;s about being &quot;a tier above frontier labs&#x27; offerings&quot;.

                It&#x27;s missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.

              2. adastra22 · · focus · HN ↗
                Just about every harness is. This is common knowledge, no? Frontier lab TUI tend to steal from the OSS harnesses not the other way around.
                1. dmos62 · · focus · HN ↗
                  Care to back that up with a benchmark? I&#x27;ve not seen that. OpenCode is on here, not exactly dominating though: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents
                  1. adastra22 · · focus · HN ↗
                    Benchmarks are a terrible judge for this.
                    1. dmos62 · · focus · HN ↗
                      Why?
                      1. adastra22 · · focus · HN ↗
                        The process by which benchmarks are setup and run does not correspond at all to how human developers engage with a coding agent. At best it is a loose proxy, and often a bad one.

                        What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.

                        1. dmos62 · · focus · HN ↗
                          I presume you&#x27;re not talking about autonomous agents, right? Because then you could give it a set of tasks and check its success rate. But, even so, are you not giving it tasks? Are you principally interacting with it through discussions that are more difficult to quantify? Even question answering has benchmarks. I&#x27;m having trouble imagining how you use them (or &quot;how people use them&quot; in your words).
                          1. adastra22 · · focus · HN ↗
                            I don&#x27;t know if you noticed, but Opus 4.6 was peak for human-computer interaction. Everything has been fairly downhill from there despite better benchmarks, at least in that one regard. Opus 5 and 5.5 are clearly a step above in capabilities than 4.6, and I don&#x27;t think anyone wants to go back, but 4.7 and 4.8 were arguably worse overall. I genuinely feel I got more done with 4.6 and often switched back, prior to 5 coming out.

                            Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.

                            Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock&#x2F;swarm situation. The benchmark just doesn&#x27;t cover this. (And the difference can be nontrivial! Sakana AI&#x27;s published results show two generations of uplifting potential from better harnesses.)

            3. crossroadsguy · · focus · HN ↗
              [delayed]
              1. adastra22 · · focus · HN ↗
                Man good luck with finding that. People who wrote their own harness tend to self select into the type of people that DON’T self-author blog posts.
            4. tikkosam · · focus · HN ↗
              I think it is very much in the first party providers&#x27; interest to chase token efficiency, considering that they are offering fixed price monthly plans, and users may go to a competitor when they hit limits.
            5. edg5000 · · focus · HN ↗
              Actually the popular harnesses achieve high cache rates, but probably the syntax and prompts are more verbose than they need to be. A simpler syntax is doable, but I found with more exotic architectures I end up with lower cache rates.
          10. pjc50 · · focus · HN ↗
            &gt; I actually just cancelled Ultra also because I couldn&#x27;t subscribe to a YouTube Family plan while I had it active (Google... :[)

            This sort of thing alarms me. Having a $bigcorp account becomes a &quot;&quot;social credit&quot;&quot; system where they can ban you from all your personal stuff if they decide that you (or your agents!) are doing stuff they don&#x27;t like.

            1. w0m · · focus · HN ↗
              I take that more as the user is mixing business with pleasure. When i joined a company using `gcp` heavily, I didn&#x27;t attach my personal gmail to it - I created a &#x27;business&#x27; account and used that. Slightly inconvenient - agree, but it&#x27;s on the individual to draw the distinctions.
              1. runamok · · focus · HN ↗
                There is still anecdata that Google &quot;knows&quot; you are the same person and will ban all accounts associated with you if you do something they don&#x27;t like. And of course they never need to explain themselves or listen to your appeal.
            2. nout · · focus · HN ↗
              And that&#x27;s why it&#x27;s very powerful to be able to run AI locally, even if less smart.
        2. eloisant · · focus · HN ↗
          There is a pi plugin to use agy directly from it.
          1. mapontosevenths · · focus · HN ↗
            You get banned if they catch you.
            1. tcoff91 · · focus · HN ↗
              Yes and I&#x27;ve seen reports of it being an ENTIRE GOOGLE ACCOUNT BAN.

              I don&#x27;t want to mess with antigravity because my google account is too entrenched in my life.

              1. shmoogy · · focus · HN ↗
                That&#x27;s why it&#x27;s a nonstarter for me.
              2. 8note · · focus · HN ↗
                which makes basically any product to build with google a nonstarter.

                without having an entirely separate google account with its own separated bans, theres just no ability to trust those

                1. Sabinus · · focus · HN ↗
                  I&#x27;ve read here in previous years about bans propagating to other accounts. I think that if Google can associate you with other accounts that they&#x27;re not above banning those too.
              3. jsw97 · · focus · HN ↗
                I was getting excited but thanks for reminding me of this. Not messing with this.
              4. rdtsc · · focus · HN ↗
                Yup. I don’t plan to do anything sneaky but one wrong question or query that looks like “cyber”, say me fixing a buffer overflow in library I maintain, and all of the sudden my gmail is blocked. Yeah, not worth the risk. I feel like even with a different account they’ll figure out it’s me because well, as ad sellers that’s their business to find out who is who and I will still be banned.
        3. spankalee · · focus · HN ↗
          There&#x27;s an API: you can use Gemini with other harnesses. Isn&#x27;t the situation exactly like Claude vs Claude Code?
          1. mapontosevenths · · focus · HN ↗
            You can&#x27;t without mortgaging your home to pay enterprise API rate pricing. It&#x27;s prevented on the plans, and if you find away around it they don&#x27;t ban you from Gemini... They ban your whole-ass Google account forever.
        4. gchamonlive · · focus · HN ↗
          I&#x27;m cowboying Gemini on oh-my-pi. Been running OK so far, hopefully I won&#x27;t get banned, and if so hopefully I&#x27;ll only lose access to the models, not the storage -- while models are a sort of commodity, my data isn&#x27;t.
          1. ForHackernews · · focus · HN ↗
            Careful, if they ban your Google account, they might ban you from everything: gmail, google drive, adwords, app engine, youtube, voice, android...
            1. gchamonlive · · focus · HN ↗
              [delayed]
              1. ajolly · · focus · HN ↗
                No, Google family now shares limits across all your accounts.
                1. gchamonlive · · focus · HN ↗
                  [delayed]
      2. moecables · · focus · HN ↗
        I use Antigravity but for some reason, `agy` in the command line feels very bad&#x2F;incapable of doing things. I can&#x27;t quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call
        1. Conscat · · focus · HN ↗
          I remember having this issue ALL THE TIME with Gemini CLI but personally I haven&#x27;t experienced that yet with agy.
      3. starfallg · · focus · HN ↗
        I found 3.8 Flash in Agy to be generally better than GPT 6, and only behind from Opus 5.5.
      4. lp92 · · focus · HN ↗
        Same! 3.8 Flash does really well for writing code as long as you give it a good design and plan to follow. I use Opus for architecture&#x2F;design&#x2F;implementation plans and let Gemini 3.8 work using those. Even on the $20 Pro plan I&#x27;ve only come down to 10% before the weekly reset.
      5. mpweiher · · focus · HN ↗
        I tried antigravity a little while ago and it was utterly useless for Objective-C code, tasks that both Claude and Codex handled just fine.

        Not only could it not complete the small task, the code was obviously wrong from looking at it and did not even compile.

        When I pointed that out it got pissy and insisted the code was perfect and I didn&#x27;t know how to use a compiler, or the compiler was buggy. Pasting the compiler errors did not help.

        Surreal.

      6. moffkalast · · focus · HN ↗
        &gt; Antigravity

        Is that a reference to <a href="https:&#x2F;&#x2F;xkcd.com&#x2F;353&#x2F;" rel="nofollow">https:&#x2F;&#x2F;xkcd.com&#x2F;353&#x2F;

      7. gchamonlive · · focus · HN ↗
        Your comment made me try agy again, but after having to approve and persist every single read tool, I got fatigued really fast. Oh-my-pi shows that these checks aren&#x27;t really necessary for harness safety, it&#x27;s best to invest in harness predictability.
      8. jmaker · · focus · HN ↗
        I couldn’t find a way to decline model training and reduce data retention for antigravity or Gemini. Apparently it’s only available on a business&#x2F;enterprise plan, not personal. Did you manage to solve it? That’s the only reason I don’t use Gemini or antigravity.
      9. pdntspa · · focus · HN ↗
        The gemini series have also been really strong on text extraction. I&#x27;ve been evaluating models to replace gemini 2.5 and its been hard to find something that performs as well as other gemini models
      10. noahmichael89 · · focus · HN ↗
        For frontend work, the model&#x2F;harness&#x2F;tool combo of 3.8 flash extended&#x2F;agy&#x2F;chrome dev tools MCP is shockingly good
    2. IndeanCondor · · focus · HN ↗
      Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
      1. seanthemon · · focus · HN ↗
        Gemini for day-to-day and top-of-head queries and claude for the real beefy work
    3. mapontosevenths · · focus · HN ↗
      Gemini is honestly amazing sometimes. If they didn&#x27;t force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I&#x27;m sure Google could take over the AI market.
      1. ody4242 · · focus · HN ↗
        what is so terrible with their harness? I&#x27;ve been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok.
      2. barrenko · · focus · HN ↗
        Google&#x27;s approach reminds me of what could have been the EU&#x27;s approach, they are really reluctant and drag they feet, but in the end they end up shipping and are competitive.
    4. alightsoul · · focus · HN ↗
      Please tell me you published your findings even as an issue on the llama.cpp GitHub
      1. warkdarrior · · focus · HN ↗
        Why? Anyone can run that prompt.
        1. aspect0545 · · focus · HN ↗
          Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it.
          1. FranzFerdiNaN · · focus · HN ↗
            The resources per prompt aren’t that much .

            Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.

            1. hexfish · · focus · HN ↗
              Checkmate. &#x2F;s
            2. qmr · · focus · HN ↗
              Yet you participate in a society.
              1. scarmig · · focus · HN ↗
                If action X takes a million times more resources than action Y, it&#x27;s silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn&#x27;t somehow erase that or make it irrelevant.
                1. qmr · · focus · HN ↗
                  It gets me 10 internet points though.
            3. articulatepang · · focus · HN ↗
              All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them.

              Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.

          2. dzhiurgis · · focus · HN ↗
            Took me 6M codex tokens yesterday to get omarchy wifi working on macbook
            1. alightsoul · · focus · HN ↗
              That is not normal. Were you able to use the arch wiki? Omarchy is a version of arch Linux.
              1. dzhiurgis · · focus · HN ↗
                IDK how it fixed it exactly, I suspect it&#x27;s macbook specific issue. It wasn&#x27;t accepting my wifi password and 5ghz radio wasn&#x27;t working. Steering it to fix 5ghz sorted it out. I&#x27;m not going to any wiki&#x27;s myself, fuck that.
          3. AuthAuth · · focus · HN ↗
            They dont understand the fix so sharing a vibe coded patch is upstream spam. They should submit a bug report and their findings. (Written by them not AI)
        2. luckydata · · focus · HN ↗
          why reinvent the wheel and spend tokens for a problem that has already been solved?
        3. baby_souffle · · focus · HN ↗
          Wouldn&#x27;t it be better if only one person had to and then we all got to benefit from the fix?

          Why would the guy who wrote curl share it? We can all build our own now...

          Why do the Linux folks need to be so selfless? We can all build our own kernel now...

        4. folkrav · · focus · HN ↗
          Why would we even ever distribute software again, by this logic?
      2. otabdeveloper4 · · focus · HN ↗
        Spoiler alert: the problem didn&#x27;t actually get fixed despite the jaw on the floor.
      3. dominotw · · focus · HN ↗
        he is still closing his jaw
      4. taylorfinley · · focus · HN ↗
        Here they are: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...
      5. imtringued · · focus · HN ↗
        If you read the code you would realize that the issue is....

        in amdkfd and hsakmt

        Yes, that means AMD is sitting on both sides. They wrote software that doesn&#x27;t work with their own software.

    5. bel8 · · focus · HN ↗
      I had a similar but less impressive experience recently with Muse Spark 1.3.

      Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.

      It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.

      1. seanthemon · · focus · HN ↗
        Godot encryption is laughably easy to break, there&#x27;s tons of packages available for it. It&#x27;s a well known drawback of using godot
        1. warkdarrior · · focus · HN ↗
          Now the LLMs know it too.
        2. bel8 · · focus · HN ↗
          I didn&#x27;t say it wasn&#x27;t.

          Still impressive that it did so much just to answer my simple question.

    6. amanguliani · · focus · HN ↗
      Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
      1. onlyrealcuzzo · · focus · HN ↗
        Sol 6.1 is quite good, but damn is it slow.

        I&#x27;m using it to run overnight tasks, and that&#x27;s it until my quota runs out.

        Canceled my subscription.

        1. amanguliani · · focus · HN ↗
          WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what&#x27;s next. Did the same for CODEX - nope still keeps effing going on and on and on
          1. tourist2d · · focus · HN ↗

            [dead]

          2. 8note · · focus · HN ↗
            astra afaict does two stage commits for everything. the first response is a plan, and the second is actually doing it.

            its a lot less chatty imo

          3. awakeasleep · · focus · HN ↗
            If you dig in the settings you can control that. Like on a remote ssh connection, in the settings, you can pick “friendly or terse”

            In the local app interface the winning choice is “efficient” and then turn off the sliders for warmth enthusiasm emoji etc.

            It makes openai models so good to talk to i really have trouble switching.

        2. unconscionable · · focus · HN ↗
          I find Opus 5.5 is better at giving a high level adversarial &quot;should you do this to begin with&quot; where GPT-6.1 is happy to go down any wrong path.

          Also canceled my ChatGPT Pro $200&#x2F;mo subscription. Their Oct 30 price hikes and slow GPT-6.1 model has me looking for alternatives.

          1. directdev · · focus · HN ↗

            [dead]

    7. gottorf · · focus · HN ↗
      My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I&#x27;m not using it for coding, but general research on different topics.
      1. staticman2 · · focus · HN ↗
        The web version of Gemini is awful at search but I don&#x27;t think that&#x27;s the models fault.
      2. MILP · · focus · HN ↗
        I&#x27;m also not using it for coding but I&#x27;ve found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
        1. robobo96 · · focus · HN ↗
          Only html or also css? Opus seems a bit more creative than most other models i&#x27;ve seen.
      3. mattjoyce · · focus · HN ↗
        Hallucination seems a very dated term.
        1. nkozyra · · focus · HN ↗
          Why? It&#x27;s the same concept and root cause it was when we first started using it.
        2. xdavidliu · · focus · HN ↗
          there are many dated expressions, including

          - AI is just a tool, like excel; it does what the human operating it tells it to

          - next token prediction cannot be true understanding

          - models can have no desires and goals, don&#x27;t anthropomorphize it

          However, &quot;hallucination&quot; is very much not one of them

        3. gottorf · · focus · HN ↗
          Hallucination is accurate for what I&#x27;m seeing -- e.g. it&#x27;s making up information about the 2nd gen Toyota Tundra that has no basis in reality. When challenged, it corrects itself.
          1. alluro2 · · focus · HN ↗
            My colleague wanted to diagnose a specific error code on his car himself, and Gemini told him that it&#x27;s simple to do with an OBD2 dongle - he asked it about the details thoroughly, to confirm, and bought the dongle.

            It didn&#x27;t work. Gemini: &quot;Oh yeah, that obviously cannot work, it&#x27;s not possible to do it through OBD2&quot; (paraphrasing)

            It was quite funny to me, but a bit less so to my colleague.

            1. Gareth321 · · focus · HN ↗
              I&#x27;ve had many similar experiences. It&#x27;s confidently incorrect to a shocking degree. Worse than ChatGPT from two years ago.
          2. mattjoyce · · focus · HN ↗
            Its always been a bad term. If we wanted an accurate term then it&#x27;s &#x27;confabulation&#x27;, but &#x27;muddled&#x27; or just &#x27;wrong&#x27; are also good.
        4. rdtsc · · focus · HN ↗
          What do we use for the “model made stuff up and claimed it as facts”? I can see hallucinations somehow anthropomorphizing LLM even more. I don’t like that we’re doing that to begin with but it’s a losing battle. I prefer “it’s broken” and “IT produced shit results” personally.
          1. mattjoyce · · focus · HN ↗
            I agree with you, except I really don&#x27;t hear that term much. People just say it wrong or confused. Good riddance, it was always a bad term.
          2. krapp · · focus · HN ↗
            Models don&#x27;t make claims. That would require a degree of interiority and intent that they don&#x27;t have.

            The bigger problem is that people expect LLMs to know what facts are. That assumption is even baked into the term &quot;hallucination.&quot; Someone who hallucinates is expected to otherwise have a grounding in objective reality, to &quot;not&quot; hallucinate, and to be able to recognize reality from fantasy. We wouldn&#x27;t allow a person who &quot;hallucinates&quot; as much as an LLM anywhere near the roles we give to LLMs. But everything an LLM does is as much a &quot;hallucination&quot; as anything else, it&#x27;s just stochastically generating grammar. Some grammar just happens to be useful because of the quality of its training data, which was probably created by humans who do possess interiority and awareness of fact.

            And it isn&#x27;t &quot;broken&quot; either. Broken assumes that the correct mode of operation is to act as a source of truth or fact generation. When LLMs &quot;apologize&quot; for bad results, for instance they aren&#x27;t actually apologizing. Try getting it to apologize for returning the correct data. It probably will. There is no cognition happening. It doesn&#x27;t know either way. It isn&#x27;t a calculator crunching numbers or a computer doing data analysis. It&#x27;s just pattern matching.

            &quot;Hallucination&quot; is no less correct than &quot;confabulation&quot; which also presupposes intent and contextual awareness. Unfortunately the way LLMs operate is so unintuitive (as opposed to the intuitive nature of the interface) that the only language we have to describe it is the language of human behavior, with all of the biases and false assumptions that brings.

        5. UpsideDownRide · · focus · HN ↗
          They still happen.
      4. WarmWash · · focus · HN ↗
        The achilles heel of 3.8 flash is it&#x27;s january 2025 knowledge cutoff date. Yes, almost 2 years ago.

        I&#x27;m assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.

        1. blinding-streak · · focus · HN ↗
          Incorrect (to some degree)

          &gt; The knowledge cutoff date for Gemini 3.8 Flash is March 2026

          <a href="https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash&#x2F;" rel="nofollow">https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash&#x2F;

          1. WarmWash · · focus · HN ↗
            &gt;The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family). For more information about known limitations, see the Gemini 3.7 Flash

            The &quot;some domains&quot; are very narrow. They likely just RL&#x27;ed popular queries.

        2. Gareth321 · · focus · HN ↗
          While that&#x27;s annoying, the other frontier models easily overcome this with appropriate tool usage. I do a lot of research with frontier models and they&#x27;re very good about identifying where their parametric knowledge is insufficient and searching for the correct knowledge on the internet. 3.8 Flash is HORRIFIC. The majority of the time it doesn&#x27;t use any tools and infers things from its parametric knowledge. Things which should clearly have implied tool calls. Historical statistics, legal precedent, economic data, etc. I think it&#x27;s incredibly clear that it has been tuned for speed and not accuracy.

          Of course, it&#x27;s called &quot;flash,&quot; and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I&#x27;m hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let&#x27;s see.

      5. esafak · · focus · HN ↗
        I would expect a flash model, with its reduced size, to suffer on tail tasks. That is the trade-off you make.
      6. Gareth321 · · focus · HN ↗
        I agree. It&#x27;s much worse than the cheap Chinese models. They appear to heavily bias parametric knowledge and discourage tool use. That&#x27;s fine for things like &quot;how do I perform CPR?&quot; but worse than useless for any kind of research. [There is one benchmark showing far lower rates of hallucination, so let&#x27;s see how accurate this is.](<a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;singularity&#x2F;comments&#x2F;1wuj72j&#x2F;gemini_4_argon_solved_hallucinations&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;singularity&#x2F;comments&#x2F;1wuj72j&#x2F;gemini...)
    8. yegle · · focus · HN ↗
      For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch&#x2F;build it properly.

      And I saw it do this twice, once for Android 14 and once for Android 16.

      I think this is just within 3.8 flash&#x27;s capabilities.

      1. p_l · · focus · HN ↗
        3.8 Flash (but also last two ones) have really strong preference for dissecting binaries with quick thrown-together bits of python in my experience.

        Including going first for decompiling AGY binary instead of searching the web for documentation...

        1. IshKebab · · focus · HN ↗
          Astra also really loves reverse engineering binaries. I guess it&#x27;s one of those things that isn&#x27;t that complicated but is super tedious, and tedium means nothing to AI.
    9. illwrks · · focus · HN ↗
      I&#x27;ve been tinkering with Gemini for several months and I think it&#x27;s great. The most complex things I&#x27;ve had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
    10. martythemaniak · · focus · HN ↗
      Adding my anecdote, because it amused me: I finished wiring up the compute&#x2F;sensor box for my robot, ssh&#x27;d in and told agy &quot;I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output&quot;. It did all the local config for the lidar, setup docker and the ROS2 environment, then told me &quot;open up this url in foxglove&quot; and sure enough everything worked. Whole thing used up 6% of my weekly limit.
    11. Grimburger · · focus · HN ↗
      &gt; I was trying to use rocm with llama.cpp

      completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.

      once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B

      I get the feeling the situation is only going to improve longer term so might be a good time to just do it

      1. esseph · · focus · HN ↗
        &gt; completely offtopic but is rolling with rocm worth it?

        It&#x27;s so fucking easy.

        From AMD: <a href="https:&#x2F;&#x2F;lemonade-server.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;lemonade-server.ai&#x2F;

        Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).

      2. clw8 · · focus · HN ↗
        Support has gotten much better in just the last couple months. I just got a 9070 XT and can&#x27;t count the number of times I&#x27;ve installed a package and the changelog made me think how much it would have sucked to be doing this a year ago.
      3. nzeid · · focus · HN ↗
        I have so much to share on this topic. Will keep it short.

        ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I&#x27;ve been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.

        The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn&#x27;t materialize and the token generation speed decreases.

        There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.

    12. gcy · · focus · HN ↗
      I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it&#x27;s not as capable but it&#x27;s fast and almost free (sufficient quota with pro account). The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.
    13. cyanydeez · · focus · HN ↗
      if you haven&#x27;t tried Qwen3.8-Flash-Next with halogen, you&#x27;re missing out: <a href="https:&#x2F;&#x2F;github.com&#x2F;peonist-ai&#x2F;halogen-flash-server#the-host-settings-these-numbers-were-measured-on" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;peonist-ai&#x2F;halogen-flash-server#the-host-...
      1. solaire_oa · · focus · HN ↗
        I coincidentally just installed this (like 30 minutes ago), and gud dayum, it&#x27;s pretty awesome.

        I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it&#x27;s slop-riddled hallmarks.

        In any case, yeah, ~55 tok&#x2F;s on a high quality model (and massive RAM savings I think?), seems dope.

        1. cyanydeez · · focus · HN ↗
          Yeah, unfortunately, it does deliver for its platform.
    14. lardo · · focus · HN ↗
      Did rocm provide any benefit over vulkan?
      1. taylorfinley · · focus · HN ↗
        Vulkan was reliably crashing after a certain point in the context window, Rocm has been stable as a rock. Tps was basically a wash.
    15. danpalmer · · focus · HN ↗
      3.8 Flash is my daily driver and produces pretty excellent results all round.
      1. dcl · · focus · HN ↗
        What harness you use? Have you tested more than 1?
        1. danpalmer · · focus · HN ↗
          This experience is with Antigravity both internally and externally, and I have done quite a few side-by-side comparisons with the same prompt across a number of different Google and non-Google models.

          I&#x27;ve tried Codex as a harness too, and that was nice. I don&#x27;t find a significant difference between Antigravity and Codex. Codex has more features but I don&#x27;t use them.

    16. iknowstuff · · focus · HN ↗
      lol I&#x27;ve hit the hipStreamCreate problem!
    17. la6479 · · focus · HN ↗

      [dead]

    18. plasticchris · · focus · HN ↗
      This is why I think llms are a killer app for Linux desktop. They’ve been trained on Linux very hard, and it cleanly solves the “how do I make it do $thing” problem since everything is open and the llm can manipulate it. For example: Sound not working? Just tell the llm.
      1. _heimdall · · focus · HN ↗
        Is that specific to Linux? For better or worse, I have felt like it&#x27;d fairly good at working through most tech stacks I throw at it.

        I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.

        1. jubilanti · · focus · HN ↗
          &gt; Is that specific to Linux?

          Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too.

        2. x-complexity · · focus · HN ↗
          Arguably, it requires the LLM to have access to the source code during training. If it&#x27;s closed off (like Windows), then the LLM can&#x27;t train on it.

          Even if the source code&#x27;s old, the fact that it is publicly available makes it much easier to train &amp; improve on than if it were walled off.

          1. augusto-moura · · focus · HN ↗
            Not only training, you can literally clone the source code locally and ask the agent to debug your problem by grepping the source code itself. Even if doesn&#x27;t &quot;know&quot; the answer it can investigate it on the fly
        3. matsemann · · focus · HN ↗
          It&#x27;s at least specific to open source, in my experience. Ask it about some commercial software: it searches some lacking documentation, social media discussions and make some guesses. Ask it about an open source tool: it searches the documentation, if not it dives into the implementation to figure out the edge case.
        4. Sammi · · focus · HN ↗
          The more open the operating system the more LLMs can help you. This is what Arch Linux was waiting for all along.

          I used to be afraid of Arch, because I don&#x27;t want a system that takes work because I&#x27;m already busy with work. But now I love it, because the LLM can tweak every knob and fix every issue for me, so it ends up being the OS that takes the least work to use. Get an error message? Tell the LLM and they fix it. Something not working exactly like you like it? Tell the LLM and they tweak it for you. This also works to great effect on Win and Mac, but not to the same extreme degree as it does on Linux and especially Arch.

          1. lukan · · focus · HN ↗
            Hell yes, I also started to tweak my XFCE arch desktop to my liking.

            Anything I want different now, I just tell the LLM to do it for me.

            (For example I can now close lots of windows of the same type with 1 click, not 3, my whisker search now finds files and folders and I am able to run games that refused before)

            I still ocasionally run into the usual linux driver issues, but not for much longer I suppose. I probably could fix some driver bugs now already if I point fable towards it and pay some attention.

            In theory I could do all this before myself, but not just like that in some minutes, but in days&#x2F;weeks&#x2F;months ..

            1. Sammi · · focus · HN ↗
              Everyone that didn&#x27;t have time to be a neck-beard themselves now has a neck-beard built right into their own computer ready to go.
      2. masto · · focus · HN ↗
        I’ve been playing with the ITS operating system on a simulated PDP-10. Despite it being from the 1960s, Claude has been able to do sysadmin work on it, including a very similar journey where it started disassembling code to identify the source of an obscure problem. I don’t want to take away from the learning experience (the whole reason I’m interested in this stuff), but it’s often almost impossible to find what I’m looking for online, and the LLM turns out to be a better resource when I want to learn how a particular subsystem works. I ended up making a custom MCP server so it could connect to the simulator more efficiently. <a href="https:&#x2F;&#x2F;github.com&#x2F;masto&#x2F;pidp10-mcp" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;masto&#x2F;pidp10-mcp

        Drifted off my point a bit, which I guess was meant to be it’s not necessarily Linux-specific training.

      3. c0n5pir4cy · · focus · HN ↗
        I got Claude to set up 5.1 Surround Streaming from my Linux Desktop to my Steam Deck using Moonlight. It needed some help (or really human ears) to make sure things were coming from the right speakers - but it was pretty much a one-shot.
      4. chinathrow · · focus · HN ↗
        My laptop resumed from sleep all too often and Claude found and fixed it within minutes. Now I still need to file that bug report though...
      5. oldandboring · · focus · HN ↗
        I used Claude to build an entire runbook for how to build my Linux desktop up from scratch to its current state, including all my installed software, configs, hacks, workarounds, etc. Then I used it to migrate from one laptop to another in an afternoon. Now it keeps itself up to date.
      6. safog · · focus · HN ↗
        Yes this might be a bit too noob linux but I have the AIs maintaining an env.md + logs &#x2F; updates on that md whenever they change anything about the env. Even things as simple as getting tmux to be fast, they do an excellent job at.

        Drivers, Coding environment setup etc. are great too and it&#x27;s nice to have everything logged so the next (more powerful) agent can come and improve the thing once in a while.

      7. UpsideDownRide · · focus · HN ↗
        Fully agree. There is so much customization that I&#x27;m doing thanks to LLMs that I just wouldnt to bother to do by myself.
      8. stefan_ · · focus · HN ↗
        Nice idea, then they want to make a quick little kernel module, oops, you secure booted, this is not your system, can&#x27;t load...

        Since LLMs have been so successful at finding exploits it&#x27;s been clearer than ever that the people so obsessed with enshittifying every system with ineffective (other than pissing you off and wasting your time) &quot;secure boot&quot; functionality were really just too ignorant to succeed at actual novel security work, so they focused on this make believe crap. Well, glad that&#x27;s over.

      9. samspot · · focus · HN ↗
        I used Claude &amp; Gemini extensively to troubleshoot Linux since I switched last year, and it&#x27;s been mostly great. But then it nearly bricked my install last week and I had to get some human help. I was getting black screens during and after boot and tried many suggestions from multiple AI&#x27;s. After many hours it turns out I just needed to power cycle my monitor. Shout out to Claude for suggesting that (its memory remembered that I have a dock built in to my primary display).

        I am still very happy with my switch to Linux. But if I didn&#x27;t have the AI help I would say linux is still unacceptable platform for those not willing, able, and excited to get their hands very dirty.

      10. tarokun-io · · focus · HN ↗
        Agreed 100%! I got some Steam games that didn&#x27;t run well (or at all) to run smoothly thanks to AI, after having spent hours trying to figure it out on my own (months prior), fix some issues with HiDPI and external monitor with my laptop, and now I have a tiny shell script that basically runs `checkupdates` (`checkupdates | claude -p`, roughly) and looks online for potential issues. I&#x27;ve been using Linux for many years but still Claude and ChatGPT helped me a ton.
    19. mirmor23 · · focus · HN ↗
      &gt; it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo.

      Most llm could do it. Claude went from firmware thread -&gt; rtos scheduler -&gt; mcu reference manual -&gt; hardware controller register interface -&gt; vendor sdk -&gt; problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.

    20. d3Xt3r · · focus · HN ↗
      That&#x27;s cool, but the real question is, why aren&#x27;t you using Hipfire[1] or HaloPFX[2]? Both are far superior to llama.cpp in terms of performance, for Strix Halo.

      [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;warpfront&#x2F;hipfire" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;warpfront&#x2F;hipfire

      [2] <a href="https:&#x2F;&#x2F;github.com&#x2F;julianmb&#x2F;halofpx" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;julianmb&#x2F;halofpx

    21. qrify_app · · focus · HN ↗

      [dead]

    22. vinzenzu · · focus · HN ↗
      I would like to use my Google AI Pro subscription included with Drive but last time I checked, their terms were not just ambiguous and confusing wrt training and ZDR, they were contradictory. Until they have that fixed and properly communicated, I&#x27;ll stick to other vendors.
    23. crossroadsguy · · focus · HN ↗
      So Gemini Pro (I use agy, because no other harness can be used for subscription plan) was supposed to get Gemini 3.5 Pro, at some point (it was planned, right?), but instead Argon arrives and I am pretty sure it will only be in Ultra sub and don&#x27;t think 3.5 Pro is coming anymore.
      1. p_l · · focus · HN ↗
        3.5 Pro has been pretty much canceled for few months now, 4.0 was in the works instead
    24. KMnO4 · · focus · HN ↗
      My theory is that Gemini 3.8 Flash was supposed to be Gemini Pro, but by the time it was ready for release, Google was embarrassed at how far behind their &quot;pro&quot; model was, so they just named it Flash. It&#x27;s a good model, but don&#x27;t let the name trick you.
      1. w0m · · focus · HN ↗
        I&#x27;d agree if it wasn&#x27;t so much faster than ~all the other models. Maybe they nerfed reasoning to get that speed.
    25. phmx · · focus · HN ↗
      ``` void* mmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) { if (!real_mmap) real_mmap = dlsym(RTLD_NEXT, &quot;mmap&quot;); ```

      hope it&#x27;s not run by multiple threads and dlsym is not allocating.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.