‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. spankalee · · focus · HN ↗
      3.8 Flash is just quite good, and so is the Antigravity harness.

      I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it&#x27;s own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

      1. mapontosevenths · · focus · HN ↗
        Even if agy was the best (it&#x27;s not, and is missing basic features) you wouldn&#x27;t rather have a choice?

        I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

        1. drusepth · · focus · HN ↗
          What basic features are missing from agy? I&#x27;ve been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

          I actually just cancelled Ultra also because I couldn&#x27;t subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

          1. nxdmum · · focus · HN ↗
            IF you havent written your own harness - you would not understand what you can do when you&#x27;re writing your own harness . the current set of harnesses - all of them are crap tier. The only clue i can give you - it&#x27;s not in the model providers interest to have token efficiency - but when you are coding the harness yourself you can shoot for that .

            In today&#x27;s world - and idea stated stated is an idea stolen .

            1. dmos62 · · focus · HN ↗
              What you&#x27;re saying is sus. If you have a harness that&#x27;s a tier above frontier labs&#x27; offerings, link to it.
              1. imtringued · · focus · HN ↗
                It&#x27;s not sus at all. You&#x27;re looking at it from the perspective of a general purpose system that is not adapted to your use case. Basically you prompt the model and then let the agent do everything. That&#x27;s the use case you have in mind when you think that it&#x27;s about being &quot;a tier above frontier labs&#x27; offerings&quot;.

                It&#x27;s missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.

              2. adastra22 · · focus · HN ↗
                Just about every harness is. This is common knowledge, no? Frontier lab TUI tend to steal from the OSS harnesses not the other way around.
                1. dmos62 · · focus · HN ↗
                  Care to back that up with a benchmark? I&#x27;ve not seen that. OpenCode is on here, not exactly dominating though: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents
                  1. adastra22 · · focus · HN ↗
                    Benchmarks are a terrible judge for this.
                    1. dmos62 · · focus · HN ↗
                      Why?
                      1. adastra22 · · focus · HN ↗
                        The process by which benchmarks are setup and run does not correspond at all to how human developers engage with a coding agent. At best it is a loose proxy, and often a bad one.

                        What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.

                        1. dmos62 · · focus · HN ↗
                          I presume you&#x27;re not talking about autonomous agents, right? Because then you could give it a set of tasks and check its success rate. But, even so, are you not giving it tasks? Are you principally interacting with it through discussions that are more difficult to quantify? Even question answering has benchmarks. I&#x27;m having trouble imagining how you use them (or &quot;how people use them&quot; in your words).
                          1. adastra22 · · focus · HN ↗
                            I don&#x27;t know if you noticed, but Opus 4.6 was peak for human-computer interaction. Everything has been fairly downhill from there despite better benchmarks, at least in that one regard. Opus 5 and 5.5 are clearly a step above in capabilities than 4.6, and I don&#x27;t think anyone wants to go back, but 4.7 and 4.8 were arguably worse overall. I genuinely feel I got more done with 4.6 and often switched back, prior to 5 coming out.

                            Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.

                            Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock&#x2F;swarm situation. The benchmark just doesn&#x27;t cover this. (And the difference can be nontrivial! Sakana AI&#x27;s published results show two generations of uplifting potential from better harnesses.)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.