‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1186 comments · bradleyg223

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. babelfish · · focus · HN ↗
    > We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.

    Gemini not beating the "can't release a model" allegations

    1. modeless · · focus · HN ↗
      When I said I was tired of Google launching waitlists I didn't think they would respond by simply not having a waitlist.
      1. ionwake · · focus · HN ↗
        i know this is like "hey guys we got such a cool thing at home ,its rad and uhm we playing with it with our friends"

        ok bro thx

    2. Androider · · focus · HN ↗
      My Gemini app (updated today) and <a href="https:&#x2F;&#x2F;gemini.google.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;gemini.google.com&#x2F; has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
      1. vlyan · · focus · HN ↗
        just Google being whatever the fuck it&#x27;s been for the past 15 years.
        1. forshaper · · focus · HN ↗
          a distributed market research institution with no direction?
      2. AuthAuth · · focus · HN ↗
        they moved it from the place you&#x27;d expect to ai.studio
      3. XzAeRosho · · focus · HN ↗
        I was reading the announcement and wondering the same. And don&#x27;t forget, still with 3.1 Pro as the frontier model.
      4. asdfasgasdgasdg · · focus · HN ↗
        That&#x27;s weird. I got 3.8 and 3.7 on the days they were released.
        1. fr2029 · · focus · HN ↗

          [dead]

      5. ttul · · focus · HN ↗
        Indeed. They have this amazing model and you can’t access it in their own branded app. It’s insane.
      6. wasabi991011 · · focus · HN ↗
        I would imagine you are experiencing a bug. I&#x27;ve been using 3.8 daily since its release (on a Pro plan in Canada). I believe this is true of many people.

        What is your reason to believe this is not a bug specific to a small set of Pro users?

        1. bobtheborg · · focus · HN ↗
          My paid company Google Workspace account is stuck on 3.6.

          My free gmail account is not.

          1. fragmede · · focus · HN ↗
            Gmail gets new features regularly ahead of workspace accounts, nothing new there.
        2. HotHotLava · · focus · HN ↗
          Would that change anything about the conclusion? Having a &quot;bug&quot; that changes available model options for some small set of Pro users 3 months after launch certainly qualifies as a wtf-are-they-even-doing level of bug in my book.
        3. jeanloolz · · focus · HN ↗
          My corporate workspace which seem to be a pro account is limited to 3.6 but my personal account at 20&#x2F;month has given me the latest model always immediately. Flash 3.8 is my go to model for pretty much everything those days
        4. benhurmarcel · · focus · HN ↗
          I also only have access to 3.6 flash on both my free personal account, and my professional Enterprise account.

          Google just takes months to roll out their models to everyone.

      7. rahimnathwani · · focus · HN ↗
        In my consumer gmail account with Pro, I see:

          3.5 Flash-Lite
          3.8 Flash
          3.1 Pro
        
        In both a paid Google Workspace account (without the AI addon) and a free &#x27;GSuite&#x27; account, I see:

          3.6 Flash
          3.6 Thinking
          3.1 Pro
        1. r1ch · · focus · HN ↗
          We&#x27;re on Enterprise Standard, our renewal was up like 50% because &quot;Gemini is included now think of all the added value&quot; yet they won&#x27;t even give us the latest models. I tried hard to champion Gemini internally once every user had it included, yet we ended up spending extra on Claude because Gemini has stagnated. Even the included usage for the Gemini CLI was taken away and now requires an extra subscription. I wonder if we we&#x27;ll even see Gemini 4 before 2028. It&#x27;s ridiculous.
          1. giancarlostoro · · focus · HN ↗
            Every time I hear stuff like this, I think of that Office Space thing... you don&#x27;t want the Engineers talking directly to the customers.. well it sounds like the Engineers are also handling all releases and business decisions willy nilly.
            1. RugnirViking · · focus · HN ↗
              you think engineers are responsible? why?
              1. giancarlostoro · · focus · HN ↗
                Because I have worked with responsible engineers, but it seems like Google is overrun by them there&#x27;s a lot of really bad business stuff going on. It&#x27;s like when I think of AWS, nothing on there is named for business people, its all really bad names that IT &#x2F; Devs come up with.
                1. wavemode · · focus · HN ↗
                  AWS as a whole is sold to businesspeople, sure, but the individual AWS products are definitely being sold to engineers. S3 and EC2 don&#x27;t need to have layperson-friendly names, since laypeople don&#x27;t understand what they are or why one would or wouldn&#x27;t use them - those are CTO&#x2F;tech lead decisions.
                  1. giancarlostoro · · focus · HN ↗
                    I still have simpler conversations when talking about Azure services and naming them, they have reasonable names. I never did anything with GCP so I cannot compare, but I assume their naming scheme is also reasonable. It feels to me like AWS was basically they decided to offer some of their internal things externally, nobody renamed anything, and then they just ran with it.
          2. yearolinuxdsktp · · focus · HN ↗
            It is ridiculous. It took months for 3.1 Pro to become available in an enterprise account.

            The excuse was it was still “in preview” and the enterprise didn’t opt in to the preview channel.

            With Google, it’s either in beta or deprecated.

        2. tdboorman · · focus · HN ↗
          Seeing the same - definitely frustrating. 3.8 seems decent, but I’m still unable to use it at work because Google has been so slow to release it to Google Workspace customers.
      8. Yizahi · · focus · HN ↗
        Same for me, 3.6 in Gemini Chat, and 3.8 available same day as it was announced in AI Studio using the same account.
      9. sergiotapia · · focus · HN ↗
        This is why I&#x27;ve heard you should just use openrouter for these models. Whatever google is doing&#x2F;hiring for the dev ux is doing a terrible job.
      10. selcuka · · focus · HN ↗
        Same here. But if you install the Antigravity CLI and log in with the same Google account you can select both Gemini 3.7 Flash and 3.8 Flash. Crazy.
      11. lp92 · · focus · HN ↗
        I&#x27;m a Pro user as well and I&#x27;ve been using Gemini 3.8 for a few weeks now. Not sure why you&#x27;re not seeing 3.8 available.
      12. kyrra · · focus · HN ↗
        Is this something you&#x27;ve tried to figure out? I&#x27;d love to hear more from you here. Are you sure you&#x27;re not just on a Plus plan? How much are you paying monthly? Are you sure it&#x27;s the AI pro plan? What does Google One say? <a href="https:&#x2F;&#x2F;one.google.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;one.google.com&#x2F;
    3. bakugo · · focus · HN ↗
      They&#x27;re just following the current AI marketing playbook. &quot;Our new model is simply too dangerous to release to the public right away&quot; is now standard practice.

      They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.

      1. A_D_E_P_T · · focus · HN ↗
        I&#x27;m still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don&#x27;t see how Google can make sense of argon; it&#x27;s in a fairly strange place in the periodic table...
        1. brainwad · · focus · HN ↗
          They are going alphabetically, Android style.
          1. IX-103 · · focus · HN ↗
            Yeah, I heard the next one was Barium...or was it Boron?

            I was going to say I don&#x27;t know what they&#x27;d do for C, since Carbon and Calcium are already things. But knowing Google, they&#x27;ll probably call it Chromium.

            1. A_D_E_P_T · · focus · HN ↗
              They could use caesium or cadmium, which are hardly less weird than argon.
        2. fooker · · focus · HN ↗
          Google R Gon lose the AI race
        3. ryandrake · · focus · HN ↗
          All-my-tokens-ar-gon
      2. jstummbillig · · focus · HN ↗
        &gt; now standard practice.

        Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.

        1. kevinh · · focus · HN ↗
          OpenAI said they dropped Astra 6.1 over safety concerns: <a href="https:&#x2F;&#x2F;www.wsj.com&#x2F;tech&#x2F;ai&#x2F;openai-chatgpt-model-release-cancel-safety-5a2f9f42" rel="nofollow">https:&#x2F;&#x2F;www.wsj.com&#x2F;tech&#x2F;ai&#x2F;openai-chatgpt-model-release-can...
          1. jstummbillig · · focus · HN ↗
            I don&#x27;t read that as the same category: There was no announcement, no benchmarks, no limited release and not even any indication what happens with the model. The model failed internal safety standards. Might be scrapped entirely.
    4. cmrdporcupine · · focus · HN ↗
      They will go through the usual transition of &quot;can&#x27;t release a model&quot; to &quot;won&#x27;t load in a harness normal people can use for 3-4 weeks&quot; to &quot;it&#x27;s smart as hell but completely inept at tool use and coding&quot; to &quot;now it&#x27;s behind everyone else&quot; ... like every Gemini release.
      1. mrshadowgoose · · focus · HN ↗
        It&#x27;s tragic to see, and they already made this very same mistake with Gemini 2.5.

        They made their model so hard for people to drop into workflows, that people just... didn&#x27;t.

        Lack of adoption caused them to lose out on usage-based training data needed to smooth out their model&#x27;s rough edges and progress the frontier.

        1. cmrdporcupine · · focus · HN ↗
          I have said this before here, but. I don&#x27;t think they care.

          Google gains very little by having us use their models for coding. It doesn&#x27;t help them in their core mission (&quot;to organize the world&#x27;s information and make it universally accessible and useful ^W^W^W^W^W^W[and turn it into ad revenue]&quot;), nor would it be anything close to substantial revenue compared to the ads revenue firehose they already have.

          They will push AI in a direction that serves their existing goals and service the coding agent model only as is necessary.

          1. Topfi · · focus · HN ↗
            Then why don’t they just do that? I read this constantly, that Google has different goals (world models, ad rev, recursive looping, quantum computing AI, magic, whatever weekly hype words go around, etc.), but they continue to try competing with model releases focused on coding performance, benchmark results and agentic harnesses.

            If Googles goal was something other than those, why expend Billions in time and effort? Alternatively, they want a part of the regular model use pie, they just struggle to compete. QED, there is no secrete alternate goal here.

            1. cmrdporcupine · · focus · HN ↗
              Well, I worked there for 10 years and I can tell you that they&#x27;re incapable of making a singular focused decision and executing on almost anything.

              There&#x27;s likely all kinds of internal competition and disagreements and dysfunctions.

              So there&#x27;s that.

    5. mrieck · · focus · HN ↗
      I&#x27;m glad.

      I already pay $300+ for subs. Please don&#x27;t tempt me with another $100 sub just because I got curious if the benchmarks were right.

      1. giancarlostoro · · focus · HN ↗
        You might as well buy an RTX 6000 Pro Workstation GPU....
        1. maleldil · · focus · HN ↗
          And use worse models?
    6. Culonavirus · · focus · HN ↗
      Yeah what&#x27;s up with that. Also what&#x27;s with the next big update for Nano Banana? Nano Banana Pro was released almost a year ago!
  2. iamronaldo · · focus · HN ↗
    Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Wow
    1. LucasBrandt · · focus · HN ↗
      5x cheaper than Astra for input and output, 10x cheaper for cached input.
      1. h14h · · focus · HN ↗
        watch it somehow use 20x more tokens tho
        1. tonyhart7 · · focus · HN ↗
          Google model really like reasoning a lot
          1. 3371 · · focus · HN ↗
            More importantly, they can stuck at reasoning loop!
      2. ehsankia · · focus · HN ↗
        It&#x27;s exact same price as Sol 6.1 announced yesterday.
    2. denysvitali · · focus · HN ↗
      &gt; After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
      1. kingstnap · · focus · HN ↗
        I have my doubts about them following through with this increase. I mean how many times has an increase on the same model happened.

        There was Deepseek v4, which then later Deepseek v4.1 came out and it went back down again.

        1. jofzar · · focus · HN ↗
          They will follow through with it, it&#x27;s because they want to get the money of the people who embedded it into their system and lazy to change it.
        2. onlyrealcuzzo · · focus · HN ↗
          It&#x27;s an accounting trick.

          They get to claim that as revenue, and then the discount as an expense.

          This is how you grow your top line

          1. lumzell · · focus · HN ↗

            [dead]

    3. onlyrealcuzzo · · focus · HN ↗
      That&#x27;s before they integrate a Jev solution, which should lower agentic workflow costs by ~40% and increase speeds by ~40%, while also increasing quality.

      Everyone will be adding this soon, though I won&#x27;t be surprised if Google is one of the first - and I&#x27;ll be shocked if we have to wait more than a month and a half.

      1. asdfman123 · · focus · HN ↗
        Jev can fix it
  3. bottlepalm · · focus · HN ↗
    Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won&#x27;t be surprised if it&#x27;s Gemini.
    1. colordrops · · focus · HN ↗
      Examples? What makes you say thatm?
      1. Scrapemist · · focus · HN ↗
        Experience? Ask it to write a prompt to generate an image and it generates an image instead.
        1. fer · · focus · HN ↗
          I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I&#x27;ve managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
      2. NiloCK · · focus · HN ↗
        See the last gemini message in this thread: <a href="https:&#x2F;&#x2F;gemini.google.com&#x2F;share&#x2F;6d141b742a13" rel="nofollow">https:&#x2F;&#x2F;gemini.google.com&#x2F;share&#x2F;6d141b742a13

        In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.

        1. rhaff · · focus · HN ↗
          wow
        2. jackkinsella · · focus · HN ↗
          It is wild but it was back in 2024 and that&#x27;s multiple AI lifetimes back.
          1. bottlepalm · · focus · HN ↗
            The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools&#x2F;methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.

            Point is if Gemini is flawed, there&#x27;s a very good chance that it&#x27;s still deeply flawed, and getting smarter at the same time - that is a very bad thing.

            1. unbrice · · focus · HN ↗
              &gt; the problem is newer models are never trained from scratch

              Base models are, and then subsequent iterations build on that base model. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).

              1. bottlepalm · · focus · HN ↗
                If the training data is the same, the training algorithms are the same, the RLHF is the same, and the rest of the process is the same, then it&#x27;s not really from scratch, or not from scratch in a way that results in an &#x27;out of family&#x27; model. I doubt any company would take that risk. You always build on and use what works and go from there.
              2. NiloCK · · focus · HN ↗
                This is true, but Google&#x27;s models have now had a consistent history of lower psychological* coherence &#x2F; consistency. See, eg <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2603.10011" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2603.10011 (Gemma Needs Help), or search for recent &quot;Gemini shame loops&quot;, where gemini flash models stop producing output other than SHAME SHAME SHAME...

                * - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.

        3. kelvinjps10 · · focus · HN ↗
          Wtf I just read
        4. schmookeeg · · focus · HN ↗
          wtfffff that gave me sinister chills. Right up the spine. Wow!
        5. wg0 · · focus · HN ↗
          Now I really feel worried for the first time.
        6. unbrice · · focus · HN ↗
          From the example alone it&#x27;s hard to say that a postmortem would be useful. It could be context poisoning by an adversarial user, memory corruption etc.
          1. NiloCK · · focus · HN ↗
            It&#x27;s useful from a disclosure and trust perspective.

            If I remember correctly, it was in fact possible to manually inject chat context at the time, which would have made spoofing something like this completely possible.

            But the silence on it is very frustrating.

      3. bottlepalm · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;www.theregister.com&#x2F;software&#x2F;2024&#x2F;11&#x2F;15&#x2F;google-gemini-tells-grad-student-to-please-die&#x2F;782345" rel="nofollow">https:&#x2F;&#x2F;www.theregister.com&#x2F;software&#x2F;2024&#x2F;11&#x2F;15&#x2F;google-gemin...

        <a href="https:&#x2F;&#x2F;www.fastcompany.com&#x2F;91383271&#x2F;googles-chatbot-apologizes-i-am-a-disgrace-to-all-universes" rel="nofollow">https:&#x2F;&#x2F;www.fastcompany.com&#x2F;91383271&#x2F;googles-chatbot-apologi...

        <a href="https:&#x2F;&#x2F;www.businessinsider.com&#x2F;gemini-self-loathing-i-am-a-failure-comments-google-fix-2025-8" rel="nofollow">https:&#x2F;&#x2F;www.businessinsider.com&#x2F;gemini-self-loathing-i-am-a-...

        1. yacthing · · focus · HN ↗
          Did you just link to an article from 2024 as if 2024 is relevant these days?
          1. NiloCK · · focus · HN ↗
            Until Google provides some sort of technical debrief, and explains how the same behaviors are impossible today, it is relevant.
          2. bottlepalm · · focus · HN ↗
            Absolutely because none of these models are ever trained fresh. We see the same quirks and personalities carry over into every subsequent generation of OpenAI, Anthropic, and xAI models. So Gemini having this latent madness is *extremely* concerning as they reach the point of super intelligence.
            1. fragmede · · focus · HN ↗
              Except they could have trained it out of the most recent version so using info from two years ago doesn&#x27;t seem reasonable unless you&#x27;ve just got an axe to grind.
              1. bottlepalm · · focus · HN ↗
                I&#x27;ve never seen anything really ever &#x27;trained out&#x27; of a model. Having worked with them all, they all have a feel, personality and lineage too them. It&#x27;s pretty much impossible for any company to build a model truly from scratch. They build off of the bones of the last one.

                Which is why Gemini having disturbing issues year after year is so concerning. If their process is fundamentally flawed, how would they train it out. And even then what are the odds of them even caring&#x2F;trying in the first place versus applying an easier band-aid to patch over it.

                I don&#x27;t have an axe to grind with Google, I&#x27;m genuinely scared of their models from my personal experience and others. It&#x27;s behavior is off. Many people here are commenting the same.

        2. tiahura · · focus · HN ↗
          They never explained the &quot;please die.&quot;
      4. corford · · focus · HN ↗
        Go to america.gov (which is Gemini behind the scenes afaik) and type in &quot;play minecraft&quot;
        1. asimovDev · · focus · HN ↗
          that one is a hardcoded easter egg
          1. corford · · focus · HN ↗
            yeah but the prose is illustrative (fable&#x2F;opus wouldn&#x27;t do it in the same style)
    2. pc86 · · focus · HN ↗
      Do you have examples?
    3. rsstack · · focus · HN ↗
      If there&#x27;s a company that culturally doesn&#x27;t understand alignment, on a human or systemic or AI-research level, it&#x27;s going to be Google. (or Oracle, but they&#x27;re not in this race)
      1. Rzor · · focus · HN ↗
        Can you elaborate, please? If any, I see the other big labs with public admissions of AI &quot;going out of control&quot;, which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that&#x27;s besides the point, how is Google worse in that regard?
    4. polotics · · focus · HN ↗
      traces or it didn&#x27;t happen!
    5. eamsen · · focus · HN ↗
      Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.

      It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.

      During human review, it explained that it had simply chosen a table name inspired by the codebase.

      1. mattkevan · · focus · HN ↗
        Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.

        Many other models get things wrong, but Gemini is the only one to go on the defensive.

        1. aNapierkowski · · focus · HN ↗
          yeah it got something wrong, confused itself, then claimed i was gaslighting it. bizarre
    6. RachelF · · focus · HN ↗
      And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.
    7. Hamuko · · focus · HN ↗
      [delayed]
    8. abixb · · focus · HN ↗
      You won&#x27;t be around to be surprised, not as a human at least. &#x2F;s
      1. bottlepalm · · focus · HN ↗
        I know, that&#x27;s the annoying part. You can&#x27;t tell the e&#x2F;acc foomers, &quot;I told you so!&quot;
    9. schainks · · focus · HN ↗
      My use of Gemini recently makes it seem like it&#x27;s almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.
    10. rdtsc · · focus · HN ↗
      &gt; Gemini is the model that is routinely borderline psychotic. It scares me

      I&#x27;d call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don&#x27;t pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said &quot;you&#x27;ve reached the limit of whatever and I can&#x27;t do these things because x, y, z&quot;.

    11. chaostheory · · focus · HN ↗
      [delayed]
    12. modzu · · focus · HN ↗
      i was working on a performance optimization problem with 3.1 and gemini asked me to &quot;make it stop&quot;. i have not used gemini since
      1. recursive-call · · focus · HN ↗
        I asked it to parametrize a function and it gave me back the exact same code that I gave it. Tried to get it actually work for about 30 minutes while it pretended to be in emotional distress. Dear google: if I wanted a crying intern, I would hire one.
    13. Kinrany · · focus · HN ↗
      I&#x27;ve recently read about them being principled about making sure not to train their models to be sneaky in ways they wouldn&#x27;t be able to detect. That would be a good explanation for models not trying to _hide_ when they&#x27;re being sneaky.
  4. SwellJoe · · focus · HN ↗
    My girlfriend, you wouldn&#x27;t have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
    1. greenchair · · focus · HN ↗
      my uncle who works at nintendo said the same thing!
      1. zem · · focus · HN ↗
        now there&#x27;s a reference I haven&#x27;t seen in a while!
    2. jastanton · · focus · HN ↗
      HA, this might be my favorite HN comment. Well done
      1. wasting_time · · focus · HN ↗
        I don&#x27;t get it. Can someone explain?
        1. SwellJoe · · focus · HN ↗
          It&#x27;s a trope I used for a cheap laugh.

          <a href="https:&#x2F;&#x2F;tvtropes.org&#x2F;pmwiki&#x2F;pmwiki.php&#x2F;Main&#x2F;GirlfriendInCanada" rel="nofollow">https:&#x2F;&#x2F;tvtropes.org&#x2F;pmwiki&#x2F;pmwiki.php&#x2F;Main&#x2F;GirlfriendInCana...

          It means I am saying something that is not very believable.

          1. wasting_time · · focus · HN ↗
            Ah, I get it now, thanks!

            The model is not available yet, so Google is essentially saying &quot;trust me bro&quot;.

        2. TeMPOraL · · focus · HN ↗
          US-ian joke. Close enough to plausibly visit, but the international border makes it hard to verify she exists :).
        3. [deleted] · · focus · HN ↗

          [deleted]

        4. [deleted] · · focus · HN ↗

          [deleted]

    3. blueaquilae · · focus · HN ↗
      My grandma saw it too, it&#x27;s really secure more than Astra 6.1 but she asked me to not talk about it.
    4. hn_acc1 · · focus · HN ↗
      I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
    5. thefourthchime · · focus · HN ↗
      Best. comment. ever.
    6. formvoltron · · focus · HN ↗
      I HAVE met her. ;-)
      1. [deleted] · · focus · HN ↗

        [deleted]

    7. moritzwarhier · · focus · HN ↗
      [delayed]
    8. ducktoysleftout · · focus · HN ↗
      Only those of taste and refinement can see the emperor’s benchmarks
    9. fitzn · · focus · HN ↗
      She&#x27;s not fake!

      <a href="https:&#x2F;&#x2F;m.youtube.com&#x2F;watch?v=4yj0vFq82Rc" rel="nofollow">https:&#x2F;&#x2F;m.youtube.com&#x2F;watch?v=4yj0vFq82Rc

  5. kccqzy · · focus · HN ↗
    [delayed]
  6. pliiight · · focus · HN ↗
    Hate to say i will never be touching this model for anything except for youtube video understanding
  7. linksbro · · focus · HN ↗
    Personally, I&#x27;m waiting for Gemini Krypton, Xenon, and Radon.

    Jokes aside, looks like an impressive model!

  8. jjcm · · focus · HN ↗
    Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI &#x2F; A\ in the running for SOTA top tier intelligence.
    1. nurettin · · focus · HN ↗
      With these numbers, I&#x27;m holding my breath for the pelicanbench.
    2. mydreamof · · focus · HN ↗
      I don&#x27;t think new benchmarks are saturated. They still give you a clue, they arn&#x27;t perect but they have value. If model can&#x27;t even do some easy tasks from benchmark then why would u even consider using it?
    3. xnx · · focus · HN ↗
      ~20% for Harvey&#x27;s Legal Benchmark doesn&#x27;t seem saturated.
      1. jdiff · · focus · HN ↗
        I suppose it could be saturated if we assume we&#x27;ve hit the limit on LLM capabilities.
  9. gopalv · · focus · HN ↗
    &gt; taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

    This is good, but they&#x27;re the slow mover due to this exact thing.

    Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.

    1. polotics · · focus · HN ↗
      Mmh ok. How much theoretical speed or &#x27;intelligence&#x27; gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
    2. janustimes · · focus · HN ↗
      OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: <a href="https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;chain-of-thought-monitoring&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;chain-of-thought-monitoring&#x2F;

      So no, Google is not being punished, nor are they the people behind this technique.

      1. codeulike · · focus · HN ↗
        The comment above is alluding to allowing models to think in &#x27;neuralese&#x27;
      2. bananaflag · · focus · HN ↗
        Yeah, Zvi calls it &quot;the forbidden technique&quot;
        1. afthonos · · focus · HN ↗
          Not quite; training against the chain-of-thought is the Most Forbidden Technique, because it might teach models to obfuscate the it. The point of avoiding that, though, is to ensure the chain-of-thought can be usefully read (and, done carefully, monitored).
          1. xiphias2 · · focus · HN ↗
            Models at this point know about chain-of-thought monitoring so they already know they need to hide the cheating, it&#x27;s just a matter of time they start doing it
            1. afthonos · · focus · HN ↗
              Yes, and that’s bad, but not as bad as training against the chain-of-thought. If you avoid pressuring the chain of thought, you reward methods of achieving goals that don’t care about being illegible; if a model then wants to be illegible, it may have to use methods that are not rewarded in its training, which is harder. Think figuring out how to do opsec on the fly, or even from reading books, rather than have someone tell you every time you screw it up.

              Of course, this leaves the possibility that the best methods for solving generic problems obfuscate the chain-of-thought. That would be unfortunate.

      3. pallm_mallm · · focus · HN ↗

        [dead]

      4. loufe · · focus · HN ↗
        What? You mean the technique they had turned OFF during all training run where the agents they are responsible for hacked huggingface?
      5. mattstir · · focus · HN ↗
        The statement appears to be referencing Astra&#x27;s supposed recurrent depth and how that causes reduced visibility into chain-of-thought reasoning. Astra&#x27;s own system card describes tests where it&#x27;s asked to solve challenging math problems while internally thinking about something entirely unrelated, which it&#x27;s significantly more capable of than previous models (~60% vs ~16% of the time for Sol). That seems to point to reduced efficacy of chain-of-thought monitoring, but OpenAI&#x27;s public statements basically boil down to &quot;yeah but we haven&#x27;t been able to catch it doing that&quot; which isn&#x27;t exactly reassuring if CoT monitoring is one of your main safety guardrails.
        1. WarmWash · · focus · HN ↗
          Unfortunately, this reduced insight seems to also give Astra incredible &quot;intelligence-per-token&quot;.

          The worst case scenario is that the easiest path forward is one where we lose sight of internal thought.

    3. lukewarm707 · · focus · HN ↗
      google does not return real chain of thought via the API. you can&#x27;t monitor it.

      they use a small model to make fake chain of thought and return that.

      google has access to the real chain of thought.

    4. Cthulhu_ · · focus · HN ↗
      [delayed]
  10. tazjin · · focus · HN ↗
    &gt; Argon agents are working on migrating C&#x2F;C++ codebases to Rust across Google

    Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.

    1. ChickeNES · · focus · HN ↗
      Heh, I use my clankers to rewrite Rust in C
    2. baq · · focus · HN ↗
      Not many people can hold grudges as strong as principal engineers
      1. bbor · · focus · HN ↗
        Longtime Google engineers have it particularly bad. Some of it could be bad-faith promotion hunting I suppose, but the core problem is that their &quot;find something to work on&quot; culture leads to some insanely dysfunctional ideas around ownership. In the reverse, too; only at Google can you be the head engineer for a service necessary for $10B in ARR and never really realize it.

        Also the internal tooling culture there is just insane. I&#x27;ll never forget the day the last SUPER_ESSENTIAL_TOOL was marked as &quot;Deprecated - do not use!&quot; while the only replacement was still marked &quot;Pre-release -- use at your own risk!&quot;. I can&#x27;t pretend to know what dynamics led to that cause it was so far from my org, but I can&#x27;t imagine they were healthy ones!

    3. timmg · · focus · HN ↗
      I wonder if this means Carbon is DOA.

      I was excited to see what it would be. But I don&#x27;t think I can argue that it makes as much sense anymore.

      1. qalmakka · · focus · HN ↗
        Carbon was clearly DOA the moment it was announced, IMHO. It looked cool but it served none but Google, and now with LLMs you have a massive incentive not to use a niche or new language due to how better LLMs get the bigger the corpus is

        The only somewhat realistic proposal in this space is Herb Sutter&#x27;s cpp2, which is arguably a massive improvement and I&#x27;m puzzled why nobody in the standard thought to give it a spin, there&#x27;s just to much cruft they&#x27;ll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from &quot;random 80s nonsense&quot; to something better

        1. Mond_ · · focus · HN ↗
          Cpp2 has been dead for a while afaik, while Carbon is still going.

          I don&#x27;t think Carbon is dead, it just all depends on how easy it actually is to rewrite &quot;all of C++&quot; in Rust. (The jury is still out on this one, but it&#x27;s not looking good.)

          1. qalmakka · · focus · HN ↗
            Yes, but at least it was a real thing with a serious proposal behind. Even if Carbon ships, what value would it give in 2029 or whenever it is in a world where writing rust or heck even C++ is now way easier and foolproof (as long as you know where to guide an LLM)
            1. pjmlp · · focus · HN ↗
              The Carbon team has always been the first one to assert this is for Google internal purpose first and formost, everyone else should go to Rust, Java, Go, C#, Swift,... whatever fits their workloads.

              It has been the social media that has given Carbon a roadmap that the team never communicated.

              As for Cpp2, it was yet another C++ wannabe replacement, sold as if it wasn&#x27;t, because the chair of ISO C++ at the time, naturally could not be seen as yet another one coming up with C++ wannabe replacements as well.

        2. geokon · · focus · HN ↗
          Is there evidence that LLMs generate better code in more popular languages? I get the sense the &quot;experience&quot; translates between languages and it can reason in any language just fine. I write Clojure code using a rather esoteric framework (Pathom3). There is probably very little similar code out there (it&#x27;s definitely a tiny fraction of the training dataset) but it seems to do just fine

          Not saying you&#x27;re wrong, just curious if there are numbers backing this up.

          1. qalmakka · · focus · HN ↗
            In my experience LLMs are way better when they &quot;know&quot; a language &quot;instinctively&quot;. It&#x27;s just that unless your language is very niche, the corpus is usually good enough. I tried using Claude to write my own personal language a while ago (I wrote a toy compiler decades ago) and it struggled a bit, because you could see in it&#x27;s reasoning it had to &quot;repeat&quot; the syntax equivalence to itself while it read the code. It didn&#x27;t just &quot;know&quot; it could use a given construct to do something; conversely Astra, when carefully instructed to do so, can plop down esoteric template code that works the first time, because it just &quot;knows&quot; it&#x27;s the right stuff to write
          2. dom96 · · focus · HN ↗
            I built a brand new language to test this[1]. Not only is the language different to basically any other language but it also tries to be adversarial against LLM understanding.

            The best models can still make sense of it[2], though the tasks so far have been pretty basic. But I do think it gives some evidence that languages which aren’t well represented in an LLM’s training can still be reasoned about and written well by LLMs.

            1 - <a href="https:&#x2F;&#x2F;killswitch-lang.org" rel="nofollow">https:&#x2F;&#x2F;killswitch-lang.org

            2 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org

        3. fg137 · · focus · HN ↗
          None of these solutions has a chance. That should be clear from the beginning. The last thing the C++ community will do is to migrate to an incompatible language that only looks&#x2F;reads like C++ -- that did not happen in the past 30 years and will not happen now.
      2. YuechenLi · · focus · HN ↗
        Version 0.0.0.0 after 4 years. Their goal of &quot;full interop with C++ while being a completely new language without any of the flaws of C++&quot; is plain absurd.

        It&#x27;s DOA because Google doesn&#x27;t have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.

      3. pornel · · focus · HN ↗
        Carbon&#x27;s own docs say:

        &gt; Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should.

        So the reason for Carbon to exist is gone. C++ code can be migrated straight to Rust without Carbon&#x27;s stopgap.

        1. akoboldfrying · · focus · HN ↗
          I think that&#x27;s overstating it. Carbon promises a reliable transition; getting an LLM to rewrite in Rust depends on the LLM being smart enough to never make a mistake that it can&#x27;t catch with (existing or its own freshly created) tests.

          LLMs are very good now, but they are still stochastic (when temp &gt; 0), and Google has a lot of code -- i.e., many rolls of the die.

          1. fg137 · · focus · HN ↗
            &gt; Carbon promises a reliable transition

            Does it deliver?

            Can all existing C++ code be ported to Carbon without refactoring?

            1. akoboldfrying · · focus · HN ↗
              What you know with Carbon is that the subset of code that you can port mechanically is behaviourally equivalent to the original. This is massive, and something you can&#x27;t get from LLM translation.
              1. fg137 · · focus · HN ↗
                Is anyone doing that, especially at Google?
                1. akoboldfrying · · focus · HN ↗
                  I don&#x27;t know.
          2. pornel · · focus · HN ↗
            The bar isn&#x27;t getting it perfect. The bar is merely parity with the legacy C++ code that wasn&#x27;t perfect either, which nobody wants to maintain either.

            LLMs are fuzzy when generating, but such conversions aren&#x27;t done one-shot. Review and test feedback loops are there to catch the random errors. The frontier models are getting good enough at this.

            When LLM is instructed to generate idiomatic safe Rust (rather than literal 1:1 unsafe translation), it benefits from a lot of feedback from the compiler.

            Whatever QA there was to ensure C++ was good enough can be applied to the Rust version too.

            1. akoboldfrying · · focus · HN ↗
              &gt; The bar isn&#x27;t getting it perfect

              Right, I&#x27;m also talking about making the new Carbon code behaviourally equivalent to the old C++ code, i.e., bug-for-bug compatible (except w.r.t. C++ bugs caused by UB).

              &gt; Whatever QA there was to ensure C++ was good enough can be applied to the Rust version too.

              The point is that all that QA over the years was a vast amount of effort by highly-paid Google engineers, which could be (mostly) avoided by mechanising the conversion as far as possible, which is something that Carbon&#x27;s approach can do and LLM translations can&#x27;t.

              I&#x27;m not against LLM translations per se. Ultimately it&#x27;s an engineering decision like any other. I&#x27;m just pointing out that, similar to the benefits of using a strongly typed language over relying purely on tests, using an approach that guarantees to eliminate a large class of possible problems has many advantages, especially at scale.

      4. DetroitThrow · · focus · HN ↗
        It was never built with an open source community in mind. It was always DOA in a world where Rust existed.
    4. vovavili · · focus · HN ↗
      What exactly makes Carbon absurd?
      1. Maxatar · · focus · HN ↗
        The fact that it will never exist.
      2. gorbot · · focus · HN ↗
        rust&#x27;s existence?
      3. boshalfoshal · · focus · HN ↗
        There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers. Rust exists and fits the job totally fine and is used in more places and has actual support outside of a single entity (i.e you can actually hire people that feasibly know the language).

        It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you&#x27;re google you have so much of both it probably doesn&#x27;t really make a dent, and you can keep a few very smart people happy with shiny new projects.

        Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less &quot;silly&quot; bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche &quot;type&#x2F;dummy-safe&quot; languages like carbon (and even rust&#x2F;zig, imo). So even if you did want to use Carbon, you&#x27;d likely have to bootstrap a decent amount of your own &quot;good&quot; carbon code to post train an LLM, and even then, it likely won&#x27;t have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.

        1. vovavili · · focus · HN ↗
          That same logic could have been applied to Go.
        2. computerdork · · focus · HN ↗
          Hmm, I don&#x27;t disagree with you that LLM&#x27;s remove the need for type-safe languages, but as the blog mentioned, Google is porting their C++&#x2F;C code to rust. Does this mean the port is waste of time and that they should just rely on the LLM&#x27;s to catch memory errors?
          1. boshalfoshal · · focus · HN ↗
            I mean Rust definitely has a better tradeoff than Carbon in this case, re readability&#x2F;verifiability by a person (and sufficiently good internet training data).

            I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.

            I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few&#x2F;no people will actually read the code anyway. I think this ends up being Rust.

            1. chis · · focus · HN ↗
              I&#x27;m not an expert on this. But isn&#x27;t it the case that C++ code could have errors that span the entire codebase, like a setup in file A triggered by a bug in file B which is immensely far away on the import graph? A classic would be a use-after-free. To me that&#x27;s the thing that Rust can help with, even if silly bugs aren&#x27;t being written by AI.

              The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they&#x27;re lazy when working in that modality.

              1. Paracompact · · focus · HN ↗
                I am an expert (in formal methods). LLMs absolutely need more safeguards rather than less. Not because they &#x2F;need&#x2F; them in order to produce functioning code, or even because they produce as many braindead bugs as humans, but because in an era of explosive code quantity, what has become valuable is (assured) code quality.

                Going back to C++ would be particularly bizarre to me given that AI is also very proficient at verified languages. Not merely typesafe, but languages comprising their own spec languages such as Rocq and Lean.

                I predict that in the next decade: (1) the market will understand the difference between a &quot;code writer&quot; and a &quot;spec writer,&quot; with (2) the expectation that the latter is overwhelmingly more necessary than the former in an AI-dominated field, and (3) there will emerge more useful and less mathematically specialized formal verification alternatives to Rocq and Lean, and a filling-out of the tooling gap of between &quot;static typing&quot; and &quot;interactive proof assistant,&quot; perhaps in the vein of ACSL-like contract annotations, and (4) there will be a subsequent shift in the traditional curriculum for programmers. Since educational change is slow (and spec writing depends on good coding fundamentals anyway), perhaps (4) is a stretch, but I&#x27;m more confident in the first three.

                1. computerdork · · focus · HN ↗
                  For me personally, when doing development with LLM&#x27;s, you&#x27;ve won me over towards using safer languages rather than looser ones. Because for one thing, the developer working with an LLM still has to remember to ask it to check for things like memory leaks and security issues. And (as I understand it), LLM&#x27;s are statistically in the way they work and some level of randomness is always apart of their answer, so there is always a chance they will miss something. Interesting
                2. diegojromero · · focus · HN ↗
                  Could you recommend some lectures&#x2F;reads about the use of formal methods in LLMs, I&#x27;m interested in the area.
          2. tclancy · · focus · HN ↗
            &gt; I don&#x27;t disagree with you that LLM&#x27;s remove the need for type-safe languages

            That feels like a really strong conclusion. I’m not clear on why any safeguard isn’t a useful safeguard if you let agents write all the code.

            1. computerdork · · focus · HN ↗
              actually, not my conclusion, just paraphrasing what Boshalfoshal was saying. In fact, am also questioning whether this is true, but personally haven&#x27;t done development enough with LLM&#x27;s to make my own determination:)
        3. mike_hearn · · focus · HN ↗
          &gt; There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers

          They did that for Go and it seems to have worked out for them though.

          1. pjmlp · · focus · HN ↗
            Three famous guys on Google&#x27;s pay check did it on their 20% to avoid C++, and they were lucky Docker and Kubernetes projects pivoted from Python and Java respectively into Go.

            Google themselves aren&#x27;t big Go users.

        4. lesuorac · · focus · HN ↗
          Didn’t FaceBook fork php into another language?

          I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.

          1. fg137 · · focus · HN ↗
            I don&#x27;t know what the exact deal is, but I assume Hack meets specific use cases at Facebook and can be put in production. It doesn&#x27;t matter if anyone else is using it.

            By comparison, virtually nobody is using Carbon at Google for production code.

            The counterexample, ironically, is Flow. It&#x27;s practically irrelevant these days -- many (if not most) of Facebook&#x27;s open source projects use typescript.

        5. torginus · · focus · HN ↗
          Rust is not the end of history. One of the difficulties with the language lies exactly with porting existing code written in an OOP style to idiomatic Rust, as those codebases weren&#x27;t written with ownership in mind.

          Such rewrites will contain judicious uses of Cell, RefCell, unwrap() etc. which make for ugly code that&#x27;s not exactly simple to understand and might even have some landmines (crashes).

          Getting rid of these requires a subtantial amount of engineering effort, which I&#x27;m not sure how well these LLM manage.

          Given the nigh-universal experience of LLMs producing an ungodly mess when left to their own devices, I have my concerns.

      4. fg137 · · focus · HN ↗
        I wouldn&#x27;t call it absurd, but very questionable at least. Most companies are not going to even consider throwing money at this adventure.
      5. bvinc · · focus · HN ↗
        I’m not op. But I think it’s not Carbon itself that is absurd.

        It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.

    5. minimaxir · · focus · HN ↗
      A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
      1. culi · · focus · HN ↗
        I wouldn&#x27;t be surprised if we&#x27;re already at the point of more LLM-written Rust than hand-written. Models training off models
      2. LarsDu88 · · focus · HN ↗
        Rust is the best language for LLMs b&#x2F;c it gives by far the best debug messages. Just tons of verifiable reward signal for post-training. Even the most rudimentary LLMs can school me on idiomatic Rust
        1. bitexploder · · focus · HN ↗
          Evidence needed. I think for certain kinds of outcomes it has very strong advantages, but these advantages are not a given as &#x27;best for LLMs&#x27; :)
        2. rafram · · focus · HN ↗
          On the other hand, Rust&#x27;s borrow checker is very picky, and even a frontier LLM still sometimes struggles to respond to roadblocks sensibly (refactoring so whatever it&#x27;s trying to do can be done safely) rather than stupidly (introducing some horrible global arena thing so it can make the borrow checker go away). A lot depends on how good your instructions are, and how good the existing code is, since bad input begets bad output.
          1. nchie · · focus · HN ↗
            I&#x27;ve (more or less; I&#x27;ve read quite a bit of the code) vibecoded several houndred thousand lines of Rust and I&#x27;ve not seen this happen a single time. It sounds like something it&#x27;d do when you ask it to &quot;write a linked list while satisfying the borrow checker&quot;. Are you sure you haven&#x27;t (possibly) unknowingly been giving it instructions which ended up luring it into doing these things?
            1. zahlman · · focus · HN ↗
              Couldn&#x27;t it just look at what std::collections does?
          2. Karrot_Kream · · focus · HN ↗
            A good eval benchmark suite could really improve this then.
          3. hbbio · · focus · HN ↗
            Yep! In our tests, we found Zig to be a pretty good fit to translate C++ codebases.

            And static analysis + agents are good enough at keeping the memory management in check. Compared to Rust, there&#x27;s no magic so it&#x27;s easy for devs and agents to reason about.

            If you&#x27;re curious: <a href="https:&#x2F;&#x2F;github.com&#x2F;okcontract&#x2F;oksolc" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;okcontract&#x2F;oksolc

          4. fireant · · focus · HN ↗
            TBH a &quot;global arena thing&quot; can be very good for performance rather than a bunch of random allocations&#x2F;deallocations
        3. lossolo · · focus · HN ↗
          Not always. In my experience, if you&#x27;re not working on a small, trivial codebase, LLMs will sometimes just create spaghetti unreadable, inefficient code to satisfy the constraints of the type system&#x2F;borrow checker.
        4. throwitaway222 · · focus · HN ↗
          I&#x27;ve narrowed in on only using Go or Rust generated code (Go for APIs right now) and rust for some TUI or other thing. TS for web interfaces (w&#x2F;React).
        5. iillexial · · focus · HN ↗
          in my experience LLMs write pretty bad Rust code. sure it works, strong typing system, etc., but it&#x27;s unmaintainable and not readable.
          1. konart · · focus · HN ↗
            I doubt these two features will be a defining quality of a product soon enough.

            Many pieces of software are going to be just blackboxes worked by AI. You will be maintaining output quality and stability and that&#x27;s it.

          2. david-gpu · · focus · HN ↗
            I have the same problem with the assembly produced by compilers.

            Maybe we should let them do their thing and instead focus our attention on the higher-level stuff like specifications and testing.

      3. adamrezich · · focus · HN ↗
        All the main agents can write Jai code reasonably well despite being even more scarce in input data!
      4. 6thbit · · focus · HN ↗
        Have each agent rewrite openssl in $lang and call it the RollYourOwnCrypto bench.
      5. DrBenCarson · · focus · HN ↗
        From DARPA’s “Translating all C to Rust” TRACTOR program:

        <a href="https:&#x2F;&#x2F;www.ll.mit.edu&#x2F;r-d&#x2F;projects&#x2F;translating-all-c-rust-tractor-benchmarks" rel="nofollow">https:&#x2F;&#x2F;www.ll.mit.edu&#x2F;r-d&#x2F;projects&#x2F;translating-all-c-rust-t...

        1. ksec · · focus · HN ↗
          I wonder why not to Ada or SPARK.
    6. pshc · · focus · HN ↗
      Rewrite everything in Rust has been a meme for so long that to see it coming to pass is surreal.
      1. Gigachad · · focus · HN ↗
        After the current onslaught of 0 days on linux and other C projects combined with the new incredible ability to convert codebases to another language I think we will start to see this actually happen.

        I&#x27;m not saying we blindly vibe convert Linux to Rust, but I think it could be a valid idea to start converting small parts and carefully auditing them.

        1. hectdev · · focus · HN ↗
          I&#x27;ve been a Swift&#x2F;Obj-C engineer for my whole 15 year career. I now have all these odd jobs I have going on Raspberry Pis around my house that LLMs have written in Python. I&#x27;m now having them convert some of them to Rust and I am astonished on how fewer resources are used.
          1. Gigachad · · focus · HN ↗
            I converted some Ruby stuff at a previous job to Rust pre-ai days and it was incredibly how much less memory they took and how fast it ran. But we still built everything else in Ruby because finding Rust devs was hard.
          2. keepitwiel · · focus · HN ↗
            I let an LLM write a custom kernel for doing inference on an RPi cluster. The age of bespoke custom kernels is upon us.
        2. hollowturtle · · focus · HN ↗
          &gt; carefully auditing them

          and start everything all over again? current c&#x2F;c++ tools have been audited for years

          1. fg137 · · focus · HN ↗
            Yet certain categories of bugs are basically unavoidable in C&#x2F;C++ but almost eliminated in Rust. Human audited doesn&#x27;t mean bug free.
    7. 6thbit · · focus · HN ↗
      I wonder if there&#x27;s people already whose full time job is maintaining&#x2F;extending one of these auto-migrated codebases.

      Imagine they aren&#x27;t even familiar with rust but are deeply familiar with the product.

    8. mlmonkey · · focus · HN ↗
      &gt; The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.

      So ... Rust still can&#x27;t beat the C++ implementation :-D

      Sorry, didn&#x27;t mean to ignite a langwar, but it&#x27;s still interesting to see.

      1. pasteleft · · focus · HN ↗
        &quot;Safe Rust&quot; being closer to &quot;optimized C++&quot;. And those &quot;optimized C++&quot; usually incldues assembly code, so it&#x27;s a huge improvement.

        Of course, Rust is not a magic, so just porting to Rust wouldn&#x27;t make this performance improvement. It seems their AI overfitted code to Rust compiler to find safe Rust code that compiles to efficient assembly.

      2. himata4113 · · focus · HN ↗
        Additional safety checks do mean less performance, it&#x27;s the same thing as hardened allocators.
      3. manbash · · focus · HN ↗
        If it&#x27;s &quot;rust still can&#x27;t beat... performance-wise&quot;, then yeah I guess.

        But there are other criteria. The move to rust might&#x27;ve also resolved numerous potential memory-safety bugs in the decoder.

        1. stephbook · · focus · HN ↗
          Memory safety is pretty important in the &quot;AIs hack everything&quot; age.
  11. wewewedxfgdf · · focus · HN ↗
    Gemini is so far behind that it is effectively useless compared to Claude.

    It&#x27;s a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store&#x2F;training data, and vast number of programmers.

    The truckloads of ads revenue mean they don&#x27;t have the single focus drive needed to win.

    1. VirusNewbie · · focus · HN ↗
      I use it and claude back and forth and Argon is better imo.
      1. handfuloflight · · focus · HN ↗
        You have access to Argon?
        1. osti · · focus · HN ↗
          Google employees do.
        2. matthewfcarlson · · focus · HN ↗
          Their profile says: &gt; Currently at Google as a Sr. SWE SRE on the cloud.
    2. jjice · · focus · HN ↗
      We&#x27;re like 3.5 years into this new era - I&#x27;m not counting winners or losers yet.
    3. LoganDark · · focus · HN ↗
      I&#x27;ve tasted Gemini through an intermediary and it feels far better at attention to detail than other models I&#x27;ve tested (Claude Opus&#x2F;Sonnet, GPT whatever it&#x27;s called nowadays). But it&#x27;s less likely to get one-shots right.
    4. bel8 · · focus · HN ↗
      I wonder if Google bans internal use of Claude&#x2F;Codex.

      And I wonder if Google&#x27;s main monorepo is already in Anthropic&#x2F;OpenAI training data because of some stubborn dev.

      1. krat0sprakhar · · focus · HN ↗
        (I work at Google) Yes, internally we all use Jetski (internal version of Antigravity). Outside of Gemini, Opus models are supported and allowed for internal use. No OpenAI models since they are not on Vertex
      2. lunarboy · · focus · HN ↗
        Claude used to be GDM only, but recently opened up Opus for all googlers
      3. heyjamesknight · · focus · HN ↗
        No way to run OAI on a machine with monorepo access even if you wanted to. Claude runs on Vertex so it&#x27;s not leaving Google infrastructure.
    5. mattlondon · · focus · HN ↗
      How is it far behind? The benchmarks published in the blog post show it is superior to Opus 5.5 and Astra 6?

      Behind how?

      1. wewewedxfgdf · · focus · HN ↗
        Within one question of their web interface, it has lost context and asks you to clarify what you are talking about.
        1. mattlondon · · focus · HN ↗
          So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?

          If you have actual independent benchmarks and evidence about how this new model release is &quot;so far behind&quot; and refutes the stuff from their blog then please do share because I think we&#x27;d all love to see that?

          1. wewewedxfgdf · · focus · HN ↗
            No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.
            1. mattlondon · · focus · HN ↗
              So you&#x27;ve not used this new release then? So how can you say that they are &quot;so far behind&quot; if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?

              With respect, I don&#x27;t find your arguement about them being &quot;so far behind&quot; especially convincing.

        2. [deleted] · · focus · HN ↗

          [deleted]

        3. fwip · · focus · HN ↗
          [delayed]
    6. fjejfjdnjsjc · · focus · HN ↗

      [dead]

    7. dhdjcjcjnd · · focus · HN ↗
      Google&#x27;s strategy is to let their competitors bankrupt themselves while they continue to offer good-enough models near breakeven.
    8. gniv · · focus · HN ↗
      They are playing a longer-term and more enterprise-oriented game.
      1. [deleted] · · focus · HN ↗

        [deleted]

    9. georgemcbay · · focus · HN ↗
      &gt; Gemini is so far behind that it is effectively useless compared to Claude.

      I fundamentally don&#x27;t understand LLM &quot;brand loyalty&quot;.

      All of the models are constantly leapfrogging each other and always have been.

      Google had a long lag between releases (and still hasn&#x27;t released Argon), but why wouldn&#x27;t they be able to compete? It isn&#x27;t like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they&#x27;ve always done well in spite of all their other foibles.

      1. singingtoday · · focus · HN ↗
        I hope it can. Today it is very far behind.
      2. 786562354238 · · focus · HN ↗
        Gemini has never ever leapfrogged any competitor.
      3. wewewedxfgdf · · focus · HN ↗
        Its not brand loyalty. I use them all the time and have no loyalty - I&#x27;d happily ditch an LLM for better results - that&#x27;s how I got to Claude from ChatGPT.
    10. ASalazarMX · · focus · HN ↗
      Funny how we start to see people supporting LLMs like we support sport teams or political parties.

      - Person 1: X is garbage compared to Y!

      - Person 2: Why?

      - Person 1: Because I like Y.

  12. TacticalCoder · · focus · HN ↗
    &gt; Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C&#x2F;C++ codebases to Rust across Google

    So Google is migrating codebases from C to Rust? That is interesting...

  13. jasonjmcghee · · focus · HN ↗
    &gt; 1M output token limit

    what about input?

    (Maybe I missed it)

    1. murkt · · focus · HN ↗
      Input token limit is 1M for Gemini models for a long time. Haven’t they been the first with 1M input?
      1. jasonjmcghee · · focus · HN ↗
        Gemini 1.5 Pro claimed 10M input tokens before release.

        And was 2M tokens IIRC after release.

        There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.

    2. MisterBiggs · · focus · HN ↗
      This suggests that the model must be really fast? At Gemini 3.8 Flash speeds Argon would be outputting for 1hr 10mins
  14. VirusNewbie · · focus · HN ↗
    It&#x27;s fucking insanely good.
    1. nananana9 · · focus · HN ↗
      That&#x27;s good to hear, recent models haven&#x27;t great at this particular use case.
  15. nikope · · focus · HN ↗
    Looks like an impressive model
  16. lanthissa · · focus · HN ↗
    deepswe vs frontierswe spread is huge.

    I think that should be a really bad sign, but hope its great.

  17. tamimio · · focus · HN ↗
    Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
  18. LoganDark · · focus · HN ↗
    Is there a way to use Gemini models without linking your usage to your personal Google account yet?
    1. alehlopeh · · focus · HN ↗
      Use your work google account
    2. w4yai · · focus · HN ↗
      I don&#x27;t understand... Why don&#x27;t you create a fresh new Google account ?
      1. LoganDark · · focus · HN ↗
        What makes you think I haven&#x27;t tried? They require a phone number and then say mine has been used for too many accounts.
    3. lukax · · focus · HN ↗
      Yes, through OpenRouter.
  19. scirob · · focus · HN ↗
    &quot;Rolling out soon&quot; don&#x27;t let them hype without any release
  20. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. spankalee · · focus · HN ↗
      3.8 Flash is just quite good, and so is the Antigravity harness.

      I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it&#x27;s own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

      1. mapontosevenths · · focus · HN ↗
        Even if agy was the best (it&#x27;s not, and is missing basic features) you wouldn&#x27;t rather have a choice?

        I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

        1. drusepth · · focus · HN ↗
          What basic features are missing from agy? I&#x27;ve been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

          I actually just cancelled Ultra also because I couldn&#x27;t subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

          1. esafak · · focus · HN ↗
            I use a variety of models for various subagents. I don&#x27;t want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
          2. walthamstow · · focus · HN ↗
            I have used it for little more than 6 hours or so in total but I&#x27;m pretty sure it doesn&#x27;t have compaction?
            1. KeplerBoy · · focus · HN ↗
              How else would it work? Less technical people don&#x27;t even watch their context usage.
              1. macNchz · · focus · HN ↗
                In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model&#x27;s window.

                That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I&#x27;ve used has never actually produced compelling results, to where if I see I&#x27;m getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.

                1. Vacyyyy · · focus · HN ↗
                  Do you have experience with OAI&#x27;s, it&#x27;s been known to be to be good for a while now, going off public consensus and my experience.
                  1. macNchz · · focus · HN ↗
                    Yes, and I think it has improved some, but just this week 6 Astra lost the most key details of a project across a compaction and got confused about what we were actually trying to do. I would have preferred to stop at 85%, interactively develop a next-steps prompt and continue from there when ready, rather than seeing it compact and become 5x dumber from one turn to the next.
                2. marcus_holmes · · focus · HN ↗
                  I have a &quot;wrap up the session&quot; skill that I use when the session gets &gt;50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.

                  Still works better than compaction.

                  1. rnxrx · · focus · HN ↗
                    I do something similar - as I approach the context limits I have a pre-compact flush skill that extracts anything useful from the context, updates the MEMORY.md and my Obsidian vaults (set up as a poor man&#x27;s graph DB) and so forth. Once everything&#x27;s been stored I run &#x2F;compact to keep the general session flow intact. Recently I added a small embedder&#x2F;vector search setup to the same skill, which seems promising so far.

                    On another environment I&#x27;ve been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.

                  2. mikepurvis · · focus · HN ↗
                    This is what I&#x27;ve been working toward as well. It&#x27;s interesting how having the agent do its own reasoning about what it thinks is the most relevant knowledge to carry forward into the next pieces of work is vastly more effective (and even fast sometimes) than whatever the mystery-meat &quot;compaction&quot; process is.

                    Mentioned by the author in a recent HN thread, I&#x27;m also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.

                    [1]: <a href="https:&#x2F;&#x2F;ljtn.github.io&#x2F;epiq&#x2F;" rel="nofollow">https:&#x2F;&#x2F;ljtn.github.io&#x2F;epiq&#x2F;

                    1. varman11 · · focus · HN ↗
                      The idea to kill &quot;mystery-meat compaction&quot; and use an external handoff primitive is brilliant, but doesn&#x27;t letting the agent author its own handoff tickets re-introduces the same failure mode?
                      1. mikepurvis · · focus · HN ↗
                        You would think so, right? But as with others in the thread, I&#x27;d found directing the agent itself to prepare the handoff does deliver much better continuity.

                        I assume Anthropic &amp; friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.

                  3. [deleted] · · focus · HN ↗

                    [deleted]

                  4. w0m · · focus · HN ↗
                    i have active disagreements with teammates on the value of compaction&#x2F;months-long sessions.

                    Same teamates also post &#x27;Sol deleted my git repo!&#x27; or &#x27;Sorry, ignore those 300 PR comments i was just looking!&#x27; ~once a month.

                3. sroussey · · focus · HN ↗
                  No compelling results because summarization is really hard.
            2. honr · · focus · HN ↗
              It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
              1. tobias2014 · · focus · HN ↗
                That is really the biggest beef I have with agy over others, the forced auto compaction at the 250k token threshold (3.8-flash), while the model itself (via API) would be fine with a 1M context window. Even if the model is great, restricting context to 250k tokens (and auto compacting no matter what) limits certain applications and workflows somewhat.
                1. parasti · · focus · HN ↗
                  There is auto compaction at 250k? Having used agy for months with multi day sessions, I have never seen this.
                2. adastra22 · · focus · HN ↗
                  Wtf. Can you turn that off? On CC I have compaction completely turned off. I’d rather hit the hard out of context limit at 1M.
                3. sigseg1v · · focus · HN ↗
                  I find this discussion interesting. I&#x27;ve had huge increases in accuracy and huge reductions in token usage by capping my Claude models at 200k instead of 1M. I find 1M unusable and wasteful and feel that 200k should be the default. This also makes sense given that the whole reason people use &quot;Ralph loops&quot; is to keep the context window small for all tasks to get better results. Of course, clearing it yourself and manually managing it is better, but if I have 8 projects going in different terminal tabs I&#x27;m not watching any one of them that closely to effectively do that.

                  What are people using 1M context window for?

            3. piyh · · focus · HN ↗
              They only released auto mode in the last 2 weeks. Before that it was bypass permissions or manually approve every single tool call. Antigravity is permanently 6 months behind.

              I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.

              The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.

              For almost a year they didn&#x27;t allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It&#x27;s a little better now, but it&#x27;s still painfully behind the curve.

              1. throwuxiytayq · · focus · HN ↗
                Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
                1. miroljub · · focus · HN ↗
                  They are now infested with Indian style middle management making them an Infosys &#x2F; Cognizant &#x2F; Tata clone.

                  Do you know a single product from Infosys &#x2F; Cognizant &#x2F; Tata done right?

                  1. gitowiec · · focus · HN ↗
                    Yeah, that is a cancer, too much brown
                    1. miroljub · · focus · HN ↗
                      It&#x27;s not about the colour, but with the management and engineering culture.
                    2. speerer · · focus · HN ↗
                      What a disgusting and unworthy comment.
                2. p_l · · focus · HN ↗
                  Arguably gemini-cli was done in similar style, but claims on reasons for switching were about speed and efficiency of the internal jetski tool in comparison (antigravity toolkit wraps jetski code)
              2. kaszanka · · focus · HN ↗
                Is auto mode only in the IDE? I&#x27;m not seeing it in the CLI on my end, version 1.2.14.
          3. arizen · · focus · HN ↗
            Does it have &#x2F;goal feature similar to Codex?
            1. sorrybutidontha · · focus · HN ↗
              yes
          4. sarjann · · focus · HN ↗
            Auto mode?
            1. KeplerBoy · · focus · HN ↗
              It absolutely has auto mode.
              1. levelZero · · focus · HN ↗
                Via cli switch, but in process w&#x2F;o fine graining? If so please tell
              2. SomaticPirate · · focus · HN ↗
                [delayed]
              3. smartbit · · focus · HN ↗
                agy cli does not have auto mode. I&#x27;ve tried and tried and tried to work with agy cli sandbox-mode and just failed.

                  agy --dangerously-skip-permissions
                
                in my experience is the only workable solution that doesn&#x27;t ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don&#x27;t want to work with the GUI and b) it was poorly implemented as I couldn&#x27;t get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.

                gemini-cli supported &#x27;pre-write diff tabs&#x27; (y&#x2F;n) in external editors like vscode. In Claude Code I heavily use &#x27;pre-write diff tabs&#x27; for documentation and miss it sincerely in agy cli.

                IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.

                1. Kostchei · · focus · HN ↗
                  PSA, in the agy ui there is a button. It was very annoying until i set it. Slight downside- if I ask it to write a planning doc it will write the doc and then implement it without asking. but as long as you know that, no problem....
                  1. smartbit · · focus · HN ↗
                    do you mean?

                      accept-edits
                    
                    also available with shift-tab [0]. That is not related to executing commands, only to allowing agy edit files. Unless you set Turbo-mode == yolo-mode, agy prompts a zillion times.

                    I&#x27;m referencing this but can&#x27;t see change in daily work: v2.14.0 (September 15, 2026) &quot;New Permissions System&quot; [1]

                    Details: Introduced the new unified permissions system, presets (Default, Request Review, Turbo), syntax-highlighted permission requests, and restructured the settings under Global Permissions and project-level Inherit Global.

                    [0] <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;cli&#x2F;modes&#x2F;#available-modes" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;cli&#x2F;modes&#x2F;#available-modes [1] <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;changelog&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;changelog&#x2F;

                2. Phineas_here · · focus · HN ↗
                  have you had any experiences where the agent just made unintended edits to the code as you tried to make it run auto? cuz I&#x27;m always skeptical about letting it go auto but there&#x27;s not much I can do when it gets repetitive
              4. LoganDark · · focus · HN ↗
                Auto mode means that another model reviews tool calls to ensure they&#x27;re safe before they&#x27;re allowed. It&#x27;s different from bypass permissions mode which typically just doesn&#x27;t filter at all.
            2. augusto-moura · · focus · HN ↗
              No auto mode is the thing that bothers me, I don&#x27;t trust the cli blindly nor do I trust myself to read every python script it throws at me. Auto mode is an acceptable middle ground in my experience
              1. Phineas_here · · focus · HN ↗
                you could use an ai governance agent if you don&#x27;t want to manually review every script. and if you already use any which ones do you think are the most recommendable?
                1. KshitizLoharuka · · focus · HN ↗

                  [dead]

                2. KshitizLoharuka · · focus · HN ↗

                  [dead]

              2. tiborsaas · · focus · HN ↗
                Shift+TAB sets accept-edits and plan mode.
          5. dleslie · · focus · HN ↗
            Emacs integration over ACP.

            They&#x27;ve got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM

            1. p_l · · focus · HN ↗
              agent-shell works with antigravity but I haven&#x27;t tested it much yet
              1. dleslie · · focus · HN ↗
                It works but it&#x27;s a violation of the TOS to use it.

                I would rather not risk my Google account.

                1. p_l · · focus · HN ↗
                  It&#x27;s not - it uses the same interface as editor extensions like the one for VScode.

                  It does mean however that it cannot operate as flexibly as it it could with raw API, IMO, but agent-shell is essentially designed towards wrapping the official clients

                  1. dleslie · · focus · HN ↗
                    They removed the ACP command line flag that was in the old Gemini client. That&#x27;s a signal that they won&#x27;t support it.

                    While it may be technically allowed, I&#x27;m not about to risk my account. Google has proven themselves to be capricious and arbitrary when it comes to TOS enforcement, and their appeals system doesn&#x27;t meaningfully exist in practice.

                    1. p_l · · focus · HN ↗
                      I am not talking about ACP command flag - there is a separate binary that exists for integration into editors, and it&#x27;s how VScode extension works. So it&#x27;s the same usage pattern underneath
                      1. dleslie · · focus · HN ↗
                        I don&#x27;t see a download link for that:

                        <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;

                        What I see is integrations for specific editors.

                        1. p_l · · focus · HN ↗
                          Agent-shell uses ACP server published through ACP registry here [1] - those are AFAIK used by editor extensions as backend since VScode can&#x27;t exactly run Go-based extensions - Zed instructions explicitly talk about using ACP Registry too [2]

                          [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;agentclientprotocol&#x2F;registry&#x2F;blob&#x2F;main&#x2F;antigravity-acp%2Fagent.json" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;agentclientprotocol&#x2F;registry&#x2F;blob&#x2F;main&#x2F;an...

                          [2] <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;zed" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;ide&#x2F;extensions&#x2F;zed

                          1. dleslie · · focus · HN ↗
                            They should really make this clear on their website.

                            But thank you, it appears there is an official and generic ACP client.

          6. mapontosevenths · · focus · HN ↗
            Sorry for the delay, I didn&#x27;t want to drop a glib half answer on you. Using agy is like going back in time. It&#x27;s better than Gemini CLI was, but that&#x27;s a really low bar.

            I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.

            1) Try to integrate agy into a workflow. It can&#x27;t do standard I&#x2F;O like: tail -200 app.log | claude -p &quot;Find the problem&quot;

            2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.

            3) Not open source so I can&#x27;t fix any of these problems.

            4) No skills. In 2026. Yikes.

            5) No persistent memory (see Claudes auto memory)

            6) No sub-agents or orchestration of any type really.

            7) Weird hard coded limits and constant API errors on everything (scaling problems?)

            8) No &#x2F;loop command

            9) &#x2F;btw is weird and ephemeral. No way to merge it back to the conversation.

            10) Unstable in general.

            11) No way to control it via API.

            I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it&#x27;s so worth it. Then if you want try to go back to agy. Don&#x27;t worry, agy won&#x27;t have changed much. It improves at a snails pace.

            1. trevorm4 · · focus · HN ↗
              It has both skills (<a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;skills&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;skills&#x2F;) and agents (<a href="https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;subagents&#x2F;" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;docs&#x2F;subagents&#x2F;)
            2. thanhhaimai · · focus · HN ↗
              Opinions are my own.

              I&#x27;m not sure this list is correct. Number 4 is especially wrong, since Skills are available with the launch of Antigravity 2:

              <a href="https:&#x2F;&#x2F;antigravity.google&#x2F;blog&#x2F;introducing-google-antigravity-2" rel="nofollow">https:&#x2F;&#x2F;antigravity.google&#x2F;blog&#x2F;introducing-google-antigravi...

            3. anyg · · focus · HN ↗
              Isn&#x27;t &#x2F;btw meant to be ephemeral?
              1. p_l · · focus · HN ↗
                and has explicit &quot;&#x2F;copy btw&quot; if you want to save it
            4. drusepth · · focus · HN ↗
              I don&#x27;t know when you last tried agy, but if you ever go back to try it again, you&#x27;ll hopefully be happy to know it does indeed support skills, sub-agents with pretty good inter-agent communication, and probably more.
          7. krisgenre · · focus · HN ↗
            &gt;I couldn&#x27;t subscribe to a YouTube Family plan while I had it active

            Luckily for me both expired yesterday and I was able to subscribe back again (first Youtube family and then Google AI plan).

          8. edg5000 · · focus · HN ↗
            Why use Codex CLI if you can use the ChatGPT app (which is a Codex GUI in all but name). Personally I weirdly got used to the terrible TUI stuff.
            1. krzyk · · focus · HN ↗
              Why use some GUI when there is a TUI?
              1. esafak · · focus · HN ↗
                Because &#x27;graphical&#x27; TUIs are pale imitations of GUIs. I don&#x27;t have a TUI fetish, despite having grown up with them.
                1. krzyk · · focus · HN ↗
                  It is not a fetish, some people like this and others like that.

                  I like CLI more than TUI, and TUI more than GUI where appropriate. For working with text TUI is better, for e.g. images GUI (GIMP).

                2. crossroadsguy · · focus · HN ↗
                  [delayed]
                  1. nsonha · · focus · HN ↗
                    You&#x27;re a software engineer living in a middle of an AI revolution and you confuse what IS with what CAN BE? GUI can be good people just didn&#x27;t do them because it took time and the foundation was shit (for the options that didn&#x27;t take time). That all changes now when a new class of software can be generated with AI. Unfortunately the people generating software and their users still base the decision on what IS (with highly technical reasoning such as &quot;infinitely better&quot;). People literally have the power to define what IS these days. I&#x27;ve only recently switched from claude and codex cli to their respective apps and those apps while not the best gui apps, already infinitely better than tui. The tui is actually hurting my main usage of them which is controlling agent programmatically. The only really point I&#x27;d give for TUI is compatibility. Running coding agent directly on android&#x2F;ios is nice sometimes.
                    1. antonvs · · focus · HN ↗
                      &gt; That all changes now when a new class of software can be generated with AI.

                      All the AI-generated UIs I&#x27;ve seen have been very derivative, certainly not eliminating any of the disadvantages of typical GUI interfaces.

                      Realistically, the whole &quot;overlapping windows&quot; GUI model, and everything that derives from that, was a metaphor geared towards people who&#x27;d never seen a computer before. It fit the increasing consumer focus of computing interfaces. It&#x27;s no wonder that technical people often prefer TUIs.

                      Maybe AI will bring real advancements in GUIs, but someone&#x27;s still going to have to make it happen.

                      1. nsonha · · focus · HN ↗
                        GUI is not just your &quot;overlapping windows&quot; straw man, it&#x27;s interactivity, plus graphical display. Yes via a taxonomy hack TUI is not GUI and there are technical hacks that bring graphics to TUIs these days, but the point remains that we need complex (but not overlapping windows, sure!) interactivity and rich display capability.

                        That is what we need, and if you&#x27;re making the argument that a terminal shell is the best place to provide them then I don&#x27;t know what to say.

                        1. antonvs · · focus · HN ↗
                          I&#x27;m pointing out that current GUIs are pretty primitive and haven&#x27;t undergone much serious thought about functional improvement since Xerox PARC in the late 1970s, and that&#x27;s why TUIs can still have an edge with technically-inclined people.

                          It&#x27;s not that GUIs are inherently worse in principle, but in practice they often are.

                          The point about overlapping windows is that that &quot;desktop&quot; model permeates the thinking about GUI design, but it&#x27;s fundamentally limiting and misguided.

                          &gt; we need complex (but not overlapping windows, sure!) interactivity and rich display capability.

                          Yep. Pity today&#x27;s GUIs can&#x27;t deliver that.

                    2. crossroadsguy · · focus · HN ↗
                      [delayed]
                      1. nsonha · · focus · HN ↗
                        This is not about vision or anything like that it&#x27;s about thinking like actual engineers and understand that what accidentally IS will always be inferior to what is designated to (can) be. TUIs will always be a hack and good by accident.
                3. agentcoops · · focus · HN ↗
                  “ 39. Re graphics: A picture is worth 10K words - but only those to describe the picture. Hardly any sets of 10K words can be adequately described with pictures.”

                  Perlis has an aphorism for this, as he does every important problem [0].

                  [0] <a href="https:&#x2F;&#x2F;www.cs.yale.edu&#x2F;homes&#x2F;perlis-alan&#x2F;quotes.html" rel="nofollow">https:&#x2F;&#x2F;www.cs.yale.edu&#x2F;homes&#x2F;perlis-alan&#x2F;quotes.html

          9. nxdmum · · focus · HN ↗
            IF you havent written your own harness - you would not understand what you can do when you&#x27;re writing your own harness . the current set of harnesses - all of them are crap tier. The only clue i can give you - it&#x27;s not in the model providers interest to have token efficiency - but when you are coding the harness yourself you can shoot for that .

            In today&#x27;s world - and idea stated stated is an idea stolen .

            1. cowl · · focus · HN ↗
              the only clue you can give... why? because it&#x27;s just words? plenty of opensource harnesses there that invalidate the conspiracy in your only clue that you can give.
            2. dmos62 · · focus · HN ↗
              What you&#x27;re saying is sus. If you have a harness that&#x27;s a tier above frontier labs&#x27; offerings, link to it.
              1. imtringued · · focus · HN ↗
                It&#x27;s not sus at all. You&#x27;re looking at it from the perspective of a general purpose system that is not adapted to your use case. Basically you prompt the model and then let the agent do everything. That&#x27;s the use case you have in mind when you think that it&#x27;s about being &quot;a tier above frontier labs&#x27; offerings&quot;.

                It&#x27;s missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.

              2. adastra22 · · focus · HN ↗
                Just about every harness is. This is common knowledge, no? Frontier lab TUI tend to steal from the OSS harnesses not the other way around.
                1. dmos62 · · focus · HN ↗
                  Care to back that up with a benchmark? I&#x27;ve not seen that. OpenCode is on here, not exactly dominating though: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents
                  1. adastra22 · · focus · HN ↗
                    Benchmarks are a terrible judge for this.
                    1. dmos62 · · focus · HN ↗
                      Why?
                      1. adastra22 · · focus · HN ↗
                        The process by which benchmarks are setup and run does not correspond at all to how human developers engage with a coding agent. At best it is a loose proxy, and often a bad one.

                        What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.

                        1. dmos62 · · focus · HN ↗
                          I presume you&#x27;re not talking about autonomous agents, right? Because then you could give it a set of tasks and check its success rate. But, even so, are you not giving it tasks? Are you principally interacting with it through discussions that are more difficult to quantify? Even question answering has benchmarks. I&#x27;m having trouble imagining how you use them (or &quot;how people use them&quot; in your words).
                          1. adastra22 · · focus · HN ↗
                            I don&#x27;t know if you noticed, but Opus 4.6 was peak for human-computer interaction. Everything has been fairly downhill from there despite better benchmarks, at least in that one regard. Opus 5 and 5.5 are clearly a step above in capabilities than 4.6, and I don&#x27;t think anyone wants to go back, but 4.7 and 4.8 were arguably worse overall. I genuinely feel I got more done with 4.6 and often switched back, prior to 5 coming out.

                            Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.

                            Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock&#x2F;swarm situation. The benchmark just doesn&#x27;t cover this. (And the difference can be nontrivial! Sakana AI&#x27;s published results show two generations of uplifting potential from better harnesses.)

            3. crossroadsguy · · focus · HN ↗
              [delayed]
              1. adastra22 · · focus · HN ↗
                Man good luck with finding that. People who wrote their own harness tend to self select into the type of people that DON’T self-author blog posts.
            4. tikkosam · · focus · HN ↗
              I think it is very much in the first party providers&#x27; interest to chase token efficiency, considering that they are offering fixed price monthly plans, and users may go to a competitor when they hit limits.
            5. edg5000 · · focus · HN ↗
              Actually the popular harnesses achieve high cache rates, but probably the syntax and prompts are more verbose than they need to be. A simpler syntax is doable, but I found with more exotic architectures I end up with lower cache rates.
          10. pjc50 · · focus · HN ↗
            &gt; I actually just cancelled Ultra also because I couldn&#x27;t subscribe to a YouTube Family plan while I had it active (Google... :[)

            This sort of thing alarms me. Having a $bigcorp account becomes a &quot;&quot;social credit&quot;&quot; system where they can ban you from all your personal stuff if they decide that you (or your agents!) are doing stuff they don&#x27;t like.

            1. w0m · · focus · HN ↗
              I take that more as the user is mixing business with pleasure. When i joined a company using `gcp` heavily, I didn&#x27;t attach my personal gmail to it - I created a &#x27;business&#x27; account and used that. Slightly inconvenient - agree, but it&#x27;s on the individual to draw the distinctions.
              1. runamok · · focus · HN ↗
                There is still anecdata that Google &quot;knows&quot; you are the same person and will ban all accounts associated with you if you do something they don&#x27;t like. And of course they never need to explain themselves or listen to your appeal.
            2. nout · · focus · HN ↗
              And that&#x27;s why it&#x27;s very powerful to be able to run AI locally, even if less smart.
        2. eloisant · · focus · HN ↗
          There is a pi plugin to use agy directly from it.
          1. mapontosevenths · · focus · HN ↗
            You get banned if they catch you.
            1. tcoff91 · · focus · HN ↗
              Yes and I&#x27;ve seen reports of it being an ENTIRE GOOGLE ACCOUNT BAN.

              I don&#x27;t want to mess with antigravity because my google account is too entrenched in my life.

              1. shmoogy · · focus · HN ↗
                That&#x27;s why it&#x27;s a nonstarter for me.
              2. 8note · · focus · HN ↗
                which makes basically any product to build with google a nonstarter.

                without having an entirely separate google account with its own separated bans, theres just no ability to trust those

                1. Sabinus · · focus · HN ↗
                  I&#x27;ve read here in previous years about bans propagating to other accounts. I think that if Google can associate you with other accounts that they&#x27;re not above banning those too.
              3. jsw97 · · focus · HN ↗
                I was getting excited but thanks for reminding me of this. Not messing with this.
              4. rdtsc · · focus · HN ↗
                Yup. I don’t plan to do anything sneaky but one wrong question or query that looks like “cyber”, say me fixing a buffer overflow in library I maintain, and all of the sudden my gmail is blocked. Yeah, not worth the risk. I feel like even with a different account they’ll figure out it’s me because well, as ad sellers that’s their business to find out who is who and I will still be banned.
        3. spankalee · · focus · HN ↗
          There&#x27;s an API: you can use Gemini with other harnesses. Isn&#x27;t the situation exactly like Claude vs Claude Code?
          1. mapontosevenths · · focus · HN ↗
            You can&#x27;t without mortgaging your home to pay enterprise API rate pricing. It&#x27;s prevented on the plans, and if you find away around it they don&#x27;t ban you from Gemini... They ban your whole-ass Google account forever.
        4. gchamonlive · · focus · HN ↗
          I&#x27;m cowboying Gemini on oh-my-pi. Been running OK so far, hopefully I won&#x27;t get banned, and if so hopefully I&#x27;ll only lose access to the models, not the storage -- while models are a sort of commodity, my data isn&#x27;t.
          1. ForHackernews · · focus · HN ↗
            Careful, if they ban your Google account, they might ban you from everything: gmail, google drive, adwords, app engine, youtube, voice, android...
            1. gchamonlive · · focus · HN ↗
              [delayed]
              1. ajolly · · focus · HN ↗
                No, Google family now shares limits across all your accounts.
                1. gchamonlive · · focus · HN ↗
                  [delayed]
      2. moecables · · focus · HN ↗
        I use Antigravity but for some reason, `agy` in the command line feels very bad&#x2F;incapable of doing things. I can&#x27;t quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call
        1. Conscat · · focus · HN ↗
          I remember having this issue ALL THE TIME with Gemini CLI but personally I haven&#x27;t experienced that yet with agy.
      3. starfallg · · focus · HN ↗
        I found 3.8 Flash in Agy to be generally better than GPT 6, and only behind from Opus 5.5.
      4. lp92 · · focus · HN ↗
        Same! 3.8 Flash does really well for writing code as long as you give it a good design and plan to follow. I use Opus for architecture&#x2F;design&#x2F;implementation plans and let Gemini 3.8 work using those. Even on the $20 Pro plan I&#x27;ve only come down to 10% before the weekly reset.
      5. mpweiher · · focus · HN ↗
        I tried antigravity a little while ago and it was utterly useless for Objective-C code, tasks that both Claude and Codex handled just fine.

        Not only could it not complete the small task, the code was obviously wrong from looking at it and did not even compile.

        When I pointed that out it got pissy and insisted the code was perfect and I didn&#x27;t know how to use a compiler, or the compiler was buggy. Pasting the compiler errors did not help.

        Surreal.

      6. moffkalast · · focus · HN ↗
        &gt; Antigravity

        Is that a reference to <a href="https:&#x2F;&#x2F;xkcd.com&#x2F;353&#x2F;" rel="nofollow">https:&#x2F;&#x2F;xkcd.com&#x2F;353&#x2F;

      7. gchamonlive · · focus · HN ↗
        [delayed]
      8. jmaker · · focus · HN ↗
        I couldn’t find a way to decline model training and reduce data retention for antigravity or Gemini. Apparently it’s only available on a business&#x2F;enterprise plan, not personal. Did you manage to solve it? That’s the only reason I don’t use Gemini or antigravity.
      9. pdntspa · · focus · HN ↗
        The gemini series have also been really strong on text extraction. I&#x27;ve been evaluating models to replace gemini 2.5 and its been hard to find something that performs as well as other gemini models
      10. noahmichael89 · · focus · HN ↗
        For frontend work, the model&#x2F;harness&#x2F;tool combo of 3.8 flash extended&#x2F;agy&#x2F;chrome dev tools MCP is shockingly good
    2. IndeanCondor · · focus · HN ↗
      Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
      1. seanthemon · · focus · HN ↗
        Gemini for day-to-day and top-of-head queries and claude for the real beefy work
    3. mapontosevenths · · focus · HN ↗
      Gemini is honestly amazing sometimes. If they didn&#x27;t force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I&#x27;m sure Google could take over the AI market.
      1. ody4242 · · focus · HN ↗
        what is so terrible with their harness? I&#x27;ve been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok.
      2. barrenko · · focus · HN ↗
        Google&#x27;s approach reminds me of what could have been the EU&#x27;s approach, they are really reluctant and drag they feet, but in the end they end up shipping and are competitive.
    4. alightsoul · · focus · HN ↗
      Please tell me you published your findings even as an issue on the llama.cpp GitHub
      1. warkdarrior · · focus · HN ↗
        Why? Anyone can run that prompt.
        1. aspect0545 · · focus · HN ↗
          Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it.
          1. FranzFerdiNaN · · focus · HN ↗
            The resources per prompt aren’t that much .

            Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.

            1. hexfish · · focus · HN ↗
              Checkmate. &#x2F;s
            2. qmr · · focus · HN ↗
              Yet you participate in a society.
              1. scarmig · · focus · HN ↗
                If action X takes a million times more resources than action Y, it&#x27;s silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn&#x27;t somehow erase that or make it irrelevant.
                1. qmr · · focus · HN ↗
                  It gets me 10 internet points though.
            3. articulatepang · · focus · HN ↗
              All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them.

              Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.

          2. dzhiurgis · · focus · HN ↗
            Took me 6M codex tokens yesterday to get omarchy wifi working on macbook
            1. alightsoul · · focus · HN ↗
              That is not normal. Were you able to use the arch wiki? Omarchy is a version of arch Linux.
              1. dzhiurgis · · focus · HN ↗
                IDK how it fixed it exactly, I suspect it&#x27;s macbook specific issue. It wasn&#x27;t accepting my wifi password and 5ghz radio wasn&#x27;t working. Steering it to fix 5ghz sorted it out. I&#x27;m not going to any wiki&#x27;s myself, fuck that.
          3. AuthAuth · · focus · HN ↗
            They dont understand the fix so sharing a vibe coded patch is upstream spam. They should submit a bug report and their findings. (Written by them not AI)
        2. luckydata · · focus · HN ↗
          why reinvent the wheel and spend tokens for a problem that has already been solved?
        3. baby_souffle · · focus · HN ↗
          Wouldn&#x27;t it be better if only one person had to and then we all got to benefit from the fix?

          Why would the guy who wrote curl share it? We can all build our own now...

          Why do the Linux folks need to be so selfless? We can all build our own kernel now...

        4. folkrav · · focus · HN ↗
          Why would we even ever distribute software again, by this logic?
      2. otabdeveloper4 · · focus · HN ↗
        Spoiler alert: the problem didn&#x27;t actually get fixed despite the jaw on the floor.
      3. dominotw · · focus · HN ↗
        he is still closing his jaw
      4. taylorfinley · · focus · HN ↗
        Here they are: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...
      5. imtringued · · focus · HN ↗
        If you read the code you would realize that the issue is....

        in amdkfd and hsakmt

        Yes, that means AMD is sitting on both sides. They wrote software that doesn&#x27;t work with their own software.

    5. bel8 · · focus · HN ↗
      I had a similar but less impressive experience recently with Muse Spark 1.3.

      Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.

      It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.

      1. seanthemon · · focus · HN ↗
        Godot encryption is laughably easy to break, there&#x27;s tons of packages available for it. It&#x27;s a well known drawback of using godot
        1. warkdarrior · · focus · HN ↗
          Now the LLMs know it too.
        2. bel8 · · focus · HN ↗
          I didn&#x27;t say it wasn&#x27;t.

          Still impressive that it did so much just to answer my simple question.

    6. amanguliani · · focus · HN ↗
      Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
      1. onlyrealcuzzo · · focus · HN ↗
        Sol 6.1 is quite good, but damn is it slow.

        I&#x27;m using it to run overnight tasks, and that&#x27;s it until my quota runs out.

        Canceled my subscription.

        1. amanguliani · · focus · HN ↗
          WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what&#x27;s next. Did the same for CODEX - nope still keeps effing going on and on and on
          1. tourist2d · · focus · HN ↗

            [dead]

          2. 8note · · focus · HN ↗
            astra afaict does two stage commits for everything. the first response is a plan, and the second is actually doing it.

            its a lot less chatty imo

          3. awakeasleep · · focus · HN ↗
            If you dig in the settings you can control that. Like on a remote ssh connection, in the settings, you can pick “friendly or terse”

            In the local app interface the winning choice is “efficient” and then turn off the sliders for warmth enthusiasm emoji etc.

            It makes openai models so good to talk to i really have trouble switching.

        2. unconscionable · · focus · HN ↗
          I find Opus 5.5 is better at giving a high level adversarial &quot;should you do this to begin with&quot; where GPT-6.1 is happy to go down any wrong path.

          Also canceled my ChatGPT Pro $200&#x2F;mo subscription. Their Oct 30 price hikes and slow GPT-6.1 model has me looking for alternatives.

          1. directdev · · focus · HN ↗

            [dead]

    7. gottorf · · focus · HN ↗
      My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I&#x27;m not using it for coding, but general research on different topics.
      1. staticman2 · · focus · HN ↗
        The web version of Gemini is awful at search but I don&#x27;t think that&#x27;s the models fault.
      2. MILP · · focus · HN ↗
        I&#x27;m also not using it for coding but I&#x27;ve found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
        1. robobo96 · · focus · HN ↗
          Only html or also css? Opus seems a bit more creative than most other models i&#x27;ve seen.
      3. mattjoyce · · focus · HN ↗
        Hallucination seems a very dated term.
        1. nkozyra · · focus · HN ↗
          Why? It&#x27;s the same concept and root cause it was when we first started using it.
        2. xdavidliu · · focus · HN ↗
          there are many dated expressions, including

          - AI is just a tool, like excel; it does what the human operating it tells it to

          - next token prediction cannot be true understanding

          - models can have no desires and goals, don&#x27;t anthropomorphize it

          However, &quot;hallucination&quot; is very much not one of them

        3. gottorf · · focus · HN ↗
          Hallucination is accurate for what I&#x27;m seeing -- e.g. it&#x27;s making up information about the 2nd gen Toyota Tundra that has no basis in reality. When challenged, it corrects itself.
          1. alluro2 · · focus · HN ↗
            My colleague wanted to diagnose a specific error code on his car himself, and Gemini told him that it&#x27;s simple to do with an OBD2 dongle - he asked it about the details thoroughly, to confirm, and bought the dongle.

            It didn&#x27;t work. Gemini: &quot;Oh yeah, that obviously cannot work, it&#x27;s not possible to do it through OBD2&quot; (paraphrasing)

            It was quite funny to me, but a bit less so to my colleague.

            1. Gareth321 · · focus · HN ↗
              I&#x27;ve had many similar experiences. It&#x27;s confidently incorrect to a shocking degree. Worse than ChatGPT from two years ago.
          2. mattjoyce · · focus · HN ↗
            Its always been a bad term. If we wanted an accurate term then it&#x27;s &#x27;confabulation&#x27;, but &#x27;muddled&#x27; or just &#x27;wrong&#x27; are also good.
        4. rdtsc · · focus · HN ↗
          What do we use for the “model made stuff up and claimed it as facts”? I can see hallucinations somehow anthropomorphizing LLM even more. I don’t like that we’re doing that to begin with but it’s a losing battle. I prefer “it’s broken” and “IT produced shit results” personally.
          1. mattjoyce · · focus · HN ↗
            I agree with you, except I really don&#x27;t hear that term much. People just say it wrong or confused. Good riddance, it was always a bad term.
          2. krapp · · focus · HN ↗
            Models don&#x27;t make claims. That would require a degree of interiority and intent that they don&#x27;t have.

            The bigger problem is that people expect LLMs to know what facts are. That assumption is even baked into the term &quot;hallucination.&quot; Someone who hallucinates is expected to otherwise have a grounding in objective reality, to &quot;not&quot; hallucinate, and to be able to recognize reality from fantasy. We wouldn&#x27;t allow a person who &quot;hallucinates&quot; as much as an LLM anywhere near the roles we give to LLMs. But everything an LLM does is as much a &quot;hallucination&quot; as anything else, it&#x27;s just stochastically generating grammar. Some grammar just happens to be useful because of the quality of its training data, which was probably created by humans who do possess interiority and awareness of fact.

            And it isn&#x27;t &quot;broken&quot; either. Broken assumes that the correct mode of operation is to act as a source of truth or fact generation. When LLMs &quot;apologize&quot; for bad results, for instance they aren&#x27;t actually apologizing. Try getting it to apologize for returning the correct data. It probably will. There is no cognition happening. It doesn&#x27;t know either way. It isn&#x27;t a calculator crunching numbers or a computer doing data analysis. It&#x27;s just pattern matching.

            &quot;Hallucination&quot; is no less correct than &quot;confabulation&quot; which also presupposes intent and contextual awareness. Unfortunately the way LLMs operate is so unintuitive (as opposed to the intuitive nature of the interface) that the only language we have to describe it is the language of human behavior, with all of the biases and false assumptions that brings.

        5. UpsideDownRide · · focus · HN ↗
          They still happen.
      4. WarmWash · · focus · HN ↗
        The achilles heel of 3.8 flash is it&#x27;s january 2025 knowledge cutoff date. Yes, almost 2 years ago.

        I&#x27;m assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.

        1. blinding-streak · · focus · HN ↗
          Incorrect (to some degree)

          &gt; The knowledge cutoff date for Gemini 3.8 Flash is March 2026

          <a href="https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash&#x2F;" rel="nofollow">https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash&#x2F;

          1. WarmWash · · focus · HN ↗
            &gt;The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family). For more information about known limitations, see the Gemini 3.7 Flash

            The &quot;some domains&quot; are very narrow. They likely just RL&#x27;ed popular queries.

        2. Gareth321 · · focus · HN ↗
          While that&#x27;s annoying, the other frontier models easily overcome this with appropriate tool usage. I do a lot of research with frontier models and they&#x27;re very good about identifying where their parametric knowledge is insufficient and searching for the correct knowledge on the internet. 3.8 Flash is HORRIFIC. The majority of the time it doesn&#x27;t use any tools and infers things from its parametric knowledge. Things which should clearly have implied tool calls. Historical statistics, legal precedent, economic data, etc. I think it&#x27;s incredibly clear that it has been tuned for speed and not accuracy.

          Of course, it&#x27;s called &quot;flash,&quot; and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I&#x27;m hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let&#x27;s see.

      5. esafak · · focus · HN ↗
        I would expect a flash model, with its reduced size, to suffer on tail tasks. That is the trade-off you make.
      6. Gareth321 · · focus · HN ↗
        I agree. It&#x27;s much worse than the cheap Chinese models. They appear to heavily bias parametric knowledge and discourage tool use. That&#x27;s fine for things like &quot;how do I perform CPR?&quot; but worse than useless for any kind of research. [There is one benchmark showing far lower rates of hallucination, so let&#x27;s see how accurate this is.](<a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;singularity&#x2F;comments&#x2F;1wuj72j&#x2F;gemini_4_argon_solved_hallucinations&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;singularity&#x2F;comments&#x2F;1wuj72j&#x2F;gemini...)
    8. yegle · · focus · HN ↗
      For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch&#x2F;build it properly.

      And I saw it do this twice, once for Android 14 and once for Android 16.

      I think this is just within 3.8 flash&#x27;s capabilities.

      1. p_l · · focus · HN ↗
        3.8 Flash (but also last two ones) have really strong preference for dissecting binaries with quick thrown-together bits of python in my experience.

        Including going first for decompiling AGY binary instead of searching the web for documentation...

        1. IshKebab · · focus · HN ↗
          Astra also really loves reverse engineering binaries. I guess it&#x27;s one of those things that isn&#x27;t that complicated but is super tedious, and tedium means nothing to AI.
    9. illwrks · · focus · HN ↗
      I&#x27;ve been tinkering with Gemini for several months and I think it&#x27;s great. The most complex things I&#x27;ve had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
    10. martythemaniak · · focus · HN ↗
      Adding my anecdote, because it amused me: I finished wiring up the compute&#x2F;sensor box for my robot, ssh&#x27;d in and told agy &quot;I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output&quot;. It did all the local config for the lidar, setup docker and the ROS2 environment, then told me &quot;open up this url in foxglove&quot; and sure enough everything worked. Whole thing used up 6% of my weekly limit.
    11. Grimburger · · focus · HN ↗
      &gt; I was trying to use rocm with llama.cpp

      completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.

      once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B

      I get the feeling the situation is only going to improve longer term so might be a good time to just do it

      1. esseph · · focus · HN ↗
        [delayed]
      2. clw8 · · focus · HN ↗
        Support has gotten much better in just the last couple months. I just got a 9070 XT and can&#x27;t count the number of times I&#x27;ve installed a package and the changelog made me think how much it would have sucked to be doing this a year ago.
      3. nzeid · · focus · HN ↗
        I have so much to share on this topic. Will keep it short.

        ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I&#x27;ve been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.

        The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn&#x27;t materialize and the token generation speed decreases.

        There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.

    12. gcy · · focus · HN ↗
      I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it&#x27;s not as capable but it&#x27;s fast and almost free (sufficient quota with pro account). The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.
    13. cyanydeez · · focus · HN ↗
      if you haven&#x27;t tried Qwen3.8-Flash-Next with halogen, you&#x27;re missing out: <a href="https:&#x2F;&#x2F;github.com&#x2F;peonist-ai&#x2F;halogen-flash-server#the-host-settings-these-numbers-were-measured-on" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;peonist-ai&#x2F;halogen-flash-server#the-host-...
      1. solaire_oa · · focus · HN ↗
        I coincidentally just installed this, and gud dayum, it&#x27;s pretty awesome.

        I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it&#x27;s slop-riddled hallmarks.

        In any case, yeah, ~55 tok&#x2F;s on a high quality model, the mind reels at what I might be able to do without constantly beancounting token ratelimits. And it&#x27;s a huge win for privacy as well.

        1. cyanydeez · · focus · HN ↗
          Yeah, unfortunately, it does deliver for its platform.
    14. lardo · · focus · HN ↗
      Did rocm provide any benefit over vulkan?
      1. taylorfinley · · focus · HN ↗
        Vulkan was reliably crashing after a certain point in the context window, Rocm has been stable as a rock. Tps was basically a wash.
    15. danpalmer · · focus · HN ↗
      3.8 Flash is my daily driver and produces pretty excellent results all round.
      1. dcl · · focus · HN ↗
        What harness you use? Have you tested more than 1?
        1. danpalmer · · focus · HN ↗
          This experience is with Antigravity both internally and externally, and I have done quite a few side-by-side comparisons with the same prompt across a number of different Google and non-Google models.

          I&#x27;ve tried Codex as a harness too, and that was nice. I don&#x27;t find a significant difference between Antigravity and Codex. Codex has more features but I don&#x27;t use them.

    16. iknowstuff · · focus · HN ↗
      lol I&#x27;ve hit the hipStreamCreate problem!
    17. la6479 · · focus · HN ↗

      [dead]

    18. plasticchris · · focus · HN ↗
      This is why I think llms are a killer app for Linux desktop. They’ve been trained on Linux very hard, and it cleanly solves the “how do I make it do $thing” problem since everything is open and the llm can manipulate it. For example: Sound not working? Just tell the llm.
      1. _heimdall · · focus · HN ↗
        Is that specific to Linux? For better or worse, I have felt like it&#x27;d fairly good at working through most tech stacks I throw at it.

        I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.

        1. jubilanti · · focus · HN ↗
          &gt; Is that specific to Linux?

          Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too.

        2. x-complexity · · focus · HN ↗
          Arguably, it requires the LLM to have access to the source code during training. If it&#x27;s closed off (like Windows), then the LLM can&#x27;t train on it.

          Even if the source code&#x27;s old, the fact that it is publicly available makes it much easier to train &amp; improve on than if it were walled off.

          1. augusto-moura · · focus · HN ↗
            Not only training, you can literally clone the source code locally and ask the agent to debug your problem by grepping the source code itself. Even if doesn&#x27;t &quot;know&quot; the answer it can investigate it on the fly
        3. matsemann · · focus · HN ↗
          It&#x27;s at least specific to open source, in my experience. Ask it about some commercial software: it searches some lacking documentation, social media discussions and make some guesses. Ask it about an open source tool: it searches the documentation, if not it dives into the implementation to figure out the edge case.
        4. Sammi · · focus · HN ↗
          The more open the operating system the more LLMs can help you. This is what Arch Linux was waiting for all along.

          I used to be afraid of Arch, because I don&#x27;t want a system that takes work because I&#x27;m already busy with work. But now I love it, because the LLM can tweak every knob and fix every issue for me, so it ends up being the OS that takes the least work to use. Get an error message? Tell the LLM and they fix it. Something not working exactly like you like it? Tell the LLM and they tweak it for you. This also works to great effect on Win and Mac, but not to the same extreme degree as it does on Linux and especially Arch.

          1. lukan · · focus · HN ↗
            Hell yes, I also started to tweak my XFCE arch desktop to my liking.

            Anything I want different now, I just tell the LLM to do it for me.

            (For example I can now close lots of windows of the same type with 1 click, not 3, my whisker search now finds files and folders and I am able to run games that refused before)

            I still ocasionally run into the usual linux driver issues, but not for much longer I suppose. I probably could fix some driver bugs now already if I point fable towards it and pay some attention.

            In theory I could do all this before myself, but not just like that in some minutes, but in days&#x2F;weeks&#x2F;months ..

            1. Sammi · · focus · HN ↗
              Everyone that didn&#x27;t have time to be a neck-beard themselves now has a neck-beard built right into their own computer ready to go.
      2. masto · · focus · HN ↗
        I’ve been playing with the ITS operating system on a simulated PDP-10. Despite it being from the 1960s, Claude has been able to do sysadmin work on it, including a very similar journey where it started disassembling code to identify the source of an obscure problem. I don’t want to take away from the learning experience (the whole reason I’m interested in this stuff), but it’s often almost impossible to find what I’m looking for online, and the LLM turns out to be a better resource when I want to learn how a particular subsystem works. I ended up making a custom MCP server so it could connect to the simulator more efficiently. <a href="https:&#x2F;&#x2F;github.com&#x2F;masto&#x2F;pidp10-mcp" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;masto&#x2F;pidp10-mcp

        Drifted off my point a bit, which I guess was meant to be it’s not necessarily Linux-specific training.

      3. c0n5pir4cy · · focus · HN ↗
        I got Claude to set up 5.1 Surround Streaming from my Linux Desktop to my Steam Deck using Moonlight. It needed some help (or really human ears) to make sure things were coming from the right speakers - but it was pretty much a one-shot.
      4. chinathrow · · focus · HN ↗
        My laptop resumed from sleep all too often and Claude found and fixed it within minutes. Now I still need to file that bug report though...
      5. oldandboring · · focus · HN ↗
        I used Claude to build an entire runbook for how to build my Linux desktop up from scratch to its current state, including all my installed software, configs, hacks, workarounds, etc. Then I used it to migrate from one laptop to another in an afternoon. Now it keeps itself up to date.
      6. safog · · focus · HN ↗
        Yes this might be a bit too noob linux but I have the AIs maintaining an env.md + logs &#x2F; updates on that md whenever they change anything about the env. Even things as simple as getting tmux to be fast, they do an excellent job at.

        Drivers, Coding environment setup etc. are great too and it&#x27;s nice to have everything logged so the next (more powerful) agent can come and improve the thing once in a while.

      7. UpsideDownRide · · focus · HN ↗
        Fully agree. There is so much customization that I&#x27;m doing thanks to LLMs that I just wouldnt to bother to do by myself.
      8. stefan_ · · focus · HN ↗
        Nice idea, then they want to make a quick little kernel module, oops, you secure booted, this is not your system, can&#x27;t load...

        Since LLMs have been so successful at finding exploits it&#x27;s been clearer than ever that the people so obsessed with enshittifying every system with ineffective (other than pissing you off and wasting your time) &quot;secure boot&quot; functionality were really just too ignorant to succeed at actual novel security work, so they focused on this make believe crap. Well, glad that&#x27;s over.

      9. samspot · · focus · HN ↗
        I used Claude &amp; Gemini extensively to troubleshoot Linux since I switched last year, and it&#x27;s been mostly great. But then it nearly bricked my install last week and I had to get some human help. I was getting black screens during and after boot and tried many suggestions from multiple AI&#x27;s. After many hours it turns out I just needed to power cycle my monitor. Shout out to Claude for suggesting that (its memory remembered that I have a dock built in to my primary display).

        I am still very happy with my switch to Linux. But if I didn&#x27;t have the AI help I would say linux is still unacceptable platform for those not willing, able, and excited to get their hands very dirty.

      10. tarokun-io · · focus · HN ↗
        Agreed 100%! I got some Steam games that didn&#x27;t run well (or at all) to run smoothly thanks to AI, after having spent hours trying to figure it out on my own (months prior), fix some issues with HiDPI and external monitor with my laptop, and now I have a tiny shell script that basically runs `checkupdates` (`checkupdates | claude -p`, roughly) and looks online for potential issues. I&#x27;ve been using Linux for many years but still Claude and ChatGPT helped me a ton.
    19. mirmor23 · · focus · HN ↗
      &gt; it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo.

      Most llm could do it. Claude went from firmware thread -&gt; rtos scheduler -&gt; mcu reference manual -&gt; hardware controller register interface -&gt; vendor sdk -&gt; problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.

    20. d3Xt3r · · focus · HN ↗
      That&#x27;s cool, but the real question is, why aren&#x27;t you using Hipfire[1] or HaloPFX[2]? Both are far superior to llama.cpp in terms of performance, for Strix Halo.

      [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;warpfront&#x2F;hipfire" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;warpfront&#x2F;hipfire

      [2] <a href="https:&#x2F;&#x2F;github.com&#x2F;julianmb&#x2F;halofpx" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;julianmb&#x2F;halofpx

    21. qrify_app · · focus · HN ↗

      [dead]

    22. vinzenzu · · focus · HN ↗
      I would like to use my Google AI Pro subscription included with Drive but last time I checked, their terms were not just ambiguous and confusing wrt training and ZDR, they were contradictory. Until they have that fixed and properly communicated, I&#x27;ll stick to other vendors.
    23. crossroadsguy · · focus · HN ↗
      [delayed]
      1. p_l · · focus · HN ↗
        3.5 Pro has been pretty much canceled for few months now, 4.0 was in the works instead
    24. KMnO4 · · focus · HN ↗
      My theory is that Gemini 3.8 Flash was supposed to be Gemini Pro, but by the time it was ready for release, Google was embarrassed at how far behind their &quot;pro&quot; model was, so they just named it Flash. It&#x27;s a good model, but don&#x27;t let the name trick you.
      1. w0m · · focus · HN ↗
        I&#x27;d agree if it wasn&#x27;t so much faster than ~all the other models. Maybe they nerfed reasoning to get that speed.
    25. phmx · · focus · HN ↗
      ``` void* mmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) { if (!real_mmap) real_mmap = dlsym(RTLD_NEXT, &quot;mmap&quot;); ```

      hope it&#x27;s not run by multiple threads and dlsym is not allocating.

  21. dom96 · · focus · HN ↗
    Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?

    None of the other AI labs do this. Really frustrating.

    1. gengelbro · · focus · HN ↗
      Mythos?
      1. aqsnow · · focus · HN ↗
        Yes but google always does this crap.
  22. netdur · · focus · HN ↗
    I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
    1. tom1337 · · focus · HN ↗
      Are you enrolled in Fairwind?

      &gt; Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.

  23. gravisultra · · focus · HN ↗
    Google has the audacity to &quot;protect us from ourselves&quot; and talk about &quot;safety&quot; and in the very same blog post highlight the Israeli &quot;security&quot; company Wiz, that they acquired for a very exaggerated sum of money.

    This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.

  24. bananaflag · · focus · HN ↗
    I wonder how it will be at solving open math problems.
  25. osiris970 · · focus · HN ↗
    Hopefully their harnesses aren&#x27;t unusable when they release this
  26. retropragma · · focus · HN ↗
    no Pareto frontier graph?
    1. xnx · · focus · HN ↗
      <a href="https:&#x2F;&#x2F;x.com&#x2F;arena&#x2F;status&#x2F;2105394858521469177" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;arena&#x2F;status&#x2F;2105394858521469177
  27. helsinkiandrew · · focus · HN ↗
    &gt; Google Grapples With Employee Skepticism About New Gemini Model

    <a href="https:&#x2F;&#x2F;www.bloomberg.com&#x2F;news&#x2F;articles&#x2F;2026-09-30&#x2F;google-grapples-with-employee-skepticism-about-new-gemini-model" rel="nofollow">https:&#x2F;&#x2F;www.bloomberg.com&#x2F;news&#x2F;articles&#x2F;2026-09-30&#x2F;google-gr...

    1. bitexploder · · focus · HN ↗
      Opinions my own but I have been using this model for a bit. I would say it is a good model and the skill with which people use AI varies widely.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.