‹ BackHN Continuity

Thread

Gemini 4 Argon (High): Intelligence, Performance and Price Analysis

112 points · 61 comments · theanonymousone

  1. aliljet · · focus · HN ↗
    It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
    1. tomrod · · focus · HN ↗
      They own their hardware. That vertical integration alone probably saves oodles because they can reconfigure to their needs as opposed to individually negotiating data centre plans. I can't imagine the complexity both OpenAI and Anthropic have to maintain for their deployments.
      1. mkotlikov · · focus · HN ↗
        It's still not enough though, they're buying compute from SpaceX's colossus data centers.
        1. tomrod · · focus · HN ↗
          They may be buying it, but is there any indication they are using it for consumer-facing stuff? I would suspect given the scale there are many different things. I wish I knew more about standardization for cloud data center workloads though
        2. notatoad · · focus · HN ↗
          i don't know if you could say they're buying it "happily". more like begrudgingly.

          the deal has a 30-day cancellation policy, and they raised a bunch of debt around the same time to fund their own datacenter expansion.

          1. fragmede · · focus · HN ↗
            and they shut down Stadia so they'd have some extra GPUs to use for it as well!
        3. genxy · · focus · HN ↗
          That is google trying to keep their 10% worth something.
          1. largbae · · focus · HN ↗
            And it totally worked too.
    2. zozbot234 · · focus · HN ↗
      It's not as smart as Claude Opus 5.5 High according to the AA benchmark. Looks like a big fat nothingburger so far, though it's possible that future fine-tuned checkpoints of the same pretrained model will do a lot better.
      1. mpyne · · focus · HN ↗
        > It's not as smart as Claude Opus 5.5 High according to the AA benchmark.

        If it's smart enough to do the job then it won't matter that Opus is smarter. At the right price and performance, at least.

        1. petesergeant · · focus · HN ↗
          While that’s true, I have found Gemini models to be exclusively good for data extraction, and absolutely terrible at everything else.

          I pay for lots of models because they’re good at different things: $20 a month each for Grok and GLM have easily paid for themselves by finding bugs that my main work models didn’t, but I’m yet to have any Gemini model find a real bug, and Gemini’s results for general work will sometimes border malicious compliance, when it’s not having a hissy fit about some imagined issue.

        2. jug · · focus · HN ↗
          I checked price per AA task and yeah Argon vs Opus 5.5 High are practically neck and neck both on "intelligence" and cost per task. So if you aren't pushing higher on Opus (which I think can quickly become very expensive and I always treat Max and those like "benchmark settings"), I think this becomes more of a matter of which platform you like more or prefer for various reasons.

          Honestly, I think this will be a trend in 2027 when all those models become "good enough" for elite coding and whatever. I predict they'll have to branch out more and build their platform to differentiate themselves from each other, maybe even in terms of branding and trust, marketing towards various demographies, youth vs elderly, students vs employees, etc.

    3. aleqs · · focus · HN ↗
      > Opus 5.5 right now because it's nearly unlimited use

      Anthropic has some of the lowest usage per $ in general, not sure what you're taking about.

      1. jjice · · focus · HN ↗
        I don't think the OP and your comments are mutually exclusive. If Anthropic usage is lower, and the OP considers it basically unlimited for their case, that just means that they would have virtually unlimited usage with other plans.
        1. aleqs · · focus · HN ↗
          Nothing about it is unlimited or even close to it. Just because you only use 1gb of your 2gb data plan, doesn't mean you have unlimited data.
          1. 8n4vidtmkvmk · · focus · HN ↗
            What if you use 100GB of your 10TB plan? Is it virtually unlimited then?

            It's all relative. Some people just can't use up their quotas with their normal usage.

            1. aleqs · · focus · HN ↗
              That just means your use is limited not that the plan is unlimited.
      2. UltraSane · · focus · HN ↗
        Opus 5.5. on medium effort is very good a writing code and provides a lot of tokens per 5 hour session. It is a fantastic value for $20/month.
      3. nl · · focus · HN ↗
        This is no longer the case.

        Opus and Sol usage levels vs the API are currently roughly the same, but Opus 5.5 outperforms at low and medium effort levels.

    4. sroussey · · focus · HN ↗
      GPT-6.1-sol costs less than half of Gemini and way less than Anthropic on the cost per task chart of the listed parent page.
      1. Melatonic · · focus · HN ↗
        Suspect they will release Gemini 4 flash not too long from now as the cheaper option
    5. nl · · focus · HN ↗
      This is completely the wrong read!

      Sol 6.1 scores one point less than Gemini 4 on intelligence AND costs less than half ($0.72 vs $1.99) per task.

      Additionally, if you are using OpenAI you have the option to pay a bit more and get Astra which - despite the benchmarks - does outperform Sol on some things.

      Also, people are - rightly - very wary of Google's benchmaxxing tendencies. I think lots of people remember Gemini 3.0 (I think?) which benchmarked amazingly, but as soon as you used it would go off-track and needed constant babysitting if you wanted to use it for agentic work.

      1. re-thc · · focus · HN ↗
        > Sol 6.1 scores one point less than Gemini 4 on intelligence AND costs less than half ($0.72 vs $1.99) per task.

        Via the API. The $200 OpenAI plan just got cut and most say general quotas got cut before that so for users on a plan the numbers might be different.

        1. nl · · focus · HN ↗
          Yes well obviously it's comparing API vs API and subscriptions are much cheaper on both.
    6. qrify_app · · focus · HN ↗

      [dead]

    7. scrollop · · focus · HN ↗
      Here are sites with ongoing measurements checking if a model has been nerfed - check opus 5.5-

      <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

      <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F; <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

      <a href="https:&#x2F;&#x2F;github.com&#x2F;ninjahawk&#x2F;livenerf" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ninjahawk&#x2F;livenerf

      <a href="https:&#x2F;&#x2F;isitnerfed.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;isitnerfed.org&#x2F;

      1. ozgung · · focus · HN ↗
        I wonder how they do this nerfing thing. One candidate is slightly decreasing the number of chain of thought tokens for each effort level. It must be something they do to meet increasing demand. Also the significant drop before a new release is because of reallocating the resources.

        Methodology section in some of these benchmarks doesn’t say if they use subscription or API. API usage may not be nerfed as much as subscription.

        They must be using the same lever to “pace the frontier”. All of the best effort models from different companies have similar scores. There is no standard definition of “max” effort level.

    8. alvarolucero · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.