‹ BackHN Continuity

Thread

Gemini 4 Argon (High): Intelligence, Performance and Price Analysis

112 points · 61 comments · theanonymousone

  1. aliljet · · focus · HN ↗
    It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
    1. zozbot234 · · focus · HN ↗
      It's not as smart as Claude Opus 5.5 High according to the AA benchmark. Looks like a big fat nothingburger so far, though it's possible that future fine-tuned checkpoints of the same pretrained model will do a lot better.
      1. mpyne · · focus · HN ↗
        > It's not as smart as Claude Opus 5.5 High according to the AA benchmark.

        If it's smart enough to do the job then it won't matter that Opus is smarter. At the right price and performance, at least.

        1. petesergeant · · focus · HN ↗
          While that’s true, I have found Gemini models to be exclusively good for data extraction, and absolutely terrible at everything else.

          I pay for lots of models because they’re good at different things: $20 a month each for Grok and GLM have easily paid for themselves by finding bugs that my main work models didn’t, but I’m yet to have any Gemini model find a real bug, and Gemini’s results for general work will sometimes border malicious compliance, when it’s not having a hissy fit about some imagined issue.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.