‹ BackHN Continuity

Thread

Gemini 4 Argon (High): Intelligence, Performance and Price Analysis

112 points · 61 comments · theanonymousone

  1. aliljet · · focus · HN ↗
    It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
    1. scrollop · · focus · HN ↗
      Here are sites with ongoing measurements checking if a model has been nerfed - check opus 5.5-

      <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

      <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F; <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

      <a href="https:&#x2F;&#x2F;github.com&#x2F;ninjahawk&#x2F;livenerf" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ninjahawk&#x2F;livenerf

      <a href="https:&#x2F;&#x2F;isitnerfed.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;isitnerfed.org&#x2F;

      1. ozgung · · focus · HN ↗
        I wonder how they do this nerfing thing. One candidate is slightly decreasing the number of chain of thought tokens for each effort level. It must be something they do to meet increasing demand. Also the significant drop before a new release is because of reallocating the resources.

        Methodology section in some of these benchmarks doesn’t say if they use subscription or API. API usage may not be nerfed as much as subscription.

        They must be using the same lever to “pace the frontier”. All of the best effort models from different companies have similar scores. There is no standard definition of “max” effort level.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.