‹ BackHN Continuity

Thread

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

333 points · 106 comments · theanonymousone

  1. breckenedge · · focus · HN ↗
    Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.
    1. tedsanders · · focus · HN ↗
      Can you share more information on your methodology?

      GPT-5.6 Sol's performance in the API should not change over time. If it has, that's a severe bug and we'll look into it.

      We do sometimes tweak ChatGPT settings (e.g., tools, system prompts, efforts) over time, but we never play games to juice evals at launch times. You should always get what's advertised.

      (I work at OpenAI.)

      1. breckenedge · · focus · HN ↗
        This is all code review runs via OpenRouter with a Pi harness, and it’s totally possible there are shenanigans going on elsewhere.

        Yesterday, I ran an identical bug identification dataset from two weeks ago, saw a 50% drop from a few weeks ago, putting Sol on the same level as Luna. Sol had been finding 40-50 bugs per set, then dropped to 25, matching Luna’s performance. Not enough to establish a pattern, but enough to raise eyebrows.

        Our review workflow is public if you want to peruse it, dataset isn’t. The process isn’t really stabilized yet either as I have to balance running this against limited budgets.

        <a href="https:&#x2F;&#x2F;github.com&#x2F;BiggerPockets&#x2F;.github&#x2F;blob&#x2F;main&#x2F;.github&#x2F;workflows&#x2F;biggiepockets-review.yml" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;BiggerPockets&#x2F;.github&#x2F;blob&#x2F;main&#x2F;.github&#x2F;w...

        1. tedsanders · · focus · HN ↗
          How many tasks were in this dataset?

          If it&#x27;s a single task where it dropped from 50 to 25, it could be random variation (not saying it is, but it could be). If it&#x27;s the mean over hundreds of tasks, that suggests a problem with either the eval code&#x2F;harness or our API.

          1. breckenedge · · focus · HN ↗
            Just about 90 PRs in that dataset to review.
            1. tedsanders · · focus · HN ↗
              Is it 90 independent tasks, or a single task with 90 pieces?

              (it matters if they are independent or dependent)

      2. doctorpangloss · · focus · HN ↗
        ArtificialAnalysis tweaks stuff until newest big proprietary model is on top, not you haha
      3. dr_kiszonka · · focus · HN ↗
        But the chat version does change over time, correct? It has been my experience that Sol&#x27;s performance has deteriorated significantly.
      4. sunaurus · · focus · HN ↗
        Do you have any hypothesis for why this experience is so consistently reported by users (anecdotally)? Seemingly across all providers.
        1. billypilgrim · · focus · HN ↗
          Not the one you asked but I genuinely think it’s possible that a part of the answer is the psychological effect of getting used to models performing well and picking up more if they fail, but also … If you have a product that works well for 95% of software engineering tasks (for the sake of this argument), with a large number of users there will inevitably be some poor schmucks who get bad results multiple times in a row. It would be highly unlikely if that _didn’t_ happen at all. Now assuming you have millions of users, there will always be groups of thousands that experience this, and if those people go online to complain, it will look like an actual issue when it can just be explained by randomness and large numbers.
        2. tedsanders · · focus · HN ↗
          My best guess is random variation and rising expectations. I once saw someone do in depth manual testing of 3 different models and write up a whole report on their perceived strengths and weaknesses… and then discover that all 3 models were identical. Very good reminder of why the scientific method is important.

          Serial testing over time is much less reliable than side by side testing, and even when I do side by side testing, I try to look at multiple attempts per prompt. Seeing multiple per prompt helps me realize how much intrinsic variation there is. My brain always wants to see patterns even when there isn’t enough data to prove them.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.