‹ BackHN Continuity

Thread

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

333 points · 106 comments · theanonymousone

  1. breckenedge · · focus · HN ↗
    Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.
    1. mnicky · · focus · HN ↗
      Well there is at least the degradation tracker from Margin labs for Sol and Opus: <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;
      1. therealdrag0 · · focus · HN ↗
        It looks pretty consistent? At least within reason for a stochastic model. Or am I missing something?
        1. mnicky · · focus · HN ↗
          What you may be missing is that they probably track the API performance.

          Subscription plans may be subject to other regime, e.g. lowering the thinking budget when the API is under heavy load, etc.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.