‹ BackHN Continuity

Thread

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

333 points · 106 comments · theanonymousone

  1. qsort · · focus · HN ↗
    I am begging you on my knees to please stop posting this cringe.

    The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?

    1. Someone1234 · · focus · HN ↗
      You forgot to include whatever you're proposing instead.

      "Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.

      Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.

      1. svachalek · · focus · HN ↗
        Yeah unfortunately the benchmarks are usually provided by the company themselves, unquantized, thinking set to extra-extra-ultra-high, best of 10 runs, etc etc. It's hard to know how that's going to map to real world users.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.