‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. eli · · focus · HN ↗
      Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
      1. abejora · · focus · HN ↗
        You're right about its real world performance, and I worded my original comment wrongly.

        I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.

        1. joeyhage · · focus · HN ↗
          Claude, is that you?
          1. bb-connor · · focus · HN ↗
            you're absolutely right to push back
        2. swiftcoder · · focus · HN ↗

          [dead]

          1. verdverm · · focus · HN ↗
            this is your brain on drugs...

            this is your brain on claude...

            corporate needs you to find the difference

        3. ramon156 · · focus · HN ↗
          Your clarification makes sense. The distinction between overall benchmark performance and why Terminal-Bench is an outlier is important
      2. chis · · focus · HN ↗
        Well presumably now it’ll fall back to Sonnet 5.5 lol
      3. spider-mario · · focus · HN ↗
        You can read “a little bit” (i.e. “not too much”) into it (it does indeed tell you about the out-of-the-box experience), but e.g. being able to know when a fallback model has been used means that in terms of pure accuracy, you might still be better off defaulting to Opus 5.5 and re-routing to Sonnet 5.5 yourself when you get the fallback.
    2. MadameMinty · · focus · HN ↗
      That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?

      I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?

      1. manojlds · · focus · HN ↗
        Fallback was usually Opus 4.8
        1. Aissen · · focus · HN ↗
          Note that it seems that it no longer falls back automatically. So the actual score of Opus 5.5 will be even lower (fail vs fallback than can succeed)
    3. radlad · · focus · HN ↗
      I believe you meant to cite the Opus 5.5 System Card which states:

      > Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).

      &gt; <a href="https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;fc1b44717c85dc068bc6ba5024219938094694bd&#x2F;Claude%20Opus%205.5%20System%20Card.pdf" rel="nofollow">https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;fc1b44717c85dc068bc6ba50242199...

      I cannot find a Sonnet 5.5 system card.

      1. abejora · · focus · HN ↗
        It was linked in another HN post: <a href="https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;870c8f525702625d2c62fc6dd04c857e3250bec1&#x2F;Claude%20Sonnet%205.5%20System%20Card.pdf" rel="nofollow">https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;870c8f525702625d2c62fc6dd04c85...
    4. Leary · · focus · HN ↗
      And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!
    5. manojlds · · focus · HN ↗
      Isn&#x27;t that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?
      1. verdverm · · focus · HN ↗
        this &quot;feature&quot; is one of the primary that caused me to cancel and move to exclusively open weight based systems
    6. oh_no · · focus · HN ↗
      it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

      AA intelegence index (agent harness doesn&#x27;t have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

      Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

    7. subscribed · · focus · HN ↗
      I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn&#x27;t really matter why.

      Anthropic made it that way, and I&#x27;d say the lower score is accurate.

      1. cromka · · focus · HN ↗
        It matters if it&#x27;s not Sonnet performing the task, doesn&#x27;t it?
        1. subscribed · · focus · HN ↗
          *Opus

          But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.

          So both you&#x27;re right, it matters because it wasn&#x27;t the model examined, and it doesn&#x27;t matter because the score reflects nerfed experience resulting in nerfed results.

      2. bitexploder · · focus · HN ↗
        If you care about the things terminal bench cares about, yes. Sonnet was probably trained aggressively on agentic coding and things that align well with deepswe and terminal bench and or tuned heavily for those tasks. Sonnet is an agent likely to do more of those tasks and be given the more grunt work tasks. Whilst Opus&#x27; wider knowledge pool means it can deal with a much higher variety of real world situations successfully. And, those benches are often timed or limited. Opus may have been running out of time. Looots of factors.
        1. subscribed · · focus · HN ↗
          I see you.

          I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it&#x27;s reflected in the score.

          (incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it&#x27;s okay. No one outside uses it :p)

          I agree with your point - IMO lower Opus score in these suggests that in general it&#x27;s worse for these tasks. Not that it&#x27;s a worse model in general.

    8. falcor84 · · focus · HN ↗
      So I suppose the easy fix for Anthropic would be to have Opus 5.5 now fall back to Sonnet 5.5, right?
    9. shmel · · focus · HN ↗
      It&#x27;d be relevant if I could disable safeguards. As long as I can&#x27;t, this is Opus 5.5 performance I have to deal with.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.