‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. subscribed · · focus · HN ↗
      I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.

      Anthropic made it that way, and I'd say the lower score is accurate.

      1. bitexploder · · focus · HN ↗
        If you care about the things terminal bench cares about, yes. Sonnet was probably trained aggressively on agentic coding and things that align well with deepswe and terminal bench and or tuned heavily for those tasks. Sonnet is an agent likely to do more of those tasks and be given the more grunt work tasks. Whilst Opus' wider knowledge pool means it can deal with a much higher variety of real world situations successfully. And, those benches are often timed or limited. Opus may have been running out of time. Looots of factors.
        1. subscribed · · focus · HN ↗
          I see you.

          I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it's reflected in the score.

          (incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it's okay. No one outside uses it :p)

          I agree with your point - IMO lower Opus score in these suggests that in general it's worse for these tasks. Not that it's a worse model in general.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.