‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. radlad · · focus · HN ↗
      I believe you meant to cite the Opus 5.5 System Card which states:

      > Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).

      &gt; <a href="https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;fc1b44717c85dc068bc6ba5024219938094694bd&#x2F;Claude%20Opus%205.5%20System%20Card.pdf" rel="nofollow">https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;fc1b44717c85dc068bc6ba50242199...

      I cannot find a Sonnet 5.5 system card.

      1. abejora · · focus · HN ↗
        It was linked in another HN post: <a href="https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;870c8f525702625d2c62fc6dd04c857e3250bec1&#x2F;Claude%20Sonnet%205.5%20System%20Card.pdf" rel="nofollow">https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;870c8f525702625d2c62fc6dd04c85...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.