‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. subscribed · · focus · HN ↗
      I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.

      Anthropic made it that way, and I'd say the lower score is accurate.

      1. cromka · · focus · HN ↗
        It matters if it's not Sonnet performing the task, doesn't it?
        1. subscribed · · focus · HN ↗
          *Opus

          But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.

          So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.