‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. eli · · focus · HN ↗
      Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
      1. spider-mario · · focus · HN ↗
        You can read “a little bit” (i.e. “not too much”) into it (it does indeed tell you about the out-of-the-box experience), but e.g. being able to know when a fallback model has been used means that in terms of pure accuracy, you might still be better off defaulting to Opus 5.5 and re-routing to Sonnet 5.5 yourself when you get the fallback.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.