‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. manojlds · · focus · HN ↗
      Isn't that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?
      1. verdverm · · focus · HN ↗
        this "feature" is one of the primary that caused me to cancel and move to exclusively open weight based systems
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.