‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. MadameMinty · · focus · HN ↗
      That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?

      I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?

      1. manojlds · · focus · HN ↗
        Fallback was usually Opus 4.8
        1. Aissen · · focus · HN ↗
          Note that it seems that it no longer falls back automatically. So the actual score of Opus 5.5 will be even lower (fail vs fallback than can succeed)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.