‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. abejora · · focus · HN ↗
    Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

    Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

    [1] Section 8.5 of the Sonnet 5.5 System Card

    1. oh_no · · focus · HN ↗
      it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

      AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

      Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.