‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. ramish94 · · focus · HN ↗
    In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.

    Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)

    FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)

    CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)

    Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family

    1. level87 · · focus · HN ↗
      This is crazy, what is the point of all these equivalent models?
      1. salviati · · focus · HN ↗
        Price going down on each release
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.