Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.
abejora · · focus · HN ↗
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
subscribed · · focus · HN ↗
Anthropic made it that way, and I'd say the lower score is accurate.
cromka · · focus · HN ↗
subscribed · · focus · HN ↗
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.