Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
You're right about its real world performance, and I worded my original comment wrongly.
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
You can read “a little bit” (i.e. “not too much”) into it (it does indeed tell you about the out-of-the-box experience), but e.g. being able to know when a fallback model has been used means that in terms of pure accuracy, you might still be better off defaulting to Opus 5.5 and re-routing to Sonnet 5.5 yourself when you get the fallback.
I believe you meant to cite the Opus 5.5 System Card which states:
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
It was linked in another HN post: <a href="https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf" rel="nofollow">https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c85...
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.
If you care about the things terminal bench cares about, yes. Sonnet was probably trained aggressively on agentic coding and things that align well with deepswe and terminal bench and or tuned heavily for those tasks. Sonnet is an agent likely to do more of those tasks and be given the more grunt work tasks. Whilst Opus' wider knowledge pool means it can deal with a much higher variety of real world situations successfully. And, those benches are often timed or limited. Opus may have been running out of time. Looots of factors.
I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it's reflected in the score.
(incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it's okay. No one outside uses it :p)
I agree with your point - IMO lower Opus score in these suggests that in general it's worse for these tasks. Not that it's a worse model in general.
abejora · · focus · HN ↗
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
eli · · focus · HN ↗
abejora · · focus · HN ↗
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
joeyhage · · focus · HN ↗
bb-connor · · focus · HN ↗
swiftcoder · · focus · HN ↗
[dead]
verdverm · · focus · HN ↗
this is your brain on claude...
corporate needs you to find the difference
ramon156 · · focus · HN ↗
chis · · focus · HN ↗
spider-mario · · focus · HN ↗
MadameMinty · · focus · HN ↗
I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?
manojlds · · focus · HN ↗
Aissen · · focus · HN ↗
radlad · · focus · HN ↗
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
> <a href="https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf" rel="nofollow">https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
I cannot find a Sonnet 5.5 system card.
abejora · · focus · HN ↗
Leary · · focus · HN ↗
manojlds · · focus · HN ↗
verdverm · · focus · HN ↗
oh_no · · focus · HN ↗
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
subscribed · · focus · HN ↗
Anthropic made it that way, and I'd say the lower score is accurate.
cromka · · focus · HN ↗
subscribed · · focus · HN ↗
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.
bitexploder · · focus · HN ↗
subscribed · · focus · HN ↗
I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it's reflected in the score.
(incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it's okay. No one outside uses it :p)
I agree with your point - IMO lower Opus score in these suggests that in general it's worse for these tasks. Not that it's a worse model in general.
falcor84 · · focus · HN ↗
shmel · · focus · HN ↗