The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
That makes sense. I'm interested in seeing where Haiku 5.5 comes in then when it gets released. It feels like the low intelligence / fast niche will be covered there.
It depends entirely on its capabilities. If it is significantly smarter than Luna, which frankly is quite likely, then a lot of people won't mind paying more.
Or if its luna at 1000 tok/sec. Speed is what most of my peers care most about these days since less intelligent models can do most grunt work just fine.
My guess is that haiku will be a mid release. They don’t seem to care to compete for the low end. Something akin to gpt6 sol level intelligence at $1 / $5 pricing. Then not release an update for 6+ months.
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
wkcheng · · focus · HN ↗
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
quatotor · · focus · HN ↗
[dead]
SubiculumCode · · focus · HN ↗
ricardobeat · · focus · HN ↗
wkcheng · · focus · HN ↗
oh_no · · focus · HN ↗
i think they see what openai charges for luna and just don't want to try and compete
canad3nse · · focus · HN ↗
ac29 · · focus · HN ↗
enraged_camel · · focus · HN ↗
conception · · focus · HN ↗
mchusma · · focus · HN ↗
pdantix · · focus · HN ↗
water-drummer · · focus · HN ↗
solenoid0937 · · focus · HN ↗
RussianCow · · focus · HN ↗
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
delillos · · focus · HN ↗
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
usaar333 · · focus · HN ↗
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
verdverm · · focus · HN ↗
NotSuspicious · · focus · HN ↗
verdverm · · focus · HN ↗
Their new Ember-1 model is pretty good, fine-tune of Kimi3 with way less thinking
Does this make it an American model or is it still Chinese?
Jcampuzano2 · · focus · HN ↗
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
dominotw · · focus · HN ↗
benjiro29 · · focus · HN ↗
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.