Claude Opus 5.5
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Claude Opus 5.5
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
stefan_ · · focus · HN ↗
Ah, I see nothing has changed. Are Anthropic aware that their models are generating tons of gibberish? In comparison, Astra is sublime.
mupuff1234 · · focus · HN ↗
Lord_Zero · · focus · HN ↗
setsewerd · · focus · HN ↗
WarmWash · · focus · HN ↗
petesergeant · · focus · HN ↗
mupuff1234 · · focus · HN ↗
Less companies involved means less pressure to go fast.
solenoid0937 · · focus · HN ↗
nozzlegear · · focus · HN ↗
roughly · · focus · HN ↗
setsewerd · · focus · HN ↗
icrbow · · focus · HN ↗
prodigycorp · · focus · HN ↗
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.
Thank GOD
SadErn · · focus · HN ↗
[dead]
Gander5739 · · focus · HN ↗
unethical_ban · · focus · HN ↗
alpineman · · focus · HN ↗
voiceeh · · focus · HN ↗
palata · · focus · HN ↗
gopalv · · focus · HN ↗
[1] - <a href="https://news.ycombinator.com/item?id=6372466">https://news.ycombinator.com/item?id=6372466
meerita · · focus · HN ↗
nozzlegear · · focus · HN ↗
nickandbro · · focus · HN ↗
keeganpoppen · · focus · HN ↗
cogythea · · focus · HN ↗
jdmoreira · · focus · HN ↗
NielsHarksen · · focus · HN ↗
calibas · · focus · HN ↗
We can't test it properly because it knows it's being tested.
johntb86 · · focus · HN ↗
actionfromafar · · focus · HN ↗
pookieinc · · focus · HN ↗
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
jbellis · · focus · HN ↗
meric_ · · focus · HN ↗
Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too
randomblock1 · · focus · HN ↗
viccis · · focus · HN ↗
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
scrollop · · focus · HN ↗
thibran · · focus · HN ↗
slacktivism123 · · focus · HN ↗
benjiro29 · · focus · HN ↗
Its funny how every new model cost less to run. But this often does not match with reality.
O, and subscriptions getting less usage, despite how the new models "costs xx% less to run".
alvis · · focus · HN ↗
aennassiri · · focus · HN ↗
ryanscio · · focus · HN ↗
benjiro29 · · focus · HN ↗
tag2103 · · focus · HN ↗
seviu · · focus · HN ↗
Gattopardo · · focus · HN ↗
iamsyr · · focus · HN ↗
jacobgold · · focus · HN ↗
km144 · · focus · HN ↗
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
CPLX · · focus · HN ↗
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be that bust worse. Happy to be proven wrong.
port3000 · · focus · HN ↗
booty · · focus · HN ↗
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
Syntaf · · focus · HN ↗
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
cbg0 · · focus · HN ↗
simianwords · · focus · HN ↗
suddenlybananas · · focus · HN ↗
booty · · focus · HN ↗
1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
Solvyx · · focus · HN ↗
[dead]
ayhanfuat · · focus · HN ↗
> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
kingstnap · · focus · HN ↗
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
hirako2000 · · focus · HN ↗
danbrooks · · focus · HN ↗
b38484848 · · focus · HN ↗
jdw64 · · focus · HN ↗
richardjennings · · focus · HN ↗
glub · · focus · HN ↗
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
somewhatjustin · · focus · HN ↗
Nice. I was starting to think Haiku was going to be abandoned.
iamsyr · · focus · HN ↗
mococa · · focus · HN ↗
ricardobeat · · focus · HN ↗
xenit_v0 · · focus · HN ↗
[dead]
greenavocado · · focus · HN ↗
anthonyrstevens · · focus · HN ↗
jatins · · focus · HN ↗
Thank you.
giancarlostoro · · focus · HN ↗
[dead]
keeeba · · focus · HN ↗
velcrovan · · focus · HN ↗
adastra22 · · focus · HN ↗
hadlock · · focus · HN ↗
rs_rs_rs_rs_rs · · focus · HN ↗
postalcoder · · focus · HN ↗
Notables:
That said, Opus 5 showed us that impressive benchmarks can only take us so far. Hoping we're not in for such disappointment again.notduckrabbit · · focus · HN ↗
manmal · · focus · HN ↗
bredren · · focus · HN ↗
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
tomhow · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
sidrag22 · · focus · HN ↗
So for that 20$ tier for the entire summer and into fall, i was on their 2nd class public model(4.8) released in May. Not surprisingly it became my grunt model, doing the simple work. By far the least I've used Anthropic models in the last 2 years.