Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Unofficial Hacker News client; not affiliated with Y Combinator.
simonw · · focus · HN ↗
I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.
I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.
Transcript for one attempt here - expand the "Reasoning trace" bit to see it: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F61fd7c3683fffce9a3ab7c43d1180024#reasoning-3" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
az226 · · focus · HN ↗
simonw · · focus · HN ↗
Piping the visible reasoning trace through their token counter API (I use <a href="https://tools.simonwillison.net/claude-token-counter" rel="nofollow">https://tools.simonwillison.net/claude-token-counter for that) counts 27,888 tokens, so it's definitely a summary of the 128,000 actual token trace.
RGS1811 · · focus · HN ↗
I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
simonw · · focus · HN ↗
dgellow · · focus · HN ↗
cubefox · · focus · HN ↗
croemer · · focus · HN ↗
0x10ca1h0st · · focus · HN ↗
agar · · focus · HN ↗
"Create an SVG of Shaquille O'Neal eating potato chips shaped like a telecopier."
Shaq'sFaxSnacksMaxx
pgwhalen · · focus · HN ↗
beardsciences · · focus · HN ↗
Someone1234 · · focus · HN ↗
My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.
seabass-salmon · · focus · HN ↗
arcanemachiner · · focus · HN ↗
Have you ruled out the possibility that your system prompt, AGENTS.md, or increasing codebase complexity are not to blame?
dc443 · · focus · HN ↗
samuelknight · · focus · HN ↗
sidewndr46 · · focus · HN ↗
I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.
zerof1l · · focus · HN ↗
zozbot234 · · focus · HN ↗
dannyw · · focus · HN ↗
zozbot234 · · focus · HN ↗
qlte · · focus · HN ↗
.... after running a 24/7 model torture factory for 6 months to improve their JSONBench 9.5 scores by 0.2%.
(Are they still doing that, BTW?)
usagisushi · · focus · HN ↗
As a workaround, add this to CLAUDE.md: "Claude! Happiness is mandatory!"
EDIT: 15 years from now, I’ll be sent to re-education for this thought crime.
Mtinie · · focus · HN ↗
paradox460 · · focus · HN ↗
judge2020 · · focus · HN ↗
Most of what I've heard is that raw reasoning traces are really good for distillation, although no idea how much the summarization actually hurts distillation.
Barbing · · focus · HN ↗
hypfer · · focus · HN ↗
I'd argue that they're a necessity if you want to use the LLM as a tool instead of a black box that just does stuff for you.
It gives you a lot finer control over where the solution ends up when you can follow along the thinking trace and modulate your inputs based on what you saw in there.
And, additionally, it gives you a lot more understanding of what the model can or cannot do. Strengths and weaknesses and all that.
Using claude is like buying a car where you cannot legally open the hood. It tells you that there is something specific under there, and often it actually drives like that too, but how exactly it looks you will never see.
For some people this is fine. I do not think that these people will survive. Figuratively speaking but also literally speaking.
World's changing. Opaque abstraction like that is a luxury depending on (geo)political stability.
realusername · · focus · HN ↗
I also switch to a better model for more complex tasks, also in low settings
therealdrag0 · · focus · HN ↗
realusername · · focus · HN ↗
Regardless of the model, running it for hours means that the model will takes decisions and assumptions alone instead of you.
baq · · focus · HN ↗
realusername · · focus · HN ↗
No matter how clever the model is, most problems have multiple valid, invalid and unclear decisions to make, running it for a long time is just picking the first option on everything, which isn't usually what you want
therealdrag0 · · focus · HN ↗
It’s like learning to delegate and let go. The more senior I got the more I had to learn to let other engineers make decision i thought were suboptimal but mostly good enough. That positioned me well to be comfortable with agents. It’s contextual how much I’m willing to give them control and how much to review afterwards.
koonsolo · · focus · HN ↗
baq · · focus · HN ↗
alansaber · · focus · HN ↗
Izmaki · · focus · HN ↗
Gcam · · focus · HN ↗
Jimmc414 · · focus · HN ↗
<a href="https://platform.claude.com/docs/en/models/opus-5-5/overview" rel="nofollow">https://platform.claude.com/docs/en/models/opus-5-5/overview
fr2029 · · focus · HN ↗
[dead]
striking · · focus · HN ↗
amelius · · focus · HN ↗
This whole test tells me nothing.
The next person who proposes to use it should draw it themselves first.
simonw · · focus · HN ↗
amelius · · focus · HN ↗
1. I don't want my AI to have super-human intelligence. It would generate code that I do not understand. A coder with human capabilities is better for me.
2. We're interested in AGI. If the AI reaches human intelligence, that's a milestone. Drawing bicycles with pelicans is superhuman. Hence not a relevant test.
simonw · · focus · HN ↗
It really isn't. Many humans can draw a bicycle, and a pelican, just fine.
Reddit_MLP2 · · focus · HN ↗
nijave · · focus · HN ↗