Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
Yes, the OpenAI GPT-6 Astra limit is 128,000 as well: <a href="https://developers.openai.com/api/docs/models/gpt-6-astra" rel="nofollow">https://developers.openai.com/api/docs/models/gpt-6-astra
Gemini 3.8 Flash is 65,536 <a href="https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash" rel="nofollow">https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Score Tokens Reason Cost
Kimi K3 Max 44 48k 32k $2.00
Half reason 44 32k ? 16k ? ?
Opus Med 51 26k 12k $1.34
Opus High 54 36k 18k $1.82
Opus Max 58 119k 84k $5.98
Sonnet Med 41 ? ? $0.59
Sonnet High 47 ? ? $1.08
Sonnet Max 56 193k 142k $7.60
Medium is Anthropic's default.
Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
Mainly because they're funny, but it's also because I try pretty hard to make the comment more interesting than just "here's a pelican". In this case I used the pelicans to talk about the 128,000 token limit bug at "max" and share comparative pricing.
In the GPT-6 comment I included full visual comparison grids: <a href="https://news.ycombinator.com/item?id=49805509#49806126">https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: <a href="https://news.ycombinator.com/item?id=49639090#49645591">https://news.ycombinator.com/item?id=49639090#49645591
Agreed. It was a creative and unique test for a while. Now, no offense to the author, it feels like every conversation about a new model is dominated by the pelican on a bike posts as they always become the top comment.
It's an easy way to compare the coding and creative strengths of models. I prefer them over reading a tabular comparison of benchmarks which you have no real insights into.
Because hn has some kind of a community and not every comment is gold (see yours for example) and people are able to skip comments if they don't enjoy them?
I can't understand it. Clearly someone cares because like you say the comments are always upvoted. But why anyone cares I simply don't know. It just feels like attention seeking behaviour to continue posting it.
Linking to <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... should be pretty inoffensive (I started habitually linking to that after people kept complaining about linking to my blog) - that page renders Markdown with SVG embedded in it, but doesn't link to the rest of my site at all.
It would have taken me a while to stumble upon this "running out of tokens" on MAX thinking issue without his trials and post. I've seen the pelicans for years now, and if they stopped coming for some reason, I would probably go to his site to catch up on recent models and findings. So I don't mind them
Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.
I’m not sure whether that’s a feature or a bug at this point though.
I feel the fact that these models always modify the body design of a pelican to fit the bike rather than the other way around represents a fundamental issue with AI.
simonw · · focus · HN ↗
<a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
Low and medium both used 0 thinking tokens.croemer · · focus · HN ↗
gumby271 · · focus · HN ↗
miki123211 · · focus · HN ↗
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
TomGarden · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
petu · · focus · HN ↗
simonw · · focus · HN ↗
croemer · · focus · HN ↗
simonw · · focus · HN ↗
Insanity · · focus · HN ↗
simonw · · focus · HN ↗
Gemini 3.8 Flash is 65,536 <a href="https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash" rel="nofollow">https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
NewJazz · · focus · HN ↗
heyjstn · · focus · HN ↗
dmd · · focus · HN ↗
aimaxxed · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
platinumrad · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
usef- · · focus · HN ↗
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Medium is Anthropic's default.Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
keeeba · · focus · HN ↗
parkersweb · · focus · HN ↗
pelicanmaxer · · focus · HN ↗
amelius · · focus · HN ↗
PS: the next human that brings up pelicans on bicycles should try to draw them.
dennisy · · focus · HN ↗
Any model release it’s the top comment, I do not understand why.
uncivilized · · focus · HN ↗
conception · · focus · HN ↗
uncivilized · · focus · HN ↗
simonw · · focus · HN ↗
In the GPT-6 comment I included full visual comparison grids: <a href="https://news.ycombinator.com/item?id=49805509#49806126">https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: <a href="https://news.ycombinator.com/item?id=49639090#49645591">https://news.ycombinator.com/item?id=49639090#49645591
kennyadam · · focus · HN ↗
simonw · · focus · HN ↗
marktolson · · focus · HN ↗
Glemmlko · · focus · HN ↗
mvdtnz · · focus · HN ↗
simonw · · focus · HN ↗
mi_lk · · focus · HN ↗
simonw · · focus · HN ↗
ceroxylon · · focus · HN ↗
I find it useful (as well as a fun art project).
mi_lk · · focus · HN ↗
permalac · · focus · HN ↗
dramebaaz · · focus · HN ↗
krzyk · · focus · HN ↗
You can also check for any kind of degradation of them - you have the prompt, it doesn't use much $.
codingisfreedom · · focus · HN ↗
I’m not sure whether that’s a feature or a bug at this point though.
mgaunard · · focus · HN ↗
hooloovoo_zoo · · focus · HN ↗
nicolamanzini · · focus · HN ↗
[dead]
mewse-hn · · focus · HN ↗