Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.
Mainly because they're funny, but it's also because I try pretty hard to make the comment more interesting than just "here's a pelican". In this case I used the pelicans to talk about the 128,000 token limit bug at "max" and share comparative pricing.
In the GPT-6 comment I included full visual comparison grids: <a href="https://news.ycombinator.com/item?id=49805509#49806126">https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: <a href="https://news.ycombinator.com/item?id=49639090#49645591">https://news.ycombinator.com/item?id=49639090#49645591
Agreed. It was a creative and unique test for a while. Now, no offense to the author, it feels like every conversation about a new model is dominated by the pelican on a bike posts as they always become the top comment.
It's an easy way to compare the coding and creative strengths of models. I prefer them over reading a tabular comparison of benchmarks which you have no real insights into.
Because hn has some kind of a community and not every comment is gold (see yours for example) and people are able to skip comments if they don't enjoy them?
I can't understand it. Clearly someone cares because like you say the comments are always upvoted. But why anyone cares I simply don't know. It just feels like attention seeking behaviour to continue posting it.
Linking to <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... should be pretty inoffensive (I started habitually linking to that after people kept complaining about linking to my blog) - that page renders Markdown with SVG embedded in it, but doesn't link to the rest of my site at all.
It would have taken me a while to stumble upon this "running out of tokens" on MAX thinking issue without his trials and post. I've seen the pelicans for years now, and if they stopped coming for some reason, I would probably go to his site to catch up on recent models and findings. So I don't mind them
simonw · · focus · HN ↗
<a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
Low and medium both used 0 thinking tokens.dennisy · · focus · HN ↗
Any model release it’s the top comment, I do not understand why.
uncivilized · · focus · HN ↗
conception · · focus · HN ↗
uncivilized · · focus · HN ↗
simonw · · focus · HN ↗
In the GPT-6 comment I included full visual comparison grids: <a href="https://news.ycombinator.com/item?id=49805509#49806126">https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: <a href="https://news.ycombinator.com/item?id=49639090#49645591">https://news.ycombinator.com/item?id=49639090#49645591
kennyadam · · focus · HN ↗
simonw · · focus · HN ↗
marktolson · · focus · HN ↗
Glemmlko · · focus · HN ↗
mvdtnz · · focus · HN ↗
simonw · · focus · HN ↗
mi_lk · · focus · HN ↗
simonw · · focus · HN ↗
ceroxylon · · focus · HN ↗
I find it useful (as well as a fun art project).
mi_lk · · focus · HN ↗
permalac · · focus · HN ↗
dramebaaz · · focus · HN ↗
krzyk · · focus · HN ↗
You can also check for any kind of degradation of them - you have the prompt, it doesn't use much $.