‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. simonw · · focus · HN ↗
    Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

    <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Here&#x27;s how the thinking effort levels compare:

      low
      27 input, 1,623 output, thinking_tokens: 0
      1.6284
      Duration: 10138ms (10s)
      
      medium
      27 input, 1,796 output, thinking_tokens: 0
      1.7914 cents
      Duration: 11266ms (11s)
    
      high
      27 input, 2,334 output, thinking_tokens: 745
      2.3394 cents
      Duration: 17376ms (17s)
    
      xhigh
      27 input, 5,730 output, thinking_tokens: 2535
      5.7354 cents
      Duration: 41882ms (41s)
    
      max (failed to return response)
      27 input, 128,000 output, thinking_tokens: 128000
      $1.28
      Duration: 940617ms (15m 40s)
    
    Low and medium both used 0 thinking tokens.
    1. croemer · · focus · HN ↗
      This is evidence that Sonnet 5.5 wasn&#x27;t yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
      1. gumby271 · · focus · HN ↗
        If it was trained on HN, there would be a 60% chance of it just saying &quot;I&#x27;m so tired of this request, can we please move on&quot;
        1. miki123211 · · focus · HN ↗
          I&#x27;d say:

          30% chance of responding with something about Enshittification and how it can&#x27;t fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it&#x27;s running on FreeBSD behind PiHole).

          30% chance of complaining that it&#x27;s being subsidized and that &quot;prices are going to go up bro.&quot;

          30% chance of some unrelated rant on ID checks for age verification.

          10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.

    2. TomGarden · · focus · HN ↗
      Where do you run sonnet&#x2F;opus where you are limited to 128k, given they are both 1M context window models?
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. petu · · focus · HN ↗
        That&#x27;s max output tokens per response limit, separate from context length
      3. simonw · · focus · HN ↗
        It&#x27;s the output token limit, which has been 128,000 for Claude models for quite a while note
        1. croemer · · focus · HN ↗
          Pretty crazy that the model doesn&#x27;t know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn&#x27;t this be possible to work into the architecture?
          1. simonw · · focus · HN ↗
            I think this is a bug. I&#x27;ve not seen this problem from any of the other frontier models.
            1. Insanity · · focus · HN ↗
              Do other models put a hard cap on the output tokens it can generate?
              1. simonw · · focus · HN ↗
                Yes, the OpenAI GPT-6 Astra limit is 128,000 as well: <a href="https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;models&#x2F;gpt-6-astra" rel="nofollow">https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;models&#x2F;gpt-6-astra

                Gemini 3.8 Flash is 65,536 <a href="https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;models&#x2F;gemini-3.8-flash" rel="nofollow">https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;models&#x2F;gemini-3.8-flas...

            2. NewJazz · · focus · HN ↗
              [delayed]
    3. heyjstn · · focus · HN ↗
      I think the next models will be benchmaxxing on the Pelican benchmark tbh
      1. dmd · · focus · HN ↗
        wow nobody but you has ever thought of this and certainly simonw has never addressed this
    4. aimaxxed · · focus · HN ↗
      “Pelicans are solved.”
      1. [deleted] · · focus · HN ↗

        [deleted]

    5. platinumrad · · focus · HN ↗
      The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks&#x27;s Ember-1, which was explicitly trained to preserve the quality of a model&#x27;s responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. usef- · · focus · HN ↗
        Anthropic&#x27;s &quot;Max&quot; modes seem like a yolo mode: &quot;use 10x the tokens to try to break the hardest possible problems&quot;. But their models don&#x27;t seem less efficient at normal reasoning modes.

        I can&#x27;t see Ember on AA&#x27;s index yet, but their post claims &quot;half the reasoning tokens for the same answers&quot; as Kimi K3.

        That would make it about so, I assume?

                     Score  Tokens  Reason  Cost 
         Kimi K3 Max    44  48k     32k     $2.00
         Half reason    44  32k ?   16k ?    ?
        
         Opus Med       51  26k     12k     $1.34
         Opus High      54  36k     18k     $1.82
         Opus Max       58  119k    84k     $5.98
        
         Sonnet Med     41  ?       ?       $0.59
         Sonnet High    47  ?       ?       $1.08
         Sonnet Max     56  193k    142k    $7.60
        
        Medium is Anthropic&#x27;s default.

        Having a less efficient mode isn&#x27;t necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.

    6. keeeba · · focus · HN ↗
      Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing&#x2F;capability range?
    7. parkersweb · · focus · HN ↗
      I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!
    8. pelicanmaxer · · focus · HN ↗
      that pelican one-pedaling
    9. amelius · · focus · HN ↗
      This is great news because it means the model has not been benchmaxxed on stupid metrics.

      PS: the next human that brings up pelicans on bicycles should try to draw them.

    10. dennisy · · focus · HN ↗
      Does anyone really still care about these pelicans?

      Any model release it’s the top comment, I do not understand why.

      1. uncivilized · · focus · HN ↗
        Karma farming by parent commenter and HNers’ tendency to upvote low quality content (not dissimilar to other social media networks)
        1. conception · · focus · HN ↗
          Is this low quality content relative to most HN comments?
          1. uncivilized · · focus · HN ↗
            Very few HN comments are high quality
      2. simonw · · focus · HN ↗
        Mainly because they&#x27;re funny, but it&#x27;s also because I try pretty hard to make the comment more interesting than just &quot;here&#x27;s a pelican&quot;. In this case I used the pelicans to talk about the 128,000 token limit bug at &quot;max&quot; and share comparative pricing.

        In the GPT-6 comment I included full visual comparison grids: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49805509#49806126">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49805509#49806126

        For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49639090#49645591">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49639090#49645591

      3. kennyadam · · focus · HN ↗
        Agreed. It was a creative and unique test for a while. Now, no offense to the author, it feels like every conversation about a new model is dominated by the pelican on a bike posts as they always become the top comment.
        1. simonw · · focus · HN ↗
          You can click the little [-] icon next to the post to collapse the entire sub-thread. I do that all the time.
      4. marktolson · · focus · HN ↗
        It&#x27;s an easy way to compare the coding and creative strengths of models. I prefer them over reading a tabular comparison of benchmarks which you have no real insights into.
      5. Glemmlko · · focus · HN ↗
        Because hn has some kind of a community and not every comment is gold (see yours for example) and people are able to skip comments if they don&#x27;t enjoy them?
      6. mvdtnz · · focus · HN ↗
        I can&#x27;t understand it. Clearly someone cares because like you say the comments are always upvoted. But why anyone cares I simply don&#x27;t know. It just feels like attention seeking behaviour to continue posting it.
        1. simonw · · focus · HN ↗
          Isn&#x27;t posting any comment on a forum like Hacker News &quot;attention seeking behavior&quot;?
          1. mi_lk · · focus · HN ↗
            Did you ignore the continue part that compounds the attention seeking?
            1. simonw · · focus · HN ↗
              Linking to <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... should be pretty inoffensive (I started habitually linking to that after people kept complaining about linking to my blog) - that page renders Markdown with SVG embedded in it, but doesn&#x27;t link to the rest of my site at all.
      7. ceroxylon · · focus · HN ↗
        This is a community and it is an inside joke at this point, it wouldn&#x27;t be a proper model release without Simon&#x27;s pelicans.

        I find it useful (as well as a fun art project).

      8. mi_lk · · focus · HN ↗
        Sick of the cheap shot
      9. permalac · · focus · HN ↗
        I would not say I care, but I do have curiosity. I expect one day they will start adding some textures or something like that.
      10. dramebaaz · · focus · HN ↗
        It would have taken me a while to stumble upon this &quot;running out of tokens&quot; on MAX thinking issue without his trials and post. I&#x27;ve seen the pelicans for years now, and if they stopped coming for some reason, I would probably go to his site to catch up on recent models and findings. So I don&#x27;t mind them
      11. krzyk · · focus · HN ↗
        I do, it gives some fun comparison between models.

        You can also check for any kind of degradation of them - you have the prompt, it doesn&#x27;t use much $.

    11. codingisfreedom · · focus · HN ↗
      Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.

      I’m not sure whether that’s a feature or a bug at this point though.

    12. mgaunard · · focus · HN ↗
      what&#x27;s most surprising is the difference between high and xhigh
    13. hooloovoo_zoo · · focus · HN ↗
      I feel the fact that these models always modify the body design of a pelican to fit the bike rather than the other way around represents a fundamental issue with AI.
    14. nicolamanzini · · focus · HN ↗

      [dead]

    15. mewse-hn · · focus · HN ↗
      max thinking means benchmaxxing i guess
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.