‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. simonw · · focus · HN ↗
    Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

    <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Here&#x27;s how the thinking effort levels compare:

      low
      27 input, 1,623 output, thinking_tokens: 0
      1.6284
      Duration: 10138ms (10s)
      
      medium
      27 input, 1,796 output, thinking_tokens: 0
      1.7914 cents
      Duration: 11266ms (11s)
    
      high
      27 input, 2,334 output, thinking_tokens: 745
      2.3394 cents
      Duration: 17376ms (17s)
    
      xhigh
      27 input, 5,730 output, thinking_tokens: 2535
      5.7354 cents
      Duration: 41882ms (41s)
    
      max (failed to return response)
      27 input, 128,000 output, thinking_tokens: 128000
      $1.28
      Duration: 940617ms (15m 40s)
    
    Low and medium both used 0 thinking tokens.
    1. dennisy · · focus · HN ↗
      Does anyone really still care about these pelicans?

      Any model release it’s the top comment, I do not understand why.

      1. simonw · · focus · HN ↗
        Mainly because they&#x27;re funny, but it&#x27;s also because I try pretty hard to make the comment more interesting than just &quot;here&#x27;s a pelican&quot;. In this case I used the pelicans to talk about the 128,000 token limit bug at &quot;max&quot; and share comparative pricing.

        In the GPT-6 comment I included full visual comparison grids: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49805509#49806126">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49805509#49806126

        For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49639090#49645591">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49639090#49645591

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.