‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. simonw · · focus · HN ↗
    Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

    <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1d85a9be7f3ecce26e7f1569161a0d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Here&#x27;s how the thinking effort levels compare:

      low
      27 input, 1,623 output, thinking_tokens: 0
      1.6284
      Duration: 10138ms (10s)
      
      medium
      27 input, 1,796 output, thinking_tokens: 0
      1.7914 cents
      Duration: 11266ms (11s)
    
      high
      27 input, 2,334 output, thinking_tokens: 745
      2.3394 cents
      Duration: 17376ms (17s)
    
      xhigh
      27 input, 5,730 output, thinking_tokens: 2535
      5.7354 cents
      Duration: 41882ms (41s)
    
      max (failed to return response)
      27 input, 128,000 output, thinking_tokens: 128000
      $1.28
      Duration: 940617ms (15m 40s)
    
    Low and medium both used 0 thinking tokens.
    1. platinumrad · · focus · HN ↗
      The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks&#x27;s Ember-1, which was explicitly trained to preserve the quality of a model&#x27;s responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
      1. usef- · · focus · HN ↗
        Anthropic&#x27;s &quot;Max&quot; modes seem like a yolo mode: &quot;use 10x the tokens to try to break the hardest possible problems&quot;. But their models don&#x27;t seem less efficient at normal reasoning modes.

        I can&#x27;t see Ember on AA&#x27;s index yet, but their post claims &quot;half the reasoning tokens for the same answers&quot; as Kimi K3.

        That would make it about so, I assume?

                     Score  Tokens  Reason  Cost 
         Kimi K3 Max    44  48k     32k     $2.00
         Half reason    44  32k ?   16k ?    ?
        
         Opus Med       51  26k     12k     $1.34
         Opus High      54  36k     18k     $1.82
         Opus Max       58  119k    84k     $5.98
        
         Sonnet Med     41  ?       ?       $0.59
         Sonnet High    47  ?       ?       $1.08
         Sonnet Max     56  193k    142k    $7.60
        
        Medium is Anthropic&#x27;s default.

        Having a less efficient mode isn&#x27;t necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.