‹ BackHN Continuity

Thread

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

333 points · 106 comments · theanonymousone

  1. simonw · · focus · HN ↗
    This is the page for the &quot;max&quot; reasoning setting. The page for xhigh is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh and the page for medium (the default setting) is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium

    I&#x27;ve failed twice to get &quot;Generate an SVG of a pelican riding a bicycle&quot; to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

    I&#x27;m suspicious that &quot;max&quot; may be virtually useless if it&#x27;s that easy to have it overthink to the point that it doesn&#x27;t get to a response.

    Transcript for one attempt here - expand the &quot;Reasoning trace&quot; bit to see it: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F61fd7c3683fffce9a3ab7c43d1180024#reasoning-3" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    1. Someone1234 · · focus · HN ↗
      For people with any kind of budget, Opus 5.5&#x27;s [Medium] actually can make sense dollar per intelligence&#x2F;dollar per task wise. Heck, it puts some other models to shame. [Max]&#x27;s cost is completely unhinged.

      My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by &quot;value&quot; than hypothetical ceilings because I&#x27;m just not in that budget category.

      1. seabass-salmon · · focus · HN ↗
        That was true for me four weeks ago, but 2-3 weeks ago Luna turned into drivel in essentially the same complexity of task. I feel it came back somewhat in recent days but does feel like it&#x27;s being manipulated.
        1. arcanemachiner · · focus · HN ↗
          Interesting, I have noticed so such collapse.

          Have you ruled out the possibility that your system prompt, AGENTS.md, or increasing codebase complexity are not to blame?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.