‹ BackHN Continuity

Thread

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

333 points · 106 comments · theanonymousone

  1. simonw · · focus · HN ↗
    This is the page for the &quot;max&quot; reasoning setting. The page for xhigh is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh and the page for medium (the default setting) is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium

    I&#x27;ve failed twice to get &quot;Generate an SVG of a pelican riding a bicycle&quot; to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

    I&#x27;m suspicious that &quot;max&quot; may be virtually useless if it&#x27;s that easy to have it overthink to the point that it doesn&#x27;t get to a response.

    Transcript for one attempt here - expand the &quot;Reasoning trace&quot; bit to see it: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F61fd7c3683fffce9a3ab7c43d1180024#reasoning-3" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    1. zerof1l · · focus · HN ↗
      This is totally a thing I noticed myself about 3 months ago. Medium thinking effort is ideal for most tasks. At high and above, models tend to generate more output in the form of comments or code for the same problem with no real benefit. Its a self-feeding loop: more output becomes more input, which then becomes more output. High is the highest I go. If I need more intelligence, it&#x27;s better to use a more powerful model with less thinking effort or break the problem into phases. Much better result.
      1. zozbot234 · · focus · HN ↗
        This version of Opus &quot;max&quot; apparently has even higher thinking output than Qwen &quot;max&quot;, which is infamous for its thinking streams where it constantly second-guesses itself, then third-guesses, fourth-guesses and generally nth-guesses itself for arbitrarily large n. Of course, we aren&#x27;t actually seeing Claude&#x27;s raw thinking output: all we get is the after-the-fact prettified &quot;summary&quot;. One wonders how much of that is a coincidence, or whether there&#x27;s a reason behind that.
        1. judge2020 · · focus · HN ↗
          &gt; Of course, we aren&#x27;t actually seeing Claude&#x27;s raw thinking output: all we get is the after-the-fact prettified &quot;summary&quot;. One wonders how much of that is a coincidence, or whether there&#x27;s a reason behind that.

          Most of what I&#x27;ve heard is that raw reasoning traces are really good for distillation, although no idea how much the summarization actually hurts distillation.

          1. Barbing · · focus · HN ↗
            I wonder how many prompts you can send asking it to think step-by-step before they cut you off. Trying to get the reasoning traces into the body of the response, essentially. Or maybe that’s been effectively nerfed somehow. Or is not very useful.
          2. hypfer · · focus · HN ↗
            That is true, because they&#x27;re good for actually understanding wtf the model is doing.

            I&#x27;d argue that they&#x27;re a necessity if you want to use the LLM as a tool instead of a black box that just does stuff for you.

            It gives you a lot finer control over where the solution ends up when you can follow along the thinking trace and modulate your inputs based on what you saw in there.

            And, additionally, it gives you a lot more understanding of what the model can or cannot do. Strengths and weaknesses and all that.

            Using claude is like buying a car where you cannot legally open the hood. It tells you that there is something specific under there, and often it actually drives like that too, but how exactly it looks you will never see.

            For some people this is fine. I do not think that these people will survive. Figuratively speaking but also literally speaking.

            World&#x27;s changing. Opaque abstraction like that is a luxury depending on (geo)political stability.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.