‹ BackHN Continuity

Thread

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

333 points · 106 comments · theanonymousone

  1. simonw · · focus · HN ↗
    This is the page for the &quot;max&quot; reasoning setting. The page for xhigh is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh and the page for medium (the default setting) is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium

    I&#x27;ve failed twice to get &quot;Generate an SVG of a pelican riding a bicycle&quot; to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

    I&#x27;m suspicious that &quot;max&quot; may be virtually useless if it&#x27;s that easy to have it overthink to the point that it doesn&#x27;t get to a response.

    Transcript for one attempt here - expand the &quot;Reasoning trace&quot; bit to see it: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F61fd7c3683fffce9a3ab7c43d1180024#reasoning-3" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    1. zerof1l · · focus · HN ↗
      This is totally a thing I noticed myself about 3 months ago. Medium thinking effort is ideal for most tasks. At high and above, models tend to generate more output in the form of comments or code for the same problem with no real benefit. Its a self-feeding loop: more output becomes more input, which then becomes more output. High is the highest I go. If I need more intelligence, it&#x27;s better to use a more powerful model with less thinking effort or break the problem into phases. Much better result.
      1. realusername · · focus · HN ↗
        Personally I use everything in low reasoning. Maybe I&#x27;m wrong but I think that the higher reasoning settings are almost never worth it, it&#x27;s marginal gains for a much higher budget.

        I also switch to a better model for more complex tasks, also in low settings

        1. therealdrag0 · · focus · HN ↗
          Low is good if you’re working in a tight loop. But more risky for more agentic stuff you want to let cook for 30 minutes or more.
          1. realusername · · focus · HN ↗
            I always work on a tight loop, I don&#x27;t think it makes sense to let agents run for hours.

            Regardless of the model, running it for hours means that the model will takes decisions and assumptions alone instead of you.

            1. baq · · focus · HN ↗
              Yup that’s quite literally what ‘agency’ is and the whole point of agentic workflows. Personally I’ve had them running for days with good results and as you can see OpenAI had them running for months, and yes indeed the things made some very questionable decisions and assumptions… but they unquestionably did a lot of stuff correctly, for some definitions of ‘technically correct’.
              1. realusername · · focus · HN ↗
                Personally I don&#x27;t believe in agentic workflow. I don&#x27;t think that&#x27;s a coincidence that both OpenAI and Anthropic chose math problems to test their long agentic workflows, they are well defined, with a clear finish line and with a 100% clear progress path, most of real life tech projects aren&#x27;t like that.

                No matter how clever the model is, most problems have multiple valid, invalid and unclear decisions to make, running it for a long time is just picking the first option on everything, which isn&#x27;t usually what you want

                1. therealdrag0 · · focus · HN ↗
                  It’s a trade off of productivity and control.

                  It’s like learning to delegate and let go. The more senior I got the more I had to learn to let other engineers make decision i thought were suboptimal but mostly good enough. That positioned me well to be comfortable with agents. It’s contextual how much I’m willing to give them control and how much to review afterwards.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.