‹ BackHN Continuity

Thread

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

333 points · 106 comments · theanonymousone

  1. simonw · · focus · HN ↗
    This is the page for the &quot;max&quot; reasoning setting. The page for xhigh is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-xhigh and the page for medium (the default setting) is <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5-medium

    I&#x27;ve failed twice to get &quot;Generate an SVG of a pelican riding a bicycle&quot; to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

    I&#x27;m suspicious that &quot;max&quot; may be virtually useless if it&#x27;s that easy to have it overthink to the point that it doesn&#x27;t get to a response.

    Transcript for one attempt here - expand the &quot;Reasoning trace&quot; bit to see it: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F61fd7c3683fffce9a3ab7c43d1180024#reasoning-3" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    1. RGS1811 · · focus · HN ↗
      &quot;This is a classic test request...&quot;

      I know there&#x27;s been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

      1. simonw · · focus · HN ↗
        See here for more discussion of that: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803892#49804881">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803892#49804881
        1. dgellow · · focus · HN ↗
          Just want to say: you’re such a legend, please do not stop sharing your pelicans, it’s always fun to see how they change over the months :)
      2. cubefox · · focus · HN ↗
        The model recognizing the task doesn&#x27;t mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text.
      3. croemer · · focus · HN ↗
        Of course it was exposed - not sure it&#x27;s explicit or not. Why wouldn&#x27;t HackerNews comments be part of the training data? And Simon&#x27;s blog and the many discussions about Pelicans? It&#x27;d be hard to miss. Doesn&#x27;t mean Anthropic has made this an explicit goal in training.
      4. 0x10ca1h0st · · focus · HN ↗
        Lets start frog riding motorcycle trend until they frogmaxx, or cat driving convertible.
        1. agar · · focus · HN ↗
          At least pick something that will result in a good name:

          &quot;Create an SVG of Shaquille O&#x27;Neal eating potato chips shaped like a telecopier.&quot;

          Shaq&#x27;sFaxSnacksMaxx

      5. pgwhalen · · focus · HN ↗
        It would be genuinely shocking at this point if any of the frontier models weren&#x27;t well exposed to the problem.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.