‹ BackHN Continuity

Thread

Claude Opus 5.5

1806 points · 1134 comments · km144

  1. GodelNumbering · · focus · HN ↗
    Finally that price drop

       Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
       Cache reads              $0.20              $0.50
       Input tokens             $4                 $5
       Output tokens            $20                $25
       Cache writes             $5                 $6.25
    
    
    Opus 5 is the model with highest spend on openrouter (<a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings#task-spend" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings#task-spend) and it seems plausible that Opus 5 is&#x2F;was the highest spend model in the world, and certainly Anthropic&#x27;s biggest moneymaker.

    If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor

    1. AJ007 · · focus · HN ↗
      It is only a price drop if price * tokens used is less
      1. mcintyre1994 · · focus · HN ↗
        They&#x27;re claiming a drop in token use too, and that it nets to 40% cheaper.
        1. drbscl · · focus · HN ↗
          Unfortunately, they&#x27;re full of it <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#token-use" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#token-u...

          It does work out to be a similar cost per task though

          1. jsnell · · focus · HN ↗
            You should probably look at the cost&#x2F;score graph by effort level instead:

            <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#intelligence-comparisons" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#intelli...

            It is most of the pareto frontier.

            1. drbscl · · focus · HN ↗
              Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort
              1. 93po · · focus · HN ↗
                Is verboseness the only measure of token efficiency towards overall task completion?
              2. persedes · · focus · HN ↗
                so don&#x27;t use it at max? The benchmarks suggest that high&#x2F;xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I&#x27;d treat that as an outlier and not how verbose the model is in general (QED I know)
                1. drbscl · · focus · HN ↗
                  You’re missing my point. I’m saying anthropic are exaggerating their results.
                  1. persedes · · focus · HN ↗
                    how are they exaggerating the results? Comparing the cost from that chart for 5 and 5.5 for medium-max effort paints a pretty clear picture:

                             mean  median
                     model
                     5      4.135   4.245
                     5.5    3.150   2.640
                    
                    
                    Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.
          2. naasking · · focus · HN ↗
            I don&#x27;t think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:

            <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=claude-opus-5-5-high%2Cclaude-opus-5-high#token-use" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=...

          3. piotrdz · · focus · HN ↗
            Disagree. Our internal company tests showed a cost per task drop from 0.35usd to 0.16usd . Opus 5low vs opus 5.5 low
            1. johnbellone · · focus · HN ↗
              Very fast you were.
              1. epolanski · · focus · HN ↗
                Even created an account to tell us just that.
                1. piotrdz · · focus · HN ↗
                  Yep, moving here from reddit. Thanks for constructive discussion.
              2. piotrdz · · focus · HN ↗
                Yes, I need 30 min to run the test suite with new model. Why the sparky comment
          4. nl · · focus · HN ↗
            5.5 is higher for max effort, slightly higher for xhigh and lower for high, medium and low effort.

            The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).

            I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don&#x27;t think most people will notice. At high effort I think it looks like it will be an improvement for most people.

            5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.

            <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=claude-opus-5-5%2Cclaude-opus-5-5-xhigh%2Cclaude-opus-5-5-high%2Cclaude-opus-5%2Cclaude-opus-5-xhigh%2Cclaude-opus-5-high%2Cclaude-opus-5-5-medium%2Cclaude-opus-5-5-low%2Cclaude-opus-5-medium%2Cclaude-opus-5-low#intelligence-index-token-use-tabs" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=...

        2. make3 · · focus · HN ↗
          parent means that they could get more client &#x2F; a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
        3. meerita · · focus · HN ↗
          I tested it with Claude Code, and I can confirm it&#x27;s way cheaper, better, faster and less verbose than Opus 5.
      2. gwd · · focus · HN ↗
        Happened to be testing a &quot;review patches on a mailing list&quot; harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:

        Opus 5.5: Found 8&#x2F;14 issues. Total cost: $15.40

        Fable 5.1: Found 7&#x2F;14 issues. Total cost: $66.34

        Opus 5: Found 6&#x2F;14 issues. Total cost: $15.19

        Sonnet 5: Found 2&#x2F;14 issues. Total cost: $19.15

        This is a relatively small sample size, but it was both the best and the cheapest.

        ETA: NB this is &quot;Equivalent API&quot; cost as reported by claude&#x27;s CLI; I was using my subscription.

        1. retinaros · · focus · HN ↗
          I doubt your test if you cant even notice that it is not the cheapest with just 4 numbers to compare.
        2. roflc0ptic · · focus · HN ↗
          Heh I’ve been doing almost the same - back testing against PR comments - and opus 5.5 matched fable 5.1. About $24 for 10 PRs.

          I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence

        3. rmunn · · focus · HN ↗
          I just told Opus 5.5 &quot;Perform a code review on the current branch&quot; to see what it would come up with. The results were not inspiring. It told me there were five issues, one of which was a test-coverage gap on line 848 of ProjectTemplateTests.cs. But ProjectTemplateTests.cs is only 160 lines long.

          I told it that it had made a mistake in the line number, and to double-check all the line numbers. It responded &quot;You were right to push on this: four of the five line numbers were wrong, and while checking them I found two findings that were overstated.&quot;

          Then I noticed in the corner of the Claude CLI UI that it was showing &quot;Effort: medium&quot;. I&#x27;m pretty sure I had set it to high effort before; I don&#x27;t know when it reverted to medium, but that&#x27;s another thing that doesn&#x27;t exactly fill me with confidence.

          I&#x27;ll try again on high effort to see if it does better, but so far I am not impressed with Opus 5.5 on my first day of using it.

          1. gwd · · focus · HN ↗
            My prompts are moving in the other direction as sashiko [1], a managed pipeline developed for the Linux Kernel mailing list like a year ago. But last year&#x27;s models needed a lot more structure and guidance; the results I posted are from the &quot;single prompt&quot; version of the same thing. The README [2] describes the difference. You can browse the contents to get an idea; basically all the prompts were actually written and iterated by Fable (and now Opus 5.5), seeing how agents failed the tests and improving them.

            [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;sashiko-dev&#x2F;sashiko" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;sashiko-dev&#x2F;sashiko

            [2] <a href="https:&#x2F;&#x2F;gitlab.com&#x2F;xen-project&#x2F;people&#x2F;gdunlap&#x2F;xen-review-prompts" rel="nofollow">https:&#x2F;&#x2F;gitlab.com&#x2F;xen-project&#x2F;people&#x2F;gdunlap&#x2F;xen-review-pro...

        4. kphorn · · focus · HN ↗
          Good data and goes to show that Fable is melting the GPUs and is priced accordingly. I&#x27;d guess that cost to serve for Opus 5.5 is meaningfully lower through architecture advances
        5. gwd · · focus · HN ↗
          UPDATE: Sorry, just noticed I typed in the Opus 5 total cost wrong -- it should be $58.19. Main point &quot;best and cheapest&quot; was from the actual numbers, not my typo.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.