‹ BackHN Continuity

Thread

Claude Opus 5.5

273 points · 2 comments · throwaway371647

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. stefan_ · · focus · HN ↗
    > a `.required` tool-calling mode is a contract, so it is sent as-is and the API error names the field.

    Ah, I see nothing has changed. Are Anthropic aware that their models are generating tons of gibberish? In comparison, Astra is sublime.

  2. mupuff1234 · · focus · HN ↗
    What happened to "slowing down"?
    1. Lord_Zero · · focus · HN ↗
      The hype train must keep chuggin or it all collapses.
    2. setsewerd · · focus · HN ↗
      They're slowing down token usage, not the path to regulatory capture.
    3. WarmWash · · focus · HN ↗
      Trump got mad and investors sued.
    4. petesergeant · · focus · HN ↗
      Slowing down only makes any sense if you can coordinate a slow-down for everyone.
      1. mupuff1234 · · focus · HN ↗
        That's just false.

        Less companies involved means less pressure to go fast.

        1. solenoid0937 · · focus · HN ↗
          Ah yes, they should just cease to exist, giving OpenAI - the champions of AI safety - control over the future of humanity. Fantastic suggestion!
      2. nozzlegear · · focus · HN ↗
        [delayed]
        1. roughly · · focus · HN ↗
          Which is one of those fun things that didn’t actually exist back when we took it for granted that our fellow person was operating under some kind of moral or ethical framework, which pretty much everyone was until the economists told us that wasn’t rational, because it turns out it’s an evolutionary advantage to operate under an ethical or moral framework because it allows the kind of coordination which facilitates better collective outcomes, which everyone knew until the economists came along to tell us we were wrong and in fact it was rational not to do so and suddenly we had the prisoner’s dilemma.
          1. setsewerd · · focus · HN ↗
            On the other hand, there's research suggesting that the most optimal behavior for the best outcomes (based on the famously dependable economist style of analysis in a vacuum) is to practice the moral/ethical framework but to also engage in tit for tat - ie, assume everyone means well but respond proportionally when they don't.
    5. icrbow · · focus · HN ↗
      If you hit wall, hit it hard.
  3. prodigycorp · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;claude-opus-5-5" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;claude-opus-5-5

    &gt; Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

    Thank GOD

  4. SadErn · · focus · HN ↗

    [dead]

  5. Gander5739 · · focus · HN ↗
    Dupe: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803892">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803892
    1. unethical_ban · · focus · HN ↗
      The ~~bot~~ throwaway account beat the more established user by a minute.
  6. alpineman · · focus · HN ↗
    So we skipped 5.1, 5.2, 5.3, and 5.4: we really are plateauing
    1. voiceeh · · focus · HN ↗
      They need to catch up to OpenAI, so it makes sense to skip a few numbers.
    2. palata · · focus · HN ↗
      Oh, that explains it. I&#x27;m still on Opus 4.8 (5.0 was too annoying), and I thought I had missed a few releases...
  7. gopalv · · focus · HN ↗
    The whole thing reminds me of the Apple feature flag story[1] from a generation ago.

    [1] - <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=6372466">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=6372466

  8. meerita · · focus · HN ↗
    As long as it&#x27;s not as verbose as Opus 5, I am quite happy with a better version that&#x27;s also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.
  9. nozzlegear · · focus · HN ↗
    [delayed]
  10. nickandbro · · focus · HN ↗
    Wow! Though need to see its token efficiency to better assess. Been hearing rumors it generates much more output tokens per task.
    1. keeganpoppen · · focus · HN ↗
      my projection is that they are still gonna be pretty far behind, but they will sew it up in the next few releases. it feels like they were caught with their pants down on how much work OpenAI has put into that area, but i doubt there is some magical secret sauce that OpenAI has that Anthropic simply cannot catch up with.
  11. cogythea · · focus · HN ↗
    Interestingly they&#x27;ve changed their approach to usage resets for this release - with previous releases I&#x27;ve had my usage instantly reset, but now in the Claude app I&#x27;ve got a &#x27;Reset for free&#x27; button that expires Oct 22, which seems to effectively be a whole new usage window I can activate whenever&#x27;s convenient
    1. jdmoreira · · focus · HN ↗
      then they copied that from codex because thats exactly how codex works
    2. NielsHarksen · · focus · HN ↗
      Where in your app do you find this button?
  12. calibas · · focus · HN ↗
    &gt; We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.

    We can&#x27;t test it properly because it knows it&#x27;s being tested.

    1. johntb86 · · focus · HN ↗
      Just make it always think it&#x27;s being tested, and problem solved.
      1. actionfromafar · · focus · HN ↗
        Would it believe that?
  13. pookieinc · · focus · HN ↗
    “It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”

    They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?

    1. jbellis · · focus · HN ↗
      Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that&#x27;s less than entirely useful.
      1. meric_ · · focus · HN ↗
        Opus does seem like a more powerful coding workhorse based on the benchmarks listed though. Good coding performance, faster and less verbose, cheaper.

        Will be interesting to see how people&#x27;s opinions of it line up IRL, but so far I&#x27;ve loved Fable so hopefully will love this one too

    2. randomblock1 · · focus · HN ↗
      &gt; On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
  14. viccis · · focus · HN ↗
    So it beats Fable 5.1, by quite a bit, on every metric? Interesting.

    Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn&#x27;t use Fable 5.1 with my tier, but this is worth trying out.

    1. scrollop · · focus · HN ↗
      Why can&#x27;t they let 20usd claude subscriptions access fable in CC, as openai allows you to use astra and max modes in codex - you just pay for it in more token use.
  15. thibran · · focus · HN ↗
    Anthropic models are ridiculously expensive. I&#x27;ve stopped using any of their models months ago.
  16. slacktivism123 · · focus · HN ↗
    Opus 5.5...

        often suspects it is being evaluated, which challenges our ability to assess how it will act.
    
        comes with our watermarking measures
    
        is no longer available with &quot;thinking&quot; mode switched off
    
        stops API users from editing Claude&#x27;s prior context in an attempt to extract Claude&#x27;s reasoning.
  17. benjiro29 · · focus · HN ↗
    &gt; costs 40% less to run

    Its funny how every new model cost less to run. But this often does not match with reality.

    O, and subscriptions getting less usage, despite how the new models &quot;costs xx% less to run&quot;.

  18. alvis · · focus · HN ↗
    $0.20 vs the old $0.5 cache read is pretty much 60% off
  19. aennassiri · · focus · HN ↗
    Let&#x27;s see how much they benchmaxxed their model!
  20. ryanscio · · focus · HN ↗
    Input $4&#x2F;MTok and output $20&#x2F;MTok is a welcome surprise. Cheaper than Opus 5&#x2F;4.8, Astra 6, Fable 5.
    1. benjiro29 · · focus · HN ↗
      The biggest one is the Cache reads going from $0.50 to $0.20 ... Read&#x2F;Writes dropping by 25% but Cache reads by 60% has a much bigger impact.
  21. tag2103 · · focus · HN ↗
    Why would anyone reward bad behavior?
  22. seviu · · focus · HN ↗
    Where is the pelican test
  23. Gattopardo · · focus · HN ↗
    I wonder if we&#x27;re going to get the usual &quot;[previous model name + version] was utter dogshit, this is the new hotness&quot; comments or if people are finally becoming sick and tired of this farce.
  24. iamsyr · · focus · HN ↗
    Lower costs, lower revenue, lower profits, and slower growth is my argument correct ?
  25. jacobgold · · focus · HN ↗
    I use the other 50% of my $200&#x2F;mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don&#x27;t have to deal with Opus directly.
  26. km144 · · focus · HN ↗
    I think this release is really going to give them a hard time selling Fable:

    &gt; On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

    In general, &quot;benchmark margins have become a less reliable guide to real-world differences&quot; sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I&#x27;m not sure what to make of this admission.

    1. CPLX · · focus · HN ↗
      Opus 5 fucking sucks. Like it&#x27;s horrible. I use Fable for coding and anything important and I use Opus 4.8 for things like recursive email categorization, transaction matching, and other stuff where I don&#x27;t want to burn as much quota.

      In my experience Opus 5 is the worst of all possible worlds, it&#x27;s dumb and headstrong. It just runs away with tasks you didn&#x27;t ask it to do, is reckless, and basically is unusable in my experience.

      Not sure why but my guess is that this will be that bust worse. Happy to be proven wrong.

      1. port3000 · · focus · HN ↗
        I believe Opus 5 isn&#x27;t meant to be spoken to by humans. It&#x27;s great at executing but I reckon it&#x27;s intended to be spoken to by other models such as Fable. I use Fable as the orchestrator, only speak with Fable, and all implementation, recon, design etc happens with Opus 5, with Fable reviewing (and translating).
        1. booty · · focus · HN ↗
          That&#x27;s interesting.

          I&#x27;ve really gone in the opposite direction: having a dumber model orchestrate. In my case, it&#x27;s usually a Luna orchestrator spawning Sol&#x2F;Astra subagents to do the &quot;big brain&quot; work of planning and reviewing.

          Reason I went with &quot;dumb orchestrator&quot; was just to save tokens. Having Opus&#x2F;Sol (let alone Fable&#x2F;Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna&#x2F;Sonnet&#x2F;Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn&#x27;t just managing context properly.

      2. Syntaf · · focus · HN ↗
        Yeah if anything Opus 5 taught me how little benchmarks mean to the actual real world performance of these models.

        &quot;Better&quot; in every sense of the benchmarks and absolutely horrible results in my day-to-day work.

        The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...

        It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....

      3. cbg0 · · focus · HN ↗
        I use it frequently with a lot of success on &quot;Medium&quot; effort, it overthinks like crazy on higher levels, but YMMV.
    2. simianwords · · focus · HN ↗
      It’s likely that they have internal benchmarks but they are communicating to people who can only gauge through external benchmarks.
      1. suddenlybananas · · focus · HN ↗
        Why wouldn&#x27;t they report these benchmarks?
    3. booty · · focus · HN ↗

          &quot;benchmark margins have become a less 
          reliable guide to real-world differences&quot; 
          sounds like a big problem.
      
      My guesses:

      1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.

      2. AFAIK &quot;success&quot; in a benchmark essentially boils down to &quot;do the tests pass and do we get the right result?&quot; which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and &quot;success&quot; involves harder to measure dimensions like &quot;maintainability&quot; and &quot;did you overengineer this?&quot; and &quot;how did you cope with a bunch of vague and maybe contradictory business requirements?&quot;

      Having said all of that, I have never ever looked inside any of these benchmarks. I&#x27;m putting my guesses out here strictly in the tradition of &quot;the quickest way to learn about something is to be wrong about it on the internet.&quot;

      1. Solvyx · · focus · HN ↗

        [dead]

  27. ayhanfuat · · focus · HN ↗
    Looks like Anthropic is starting to give bank reset as well:

    &gt; Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

  28. kingstnap · · focus · HN ↗
    &gt; It’s good at finding and fixing inefficiencies in software

    Holy shit! Its happening!

    Now if we can the AI to understand this *implicitly* so that it doesn&#x27;t need to be stated upfront, we might be able to undo years of &quot;premature optimization is the root of all evil&quot;.

  29. hirako2000 · · focus · HN ↗
    [delayed]
    1. danbrooks · · focus · HN ↗
      Many people knew this announcement was coming. The betting markets suggested a very high likelihood of Opus dropping today. I was anticipating this quite a bit!
      1. b38484848 · · focus · HN ↗
        Exciting! Thanks for sharing these news.
  30. jdw64 · · focus · HN ↗
    Finally, it seems like a good time to do some &#x27;load-bearing&#x27; work on my project for a while
  31. richardjennings · · focus · HN ↗
    My 20x plan was set to end tomorrow. The writing style and insistence on word vomit just became too annoying. Is Opus 5.5 worth sticking around for ?
  32. glub · · focus · HN ↗
    &gt; For users with cybersecurity use cases that may be blocked by our cyber safeguards, we recommend accessing our models with reduced cyber blocking classifiers via our Cyber Verification Program. Claude Opus 5.5 will be available through this program in the near future.

    Anthropic has used &quot;in the near future&quot; for Mythos-class models too, but CVP is still Opus 5 only.

    Why even have the program designed for trusted access to cyber capabilities if you&#x27;re not providing access to cyber capable models via the program?

  33. somewhatjustin · · focus · HN ↗
    &gt; Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

    Nice. I was starting to think Haiku was going to be abandoned.

  34. iamsyr · · focus · HN ↗
    I don&#x27;t yet have any reason to leave Haiku 4.5 and switch to Opus 5.5.
  35. mococa · · focus · HN ↗
    I knew it. <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49801266">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49801266
  36. ricardobeat · · focus · HN ↗
    Yes, they listened! The improvement in communication style looks fantastic. I was on the verge of cancelling my subscription.
  37. xenit_v0 · · focus · HN ↗

    [dead]

  38. greenavocado · · focus · HN ↗
    Enjoy it for the next 2 weeks until its silently quanted to 4.8 level
    1. anthonyrstevens · · focus · HN ↗
      Evidence for this?
  39. jatins · · focus · HN ↗
    &gt; We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.

    Thank you.

  40. giancarlostoro · · focus · HN ↗

    [dead]

  41. keeeba · · focus · HN ↗
    Opus 5.1 came out about a month ago, what gives?
    1. velcrovan · · focus · HN ↗
      Just a guess but maybe they decided 5.1 wasn&#x27;t their last Opus model. Like they would keep developing new versions of it or something.
      1. adastra22 · · focus · HN ↗
        Crazy!
    2. hadlock · · focus · HN ↗
      Frontier model labs release some kind of update every 6 weeks on average.
    3. rs_rs_rs_rs_rs · · focus · HN ↗
      That was Fable. Last version of Opus was at the end of July.
  42. postalcoder · · focus · HN ↗
    Wow, this may be the first time I&#x27;m excited about using an Opus model in a while.

    Notables:

      - Improved ergonomics: better communication, 30% faster, and increased 5-hour limits
      - Better at sprawling jobs (Always considered codex to be superior to claude at this)
      - Pareto improvement at every thinking level
    
    That said, Opus 5 showed us that impressive benchmarks can only take us so far. Hoping we&#x27;re not in for such disappointment again.
  43. notduckrabbit · · focus · HN ↗
    They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.
    1. manmal · · focus · HN ↗
      That cost reduction seems to stem from cheaper cache reads, mostly.
  44. bredren · · focus · HN ↗
    Notes on communication:

    &quot;Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5&quot;

    and

    &quot;We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5.&quot;

    and

    &quot;In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.&quot;

    I realize it is corporate communications but &quot;most common areas of feedback&quot; and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.

    If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.

  45. tomhow · · focus · HN ↗
    Comments moved to <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803892">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803892.
  46. [deleted] · · focus · HN ↗

    [deleted]

  47. sidrag22 · · focus · HN ↗
    Cool maybe this makes the 20$ sub less of a joke. I deemed 5.0 unworthy of spending time fighting with and personally considered it by far the worst release of 2026 by either of the two major labs(promising model, but obviously not even close to ready for the general public).

    So for that 20$ tier for the entire summer and into fall, i was on their 2nd class public model(4.8) released in May. Not surprisingly it became my grunt model, doing the simple work. By far the least I&#x27;ve used Anthropic models in the last 2 years.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.