‹ BackHN Continuity

Thread

Getting the most out of Opus 5.5 in Claude and Claude Code

232 points · 156 comments · saikatsg

  1. rdli · · focus · HN ↗
    It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.

    9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.

    1. chewchewchew · · focus · HN ↗
      9 hours?!
      1. rdli · · focus · HN ↗
        Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
        1. Tade0 · · focus · HN ↗
          I dare not ask about the cost, having burned $60 on a task running for 1h 16min once.
          1. rdli · · focus · HN ↗
            I’m on the $100/month subscription; this session took about $500 in token-equivalent costs.

            (Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)

            1. atif089 · · focus · HN ↗
              Does it resume automatically on higher subscriptions?

              I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.

              1. Guillaume86 · · focus · HN ↗
                Instruct it to arm a monitor (every hour or so) to wake him up in case of quota or api issue.
                1. satvikpendem · · focus · HN ↗
                  Interesting how the French call Claude a "him" and not an "it", as French and many other languages don't have a word for a neuter pronoun.
                  1. Guillaume86 · · focus · HN ↗
                    Eh I know it's a miskate (used both here), but yeah it's a conscious effort for people at my level I guess.
              2. [deleted] · · focus · HN ↗

                [deleted]

            2. tripleee · · focus · HN ↗
              God I hope the prices drop quick. Once they stop subsidizing it these kinds of workflows will be unaffordable for anyone who isn't already wealthy
              1. satvikpendem · · focus · HN ↗
                They have dropped, in Chinese models.
              2. wyre · · focus · HN ↗
                Pricing is dropping quick. Inference is so cheap, I think they are losing a lot less money selling subscriptions than you think. It might even be more expensive managing the load, than actually selling the tokens at subscription prices.

                We are seeing with OpenAI, allegedly through their new pricing scheme, as intelligence and model efficiency increases they offer the same throughput while advertising 1/2 as much usage, letting Astra consume more usage, essentially only being available to those wealthy enough to afford it while still offering essentially unlimited Sol and Luna to their subscription tiers.

                Also if you're cache hit rate is high enough a billion tokens tokens from Deepseek 4.1 Flash costs less than $15.

              3. debatem1 · · focus · HN ↗
                Assuming this isn't some toy CI a 60% drop in billable minutes will make $500 back pretty quick. Github is wildly expensive.
                1. simon-b · · focus · HN ↗
                  The cost of the standard `actions_linux` is $0.006/minute, so spending $500 to save 6m per invocation, break-even is at ~14k invocations. But, if they're using larger machines and/or parallel jobs the $$$ saving accrues faster. IMO the wall-time saving shortening feedback loop may be a bigger win, but harder to value.
              4. miroljub · · focus · HN ↗
                Subsidies are a lie planted by Misanthropic and ClosedAI to milk their users and let them think it's the other way round.

                Inference is highly profitable business, even for third parties with much less resources and expertise.

                1. azan_ · · focus · HN ↗
                  Inference is highly profitable, but you need to recoup losses from training.
                  1. miroljub · · focus · HN ↗
                    Only if your training costs are overblown because you are bruteforcing it. Otherwise it's an equivalent of printing money.
              5. Sevii · · focus · HN ↗
                An insane amount of compute manufacturing comes online in 2028. Compute is a commodity. It's not going to stay expensive for long.
            3. cromka · · focus · HN ↗
              Curious how did you set it up like that?
          2. dan-robertson · · focus · HN ↗
            If you compare the cost to the price of dinner or whatever else you spend disposable income on, it can seem high but if you compare the cost to employing an engineer (don’t forget costs for payroll taxes, office space and equipment, etc) and consider the fact that the models often seem to be much faster than even expert humans, the costs don’t seem so terrible.
            1. Tade0 · · focus · HN ↗
              My concern is that there might come a time when this cost is passed onto employees.

              Suppose everyone starts moving faster thanks to LLMs and it becomes an expectation to use them. Budgets aren't infinite, so one of the two has to happen:

              1. People get laid off.

              2. Costs are shifted onto employees - either through lower salaries or having them bring their own subscriptions. I don't even make $500 a day!

    2. GroksBarnacles · · focus · HN ↗
      Is "wall-clock" an actual term you used before Claude? I had never heard it before the model used it and I can't stand it.
      1. woodruffw · · focus · HN ↗
        “Wall time” is a pretty common systems term (you see it when comparing total runtime to kernel time, for example). I wouldn’t have indexed on “wall clock time” or other variants as an LLMism.
      2. hazard · · focus · HN ↗
        It's a pretty old term, to distinguish from e.g. CPU time. This was in common usage even 30 years ago.

        Example from 15 years ago: <a href="https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;7335920&#x2F;what-specifically-are-wall-clock-time-user-cpu-time-and-system-cpu-time-in-uni" rel="nofollow">https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;7335920&#x2F;what-specificall...

      3. Klathmon · · focus · HN ↗
        I&#x27;ve used wall clock for many years, normally when compared to CPU time when talking about parallelizing some process.

        CPU time might go up while wall clock time goes down

      4. sdthjbvuiiijbb · · focus · HN ↗
        It&#x27;s a standard term and has been for ages. It distinguishes end to end time vs eg the amount of CPU time an individual process uses (which excludes time spent waiting for the system or time when the process was otherwise not scheduled on the CPU).
      5. kgwgk · · focus · HN ↗
        As old as time.

        <a href="https:&#x2F;&#x2F;ss64.com&#x2F;bash&#x2F;time.html" rel="nofollow">https:&#x2F;&#x2F;ss64.com&#x2F;bash&#x2F;time.html

        1. drivers99 · · focus · HN ↗
          [delayed]
      6. dolebirchwood · · focus · HN ↗
        Trying to make us feel old? Very common term among people around my age and higher (40+). Please don&#x27;t turn &quot;I&#x27;ve never heard that expression before&quot; into &quot;must be AI saying it&quot;.
      7. jghn · · focus · HN ↗
        Lolwut? You seriously have never heard anyone say this before?
      8. pertymcpert · · focus · HN ↗
        Wow...it&#x27;s super common. For example, timing a program you can get user CPU time and then wall clock time, which can be two very different numbers.
      9. saghm · · focus · HN ↗
        I&#x27;ve heard that term for years before LLMs
    3. whatsThisBtn4 · · focus · HN ↗
      Meh... There&#x27;s a reason Opus 4.6 is still an option.

      Pros know these are lower cost models.

      1. oidar · · focus · HN ↗
        I do like Opus 4.6, but I think 5.5 on medium or low is a better value. I have to steer 4.6 more and build more scaffolding around the tasks. 5.5 just does what I ask. Visual spatial reasoning greatly improved in 5.5 as well.
        1. whatsThisBtn4 · · focus · HN ↗
          I pass costs to my customers, so it doesn&#x27;t really matter. They are getting 30k in value for $1000.
    4. tamimio · · focus · HN ↗
      It’s great, but you still need to know what you are doing, not just the goal&#x2F;s. I built a platform years ago from scratch, and now I am remaking it with more features and more polished design, I know exactly what needs to be done to tiniest details. The first prompt was very well detailed about the architecture and how everything should work, after an hour work at Xhigh, it did create the blueprint artifacts that I asked for, then I spent 3 days reading every single thing and writing notes, turned out it made the system overly complicated without adding extra value, plus I can see how some of the architecture design will be potentially a security risk. So after few days I fed my notes, this time took 4hours and 1M tokens! Later I spent more few days reviewing and writing notes, it was closer to what I want but still made architecture errors, the third run took around an hour and finally made it how it supposed to be, although there are still more notes on non critical stuff. So I don’t think we are yet at the stage where sitting goals and some high level is enough to produce quality results.
    5. slaser79 · · focus · HN ↗
      Agree Opus 5.5 is incredible and efficient with Claude Max Usage.What a jump after the writing slop you got from Opus 5. I had moved to using Fable for my orchestration workflow mainly due to the communication issue (I think Opus 5 was capable enough but spoke in riddles so you lost confidence quickly)..With 5.5 it needs less steering now and communicates well, and I have had the same CC session running for the last week (obviously compacting with durable plans etc as the post), with it PM&#x27;ing my home built agent orchestration of the other coding agents(antigravity, codex, pi etc) and making decent decisions and all the recommendations are normally usually good.
    6. Betelbuddy · · focus · HN ↗
      &gt;&gt; It’s a really good model.

      In the meantime, I have cancelled my Anthropic subscription...

      I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.

      I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.

      Then I ask each model to review and critique the other proposals.

      By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.

      Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...

      1. karp773 · · focus · HN ↗
        It&#x27;s not even funny any more. Chinese model, Chinese vendor, Chinese, Chinese... Did I say Chinese? Chinese!

        Nobody in his right mind will use a Chinese clone when you have models like Opus 5.5 for peanuts.

        1. verdverm · · focus · HN ↗
          &gt; Nobody in his right mind will...

          let a few valley elites decide how humanity can use this technology

          open and transparent is the way, China is showing how

          1. christophilus · · focus · HN ↗
            I’m rooting for open models, but SOL 6.1 and Opus 5.5 are absolute workhorses on a $100&#x2F;mo sub. I share your fears, though, and really hope an open model catches up and can somehow compete with the subscription prices of the big 2.
            1. verdverm · · focus · HN ↗
              I have workhorse models, spend far less, the model matters less than people like to claim

              there&#x27;s no money long term in being a token vendor

              1. K0balt · · focus · HN ↗
                That sounds interesting.

                I spin up “offices” for different projects, using a documentation heavy approach with procedures, policies, standards, and processes. Agent onboarding and orientation, etc. I usually have an engineer for each separate part, (one for a simulator to simulate the hardware, one for the user application, one for the data analysis and evaluation tool, one for the firmware on each type of device, one for schematic and board reviews, etc. ) then I’ll have an office manager in charge of policy and issue boards, agent rosters, etc, and a engineering governance agent that makes sure code is compliant and documentation &#x2F; code is coherent before any merges. 6-15 agents in each office depending on the complexity of the task.

                It sounds like open code could be pretty handy but it would nerf my Claude subscription (api!=subscription rates). Understanding my workflow, what models do you think might be suitable for those tasks outside of OAI and Anthropic?

                1. verdverm · · focus · HN ↗
                  I&#x27;ve outlined the models I&#x27;ve used in a recent HN comment, check my history, and some spicy opinions too :]
            2. verdverm · · focus · HN ↗
              I&#x27;d be remiss if I did not point out subscription plans like OpenCode Go, with generous quotas at $10&#x2F;40 month (you can have more than one), which is a way better deal than anything you&#x27;ll find with Big Ai
        2. Betelbuddy · · focus · HN ↗
          Yes ...I was so impressed I cancelled my subscription. I could not stand all the winning. I offer cheap hourly rates of $1000 for debugging AI slop.

          Contact me at : prompt.plumber@gmail.com

        3. OtomotO · · focus · HN ↗
          &gt; Nobody in his right mind will use

          Nobody AMERICAN in his right mind will use... Wait, actually a lot of them will.

          But for me, as a non american, non chinese person: I&#x27;ll use whatever the fuck is the best and cheapest for my task, because that&#x27;s how fucking Capitalism works.

          If that means that a (proclaimed) &quot;communist&quot; country cleans the carpet with the self-proclaimed land of the free: so be it!

          1. nozzlegear · · focus · HN ↗
            [delayed]
        4. Betelbuddy · · focus · HN ↗

          [dead]

    7. rdli · · focus · HN ↗
      I would also say: I think this works great because the goal is well-defined and measurable.

      I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.

    8. Waterluvian · · focus · HN ↗
      Every time weird stuff happened this week it was because the model choice in VSCode got set back to “auto” and some other model was trying its best. 5.5 is what I just set it to. Even Fable feels worse for my use cases.
      1. AnotherGoodName · · focus · HN ↗
        Even Anthropic rates Fable lower than 5.5 on pretty much all benchmarks.

        &quot;Why does Fable even exist&quot; is a very very reasonable question right now.

        1. wyre · · focus · HN ↗
          Because Anthropic releases their different model level&#x27;s at very different points in time, they seem to always have one model that is by far an away the best to use for everything. Haiku 4.5 is almost a year old. Sonnet is fine, but idk if it has any real benefits over Opus. It&#x27;s only been since Fable has been released that you get to choose between Fable and Opus, but not with 5.5 there is no reason to use Fable.

          I feel like instead of releasing fable, they should have released it as Opus 5, then their next Opus release they would call Sonnet, and their next Sonnet release they would have called Haiku. I don&#x27;t know if their pricing structure would have been able to support that, but Anthropic has always been the least competitive regarding token pricing.

          1. comradesmith · · focus · HN ↗
            Fable came with a new set of api prices, and a special allowance for subscriptions.

            If they did what you suggested either they eat a ton of additional costs, or send a signal to the market that they’re increasing costs more generally.

            Also Fable and Opus have different specialties so they really are best presented as different models

        2. satvikpendem · · focus · HN ↗
          It&#x27;s a class of model not a static one. There&#x27;ll be Fable 5.5 that&#x27;s even better than Opus 5.5.
          1. Wowfunhappy · · focus · HN ↗
            Although we never had a Claude Opus 5.1. Fable went 5 → 5.1 and Opus went 5 → 5.5.
            1. satvikpendem · · focus · HN ↗
              Opus 5.5 is probably a distilled version of a 6 model.
              1. Wowfunhappy · · focus · HN ↗
                If they had 6, don&#x27;t you think they&#x27;d be serving it? Even if it was too expensive for most users, some big companies would likely be interested.
                1. satvikpendem · · focus · HN ↗
                  &gt; If they had 6, don&#x27;t you think they&#x27;d be serving it?

                  No I do not. OpenAI has some hidden model that&#x27;s apparently 5x or so better at certain benchmarks than GPT 6 but they&#x27;re not and have no plans to release it. It is increasingly likely that these AI companies keep the best models for themselves and then release smaller, cheaper distilled models for everyone else, especially since the AI companies are vertically integrating into many fields.

                2. K0balt · · focus · HN ↗
                  You’re assuming the goal of the big AI companies is to serve models for money. That is not the goal. They have a lot of incentives not to share their best models.
        3. adastra22 · · focus · HN ↗
          You are making the mistake of classifying models on a single linear axis, or even a multi axis basis set of all benchmarks. That just isn’t true. Each model is unique in its skills and capabilities and the way it approaches problems, in a way that is not represented in benchmarks. Fable is better at reviewing things. I don’t know how to explain it well but it is true. I trust Fable to do thorough reviews (sometimes too thorough) and to present its information in a dense but ordered way. Its output is equivalent to what you used to get from security firms doing code reviews. Having Opus do the work, and have fable do reviews (of the plan and implementation) is a good combo.
        4. sutterd · · focus · HN ↗
          I find fable better at high level planning, as in planning a task without filling in all the details. Opus doesn’t seem to be great at this, but other models aren’t either.
          1. BobbyJo · · focus · HN ↗
            I find the same. Fable is better at hashing out a plan with some back and forth and opus 5.5 is better (cheaper certainly) at sterile execution.
            1. cromka · · focus · HN ↗
              My exact experience as well.
        5. cromka · · focus · HN ↗
          I think of Fable as more knowledgeable, smart, erudite. Opus is a skilled technician.

          Benchmarks do not measure the first aspect.

    9. AnotherGoodName · · focus · HN ↗
      To back this up we have a discord chat for the board game terraforming mars where the agent takes input and vibe codes an open source implementation of the game.

      <a href="https:&#x2F;&#x2F;tfmbot.com" rel="nofollow">https:&#x2F;&#x2F;tfmbot.com is the link (discord and source links on the splash screen).

      The results are fucking incredible to the point where people in discord are stating &quot;I&#x27;m surprised this is working so well&quot;. I am too.

      I feel like there&#x27;s a group online that missed the boat. Anything negative towards AI capabilities is still upvoted but I&#x27;ve been in the industry for over 25years, highly respected and can&#x27;t fathom the &quot;AI dumb lololol&quot; type of comments i see on HN. AI is superseeding all other ways to develop.

      1. TeMPOraL · · focus · HN ↗
        I&#x27;m restoring a game I played as a kid, that I couldn&#x27;t reliably get to even start on modern Windows. The skill and speed with which Opus 5.5 got it running, while patching a bunch of bugs in the binary along the way (only some of them I knew about), has my jaw still on the floor - and few hours in, I already have whole campaign mapped out as state machine graph, and we&#x27;re upscaling graphics now.
      2. sashank_1509 · · focus · HN ↗
        Pretty sure no one on HN says AI dumb
        1. AnotherGoodName · · focus · HN ↗
          From just this morning i read this thread: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49946321">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49946321 about a linux distro no longer allowing AI code and the comments are mostly along the lines of &quot;This will be great for maintainability, AI can&#x27;t code any complexity&quot; which is essentially what I&#x27;m getting at here.

          AI writes clean code and can do so in a very maintainable way honestly.

          1. tripleee · · focus · HN ↗
            One of the last things developers had to offer was being able to guide the AI to write maintainable code. Now that it can do that there&#x27;s not a whole lot left

            It&#x27;s understandable why that&#x27;s hard to accept

      3. CoolestBeans · · focus · HN ↗
        I just see the extremes. You either see people unable to recreate results and them calling people idiots for claiming those results. Or you see people saying that AI will supersede all other ways to develop and calling anyone who doesn&#x27;t full embrace AI an idiot. Reality is that nobody knows nothing. There are a million factors that could cause the end result to be anywhere between both extremes. I don&#x27;t know, you don&#x27;t know, AI doesn&#x27;t know, least of all the people inside the AI companies don&#x27;t know. And really the end result will be extremely nuanced I&#x27;m sure.
      4. tripleee · · focus · HN ↗
        One of the last things developers had to offer was being able to guide the AI to write maintainable code. Now that it can do that there&#x27;s not a whole lot left It&#x27;s understandable why that&#x27;s hard to accept
    10. slopinthebag · · focus · HN ↗
      this sounds great until you dig into it and realize it achieved those numbers through cheating
      1. weird-eye-issue · · focus · HN ↗
        Not necessarily. I&#x27;ve had it help speed up relatively complex code by profiling it and then adding in parallelization and caching where appropriate
    11. truejaian · · focus · HN ↗

      [dead]

    12. moltar · · focus · HN ↗
      But was the code of high quality or a mess?

      Because I noticed people at my work did similar requests to improve CI. The result was a faster CI, but full of cludges huge inline bash scripts in workflow YAML files, and effectively unmaintainable, unreviewable mess. After just a few rounds of these optimizations the entire CI setup is basically a Rube Goldberg machine but made of duct tape.

      1. delusional · · focus · HN ↗
        That&#x27;s always the question. Any mid-level engineer who comes into a CI setup can trivially spot several inefficiencies, that they could solve. They also know that solving those requires implementing some specific tweaks that won&#x27;t generalize outwards from that project, or will require extra maintenance. Which they will now be the only ones aware of. 6 minutes in CI time is rarely worth the trade-off.
      2. Shorn · · focus · HN ↗
        &gt; Less than an hour of my attention.

        There&#x27;s your answer right there, no need to ask.

        Without human advice and a metric shit-tonne of guidance (either pre or post) - anything that requires active consideration of future maintenance burden will cost you 2x-10x of the time you think &quot;saved&quot; over a medium-term time-frame.

        The alternative is to spend at least 50% (or sometimes up to %100) of the &quot;saved&quot; time in active review and iterative improvement of the solution design.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.