‹ BackHN Continuity

Thread

Getting the most out of Opus 5.5 in Claude and Claude Code

232 points · 156 comments · saikatsg

  1. rdli · · focus · HN ↗
    It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.

    9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.

    1. chewchewchew · · focus · HN ↗
      9 hours?!
      1. rdli · · focus · HN ↗
        Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
        1. Tade0 · · focus · HN ↗
          I dare not ask about the cost, having burned $60 on a task running for 1h 16min once.
          1. rdli · · focus · HN ↗
            I’m on the $100/month subscription; this session took about $500 in token-equivalent costs.

            (Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)

            1. atif089 · · focus · HN ↗
              Does it resume automatically on higher subscriptions?

              I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.

              1. Guillaume86 · · focus · HN ↗
                Instruct it to arm a monitor (every hour or so) to wake him up in case of quota or api issue.
                1. satvikpendem · · focus · HN ↗
                  Interesting how the French call Claude a "him" and not an "it", as French and many other languages don't have a word for a neuter pronoun.
                  1. Guillaume86 · · focus · HN ↗
                    Eh I know it's a miskate (used both here), but yeah it's a conscious effort for people at my level I guess.
              2. [deleted] · · focus · HN ↗

                [deleted]

            2. tripleee · · focus · HN ↗
              God I hope the prices drop quick. Once they stop subsidizing it these kinds of workflows will be unaffordable for anyone who isn't already wealthy
              1. satvikpendem · · focus · HN ↗
                They have dropped, in Chinese models.
              2. wyre · · focus · HN ↗
                Pricing is dropping quick. Inference is so cheap, I think they are losing a lot less money selling subscriptions than you think. It might even be more expensive managing the load, than actually selling the tokens at subscription prices.

                We are seeing with OpenAI, allegedly through their new pricing scheme, as intelligence and model efficiency increases they offer the same throughput while advertising 1/2 as much usage, letting Astra consume more usage, essentially only being available to those wealthy enough to afford it while still offering essentially unlimited Sol and Luna to their subscription tiers.

                Also if you're cache hit rate is high enough a billion tokens tokens from Deepseek 4.1 Flash costs less than $15.

              3. debatem1 · · focus · HN ↗
                Assuming this isn't some toy CI a 60% drop in billable minutes will make $500 back pretty quick. Github is wildly expensive.
                1. simon-b · · focus · HN ↗
                  The cost of the standard `actions_linux` is $0.006/minute, so spending $500 to save 6m per invocation, break-even is at ~14k invocations. But, if they're using larger machines and/or parallel jobs the $$$ saving accrues faster. IMO the wall-time saving shortening feedback loop may be a bigger win, but harder to value.
              4. miroljub · · focus · HN ↗
                Subsidies are a lie planted by Misanthropic and ClosedAI to milk their users and let them think it's the other way round.

                Inference is highly profitable business, even for third parties with much less resources and expertise.

                1. azan_ · · focus · HN ↗
                  Inference is highly profitable, but you need to recoup losses from training.
                  1. miroljub · · focus · HN ↗
                    Only if your training costs are overblown because you are bruteforcing it. Otherwise it's an equivalent of printing money.
              5. Sevii · · focus · HN ↗
                An insane amount of compute manufacturing comes online in 2028. Compute is a commodity. It's not going to stay expensive for long.
            3. cromka · · focus · HN ↗
              Curious how did you set it up like that?
          2. dan-robertson · · focus · HN ↗
            If you compare the cost to the price of dinner or whatever else you spend disposable income on, it can seem high but if you compare the cost to employing an engineer (don’t forget costs for payroll taxes, office space and equipment, etc) and consider the fact that the models often seem to be much faster than even expert humans, the costs don’t seem so terrible.
            1. Tade0 · · focus · HN ↗
              My concern is that there might come a time when this cost is passed onto employees.

              Suppose everyone starts moving faster thanks to LLMs and it becomes an expectation to use them. Budgets aren't infinite, so one of the two has to happen:

              1. People get laid off.

              2. Costs are shifted onto employees - either through lower salaries or having them bring their own subscriptions. I don't even make $500 a day!

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.