Getting the most out of Opus 5.5 in Claude and Claude Code
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Getting the most out of Opus 5.5 in Claude and Claude Code
Unofficial Hacker News client; not affiliated with Y Combinator.
rdli · · focus · HN ↗
9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
chewchewchew · · focus · HN ↗
rdli · · focus · HN ↗
Tade0 · · focus · HN ↗
rdli · · focus · HN ↗
(Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)
atif089 · · focus · HN ↗
I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.
Guillaume86 · · focus · HN ↗
satvikpendem · · focus · HN ↗
Guillaume86 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
tripleee · · focus · HN ↗
satvikpendem · · focus · HN ↗
wyre · · focus · HN ↗
We are seeing with OpenAI, allegedly through their new pricing scheme, as intelligence and model efficiency increases they offer the same throughput while advertising 1/2 as much usage, letting Astra consume more usage, essentially only being available to those wealthy enough to afford it while still offering essentially unlimited Sol and Luna to their subscription tiers.
Also if you're cache hit rate is high enough a billion tokens tokens from Deepseek 4.1 Flash costs less than $15.
debatem1 · · focus · HN ↗
simon-b · · focus · HN ↗
miroljub · · focus · HN ↗
Inference is highly profitable business, even for third parties with much less resources and expertise.
azan_ · · focus · HN ↗
miroljub · · focus · HN ↗
Sevii · · focus · HN ↗
cromka · · focus · HN ↗
dan-robertson · · focus · HN ↗
Tade0 · · focus · HN ↗
Suppose everyone starts moving faster thanks to LLMs and it becomes an expectation to use them. Budgets aren't infinite, so one of the two has to happen:
1. People get laid off.
2. Costs are shifted onto employees - either through lower salaries or having them bring their own subscriptions. I don't even make $500 a day!
GroksBarnacles · · focus · HN ↗
woodruffw · · focus · HN ↗
hazard · · focus · HN ↗
Example from 15 years ago: <a href="https://stackoverflow.com/questions/7335920/what-specifically-are-wall-clock-time-user-cpu-time-and-system-cpu-time-in-uni" rel="nofollow">https://stackoverflow.com/questions/7335920/what-specificall...
Klathmon · · focus · HN ↗
CPU time might go up while wall clock time goes down
sdthjbvuiiijbb · · focus · HN ↗
kgwgk · · focus · HN ↗
<a href="https://ss64.com/bash/time.html" rel="nofollow">https://ss64.com/bash/time.html
drivers99 · · focus · HN ↗
dolebirchwood · · focus · HN ↗
jghn · · focus · HN ↗
pertymcpert · · focus · HN ↗
saghm · · focus · HN ↗
whatsThisBtn4 · · focus · HN ↗
Pros know these are lower cost models.
oidar · · focus · HN ↗
whatsThisBtn4 · · focus · HN ↗
tamimio · · focus · HN ↗
slaser79 · · focus · HN ↗
Betelbuddy · · focus · HN ↗
In the meantime, I have cancelled my Anthropic subscription...
I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.
I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.
Then I ask each model to review and critique the other proposals.
By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.
Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...
karp773 · · focus · HN ↗
Nobody in his right mind will use a Chinese clone when you have models like Opus 5.5 for peanuts.
verdverm · · focus · HN ↗
let a few valley elites decide how humanity can use this technology
open and transparent is the way, China is showing how
christophilus · · focus · HN ↗
verdverm · · focus · HN ↗
there's no money long term in being a token vendor
K0balt · · focus · HN ↗
I spin up “offices” for different projects, using a documentation heavy approach with procedures, policies, standards, and processes. Agent onboarding and orientation, etc. I usually have an engineer for each separate part, (one for a simulator to simulate the hardware, one for the user application, one for the data analysis and evaluation tool, one for the firmware on each type of device, one for schematic and board reviews, etc. ) then I’ll have an office manager in charge of policy and issue boards, agent rosters, etc, and a engineering governance agent that makes sure code is compliant and documentation / code is coherent before any merges. 6-15 agents in each office depending on the complexity of the task.
It sounds like open code could be pretty handy but it would nerf my Claude subscription (api!=subscription rates). Understanding my workflow, what models do you think might be suitable for those tasks outside of OAI and Anthropic?
verdverm · · focus · HN ↗
verdverm · · focus · HN ↗
Betelbuddy · · focus · HN ↗
Contact me at : prompt.plumber@gmail.com
OtomotO · · focus · HN ↗
Nobody AMERICAN in his right mind will use... Wait, actually a lot of them will.
But for me, as a non american, non chinese person: I'll use whatever the fuck is the best and cheapest for my task, because that's how fucking Capitalism works.
If that means that a (proclaimed) "communist" country cleans the carpet with the self-proclaimed land of the free: so be it!
nozzlegear · · focus · HN ↗
Betelbuddy · · focus · HN ↗
[dead]
rdli · · focus · HN ↗
I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.
Waterluvian · · focus · HN ↗
AnotherGoodName · · focus · HN ↗
"Why does Fable even exist" is a very very reasonable question right now.
wyre · · focus · HN ↗
I feel like instead of releasing fable, they should have released it as Opus 5, then their next Opus release they would call Sonnet, and their next Sonnet release they would have called Haiku. I don't know if their pricing structure would have been able to support that, but Anthropic has always been the least competitive regarding token pricing.
comradesmith · · focus · HN ↗
If they did what you suggested either they eat a ton of additional costs, or send a signal to the market that they’re increasing costs more generally.
Also Fable and Opus have different specialties so they really are best presented as different models
satvikpendem · · focus · HN ↗
Wowfunhappy · · focus · HN ↗
satvikpendem · · focus · HN ↗
Wowfunhappy · · focus · HN ↗
satvikpendem · · focus · HN ↗
No I do not. OpenAI has some hidden model that's apparently 5x or so better at certain benchmarks than GPT 6 but they're not and have no plans to release it. It is increasingly likely that these AI companies keep the best models for themselves and then release smaller, cheaper distilled models for everyone else, especially since the AI companies are vertically integrating into many fields.
K0balt · · focus · HN ↗
adastra22 · · focus · HN ↗
sutterd · · focus · HN ↗
BobbyJo · · focus · HN ↗
cromka · · focus · HN ↗
cromka · · focus · HN ↗
Benchmarks do not measure the first aspect.
AnotherGoodName · · focus · HN ↗
<a href="https://tfmbot.com" rel="nofollow">https://tfmbot.com is the link (discord and source links on the splash screen).
The results are fucking incredible to the point where people in discord are stating "I'm surprised this is working so well". I am too.
I feel like there's a group online that missed the boat. Anything negative towards AI capabilities is still upvoted but I've been in the industry for over 25years, highly respected and can't fathom the "AI dumb lololol" type of comments i see on HN. AI is superseeding all other ways to develop.
TeMPOraL · · focus · HN ↗
sashank_1509 · · focus · HN ↗
AnotherGoodName · · focus · HN ↗
AI writes clean code and can do so in a very maintainable way honestly.
tripleee · · focus · HN ↗
It's understandable why that's hard to accept
CoolestBeans · · focus · HN ↗
tripleee · · focus · HN ↗
slopinthebag · · focus · HN ↗
weird-eye-issue · · focus · HN ↗
truejaian · · focus · HN ↗
[dead]
moltar · · focus · HN ↗
Because I noticed people at my work did similar requests to improve CI. The result was a faster CI, but full of cludges huge inline bash scripts in workflow YAML files, and effectively unmaintainable, unreviewable mess. After just a few rounds of these optimizations the entire CI setup is basically a Rube Goldberg machine but made of duct tape.
delusional · · focus · HN ↗
Shorn · · focus · HN ↗
There's your answer right there, no need to ask.
Without human advice and a metric shit-tonne of guidance (either pre or post) - anything that requires active consideration of future maintenance burden will cost you 2x-10x of the time you think "saved" over a medium-term time-frame.
The alternative is to spend at least 50% (or sometimes up to %100) of the "saved" time in active review and iterative improvement of the solution design.