I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
For simple tasks that require a lot of input tokens (e.g, "Figure out the full path this function calls through and give me a map", or "Summarize these 5 PDFs and give me the main ideas I should explore") you can fire off a Sonnet agent at low-mid reasoning at it'll be cheaper than if you had Opus do that summarization itself.
Now imagine you have a set of twenty of those tasks. You launch an Opus agent, give it the task list, and tell it not to do the work itself, but orchestrate agents to perform all of the tasks and do small spot checks to verify the work.
The overall task is completed much faster at a similar or cheaper cost with an extra verification layer inserted that wouldn't have been there if you just used Opus.
Not sure how you’re seeing that. I just ran an experiment trying this and Opus was half the cost ($0.11 avg per task for opus compared with $0.21 for sonnet), mostly due to lower token usage. Opus 5.5 is really token efficient. This is without taking advantage of the further KV savings you’d get from forking.
It has to be simple tasks. Opus is the orchestrator, and it tells 10 sub-agents to go out and look for one specific pattern in code, for instance. And it's not straight-up grepable. You have to actually inspect the code and just make an assessment as to whether or not the code fits the pattern one is looking for. In instances like that, Sonnet is really good for just plowing through lots of text and doing simple cognitive work. No need to pay double the price with Opus for large-scale batch work.
But yeah, it's completely true that you sometimes have the ironic situation where you actually pay more with Sonnet because it's worse at reasoning itself down a rabbit hole. Sonnet really should be capped at medium or high reasoning.
Ironically, one of the worst things you can do is to use Sonnet to organize sub-agents. It seems to be completely bonkers with what it asks agents to do. I tried to make a colleague of mine test the feature, and he, by accident, started a large bug hunt with Sonnet. It spawned 250 sub-agents and spent 5 hours looking through everything. It actually did find a couple of useful bugs, but not the one we were looking for, which is a stupid race condition probably.
Jcampuzano2 · · focus · HN ↗
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
jtrn · · focus · HN ↗
adastra22 · · focus · HN ↗
Silagi · · focus · HN ↗
Now imagine you have a set of twenty of those tasks. You launch an Opus agent, give it the task list, and tell it not to do the work itself, but orchestrate agents to perform all of the tasks and do small spot checks to verify the work.
The overall task is completed much faster at a similar or cheaper cost with an extra verification layer inserted that wouldn't have been there if you just used Opus.
adastra22 · · focus · HN ↗
jtrn · · focus · HN ↗
But yeah, it's completely true that you sometimes have the ironic situation where you actually pay more with Sonnet because it's worse at reasoning itself down a rabbit hole. Sonnet really should be capped at medium or high reasoning.
Ironically, one of the worst things you can do is to use Sonnet to organize sub-agents. It seems to be completely bonkers with what it asks agents to do. I tried to make a colleague of mine test the feature, and he, by accident, started a large bug hunt with Sonnet. It spawned 250 sub-agents and spent 5 hours looking through everything. It actually did find a couple of useful bugs, but not the one we were looking for, which is a stupid race condition probably.