Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
If you have well-specified tasks, you can easily reach 10 simultaneous agents working on disjoint parts of the code in worktrees. That consumes tokens pretty quickly!
I created an orchestration skill for myself (using herdr but any persistent mechanism works). So I then only interact with a front session and it will triage and dispatch each request to the relevant spaces (each of them can have multiple worktrees of the same project), summarize movements and pending decisions for me all at once. I do not directly interact with a multiplexer or any dashboard.
This seems complex and expensive, and I suppose the only reason to do this is because you want to generate code faster? Do you really have so much code to write that a single LLM is too slow?
I don't do this at my day job (coworkers would be pretty mad). This is for software ideas that comes up weekly that I need to execute to at least MVP before I ever need branching/merging.
Not complex at all, only one extra session other than the ones doing work and it's on a dumb model and can be thrown away & restarted because it only dispatches work, not doing anything.
I do everything in there, collecting requirements, kick off research, branching, merging, not one other agent on top. I considered making that orchestration command llm-powered but it's not justified at my current use.
It's not more expensive, in fact I could have just chugged along with the slow and manual session by session work but I have a claude subscription and another GLM one (the most low cost basic tier, not even much), that just sit there collecting dust if I don't put them to use in a more efficient way.
And doing session by session would face your problem when context switching too much become unscalable.
Personally, it takes me longer to write the specifications than it takes the model to implement them (and it takes me much longer to review the resulting code, although maybe that makes me old-fashioned). Consequently I do not have enough tasks to run more than one agent at a time.
I'm exactly in this situation, and at the same time I get so many jumpscares when reviewing the code that I'm not going to stop anytime soon.
Exactly. I don't understand how so many developers seem to have a long tail of well written task specifications ready to submit to the LLM.
Who produces them?
LLMs are good at cleaning up tech debt. Give them lots of small refactorings or dig out those crusty old P4 tickets.
There's a lot of easy stuff for them that requires very little specification and very high probability they'll get it right the first time, especially if you can point them at an example done right.
I have conversations with a smart model, and then the model writes the spec. I review and approve the spec, and it dispatches.
For a concrete example, check out this random plan [0]. A detailed spec followed by the exact implementation tasks that will be executed by the subagents.
Not sure why I'm being downvoted for stating the obvious. To the parallel agent skeptics: I was also a skeptic until a month or two ago. I would run one agent, watch it carefully, and check its work. However the models got good enough that it was more efficient to do more work in parallel, then have a single agent integrate the changes with a critical eye, then run a code quality pass, and then I would take a look a it and kick the tires behaviorally.
It does involve letting go and not micromanaging every code convention and implementation detail, but that is the same skill you need when leading engineering teams.
Sol- · · focus · HN ↗
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
miki123211 · · focus · HN ↗
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
bensyverson · · focus · HN ↗
thunky · · focus · HN ↗
bensyverson · · focus · HN ↗
nsonha · · focus · HN ↗
thunky · · focus · HN ↗
nsonha · · focus · HN ↗
Not complex at all, only one extra session other than the ones doing work and it's on a dumb model and can be thrown away & restarted because it only dispatches work, not doing anything.
I do everything in there, collecting requirements, kick off research, branching, merging, not one other agent on top. I considered making that orchestration command llm-powered but it's not justified at my current use.
It's not more expensive, in fact I could have just chugged along with the slow and manual session by session work but I have a claude subscription and another GLM one (the most low cost basic tier, not even much), that just sit there collecting dust if I don't put them to use in a more efficient way.
And doing session by session would face your problem when context switching too much become unscalable.
daemonologist · · focus · HN ↗
stymaar · · focus · HN ↗
gbalduzzi · · focus · HN ↗
8n4vidtmkvmk · · focus · HN ↗
lobocinza · · focus · HN ↗
bensyverson · · focus · HN ↗
For a concrete example, check out this random plan [0]. A detailed spec followed by the exact implementation tasks that will be executed by the subagents.
[0]: <a href="https://github.com/bensyverson/woodcase/blob/main/project/2026-09-07-scripting-host.md" rel="nofollow">https://github.com/bensyverson/woodcase/blob/main/project/20...
bensyverson · · focus · HN ↗
It does involve letting go and not micromanaging every code convention and implementation detail, but that is the same skill you need when leading engineering teams.