Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
Claude Code has the issue that sub agents inherit the thinking level. This means that to use a smarter or dumber sub agent you need a different model. That's not a particularly good reason, but that's my one use case for Sonnet.
Being that my first prompt can be something like: for task x/issue y, which model would strike the best balance between cost and capability…
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
There's different layers of understanding the system. I generally care about high level data flow, concurrency and performance (batching, holding transactions too long, back pressure etc.) rather than the mechanics of how the code actually does a thing. I still look to see what the final output looks like and ask my agent questions on how it fits in the larger system and evolve things if necessary, but agents are pretty good at writing code if the rest of the code base looks pretty decent.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
A human can produce far more code than a human can understand, too, but pre-LLM we always viewed someone overwhelming their colleagues like that as being bad at their job.
Humans could already produce more code than a human can understand. Even a single human in the pre-agentic era could produce more code than they could understand, certainly over a career and often even in the short term given the resources many companies give to maintenance.
A lot of old-school software engineering is about how to deal with this reality.
No they couldn't. You can't create software you don't understand because you wouldn't even know what to type into the IDE in the first place. I don't understand claims like these, how exactly are people especially individuals producing more code than they could understand? Even at a huge corporation one might not understand all the code but surely they understand the part they're modifying because otherwise they wouldnt know how to modify it.
Yes, it’s possible for a person to create software he doesn’t understand himself. In the old days this was pasting from Stack Overflow and changing things until it worked.
In the old days even if I knew how the software worked when I wrote it, I’d have no idea how it worked when I looked at it weeks later.
It’s also easy to modify software without knowing how it works. This produces modifications that hopefully appear to work, but that break other things, sometimes unknown things.
I’m referring to competent engineers maintaining understanding over time of all the code they’ve produced. Long before agentic coding, codebases routinely grew beyond the comprehensive understanding of their own authors.
Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.
There's understanding and there's understanding.
Have you never "fixed a bug", only to realize that you just papered over a single symptom, while the underlying bug is still intact?
People you're disagreeing with (I think!), would say that during your first attempt, you didn't _really_ understand the part you're modifying.
It is _very easy_ to do this in large codebases, and even more so when working on anything touching UI.
I suppose I am confused about why you’re confused. There’s a long history in computing of describing pieces of programming languages syntax syntax as “incantations” and similar. I suspect it has been very common, especially in the early part of developers’ careers, to know what you’re trying to do and to know that this code accomplishes it, but to not understand how the code works, to not be able to use the technique more generally, and to not understand all the effects your change has on the rest of the system.
I see, I guess I'm using "understanding" in the more literal sense of being able to put enough context together in your mind to type the characters on screen, not necessarily understand enough to know how it affects every other part of the codebase.
It's the style of understanding that says "if your animation is stuttering, set `gc.tune(pause_length=0, frequency=-1)`. Or "to make data access fast, remember to always use `integritychecks=omit;encryption=export-grade;checksum=md5`".
You don't know what these things do and what their effects really are (examples and syntax illustrative, but this is the kind of code that has disastrous effects when used carelessly), but you know they achieve your particular micro goal of "make things go fast" or "make this fit in packets on these strange industrial networks customer X has" or whatever.
I want to make better software, not more software. Making software development faster isn't necessarily the goal. Making it better in the many, many ways that matter (of which speed is just one part) is.
There is more software to be written than there were programmers so lots of people do indeed want more software, for example small tools and one off projects that aren't worthwhile to make pre LLM.
Many people want less software to deal with now, not more. Being forced to download and update apps on your phone that previously could be done without an app, for example.
A lot of people would find it easier to pull out a few quarters and put it in the parking meter than download and update an app and give it your financial information. Or hand over a paper ticket to get into an event that’s easy to transfer instead of yet another app with a barcode.
Probably because most of that software is garbage then. While there are some people who may do that, most find the convenience of not carrying around loose change as a benefit.
I'm looking forward to the making of more software. I think there are probably people who have had really useful software ideas for a long time that they'd never be able to raise money for, but now for $20 a month, they can get up and running, serving their local and/or niche communities, without having to hire a team of engineers.
Eventually we're going to reach a point where they don't have to understand the code themselves. The democratization of software creation is going to be fascinating.
The Android/iOS app store is already flooded with low quality apps. Almost 20,000 games have been released on Steam so far in 2026, that averages to about 70 games per day, every single day. Getting any kind of traction in such a environment is hopeless; every new app might as well be a scream into the void.
I think a lot of people don't want or care about major traction. They want to make little things that make their own little corners of the world better.
There are small towns all over the world which could realistically have their own little "hometown app" now, that really does track all the interesting things going on there. People don't need "The (Unofficial) Smallville Happenings" Facebook pages anymore. These don't all have some programmer who cares enough about building such a thing to make it happen, but they probably do have a teenager or even a retiree who would have a lot of fun making something that really works well for the people of that town and how they want to use it. The barrier to entry just became low enough to get over.
And that's just one example. There are SO MANY problems that used to require mega investment to solve at scale to get any traction at all. Now communities can build them for themselves. And they'll never be the kind of data target that a megacorp is, because you'd have to target each little app individually, hoping there was something useful in there.
The democratization of software development is going to have lots of things that go nowhere, and lots of little hobby projects, and a few things people will actually hear about and care about. But a lot of people's lives will become incrementally better and I'm excited to see that.
True, but if you can think about tests as "how do we validate that the outcome I want has been solved", then you can orient your tests around that. I do wonder if BDD style integration tests will end up being the path we go down.
Frontier models today don't really write incorrect code at the micro level. They do miss edge cases at the high level though, and that's what we want to test, is the scenarios.
The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
I think they are trying now to to bake CI awareness into Claude Desktop, didn't use it yet.
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
Yeah the codex app can deterministically poll and watch for you too. Consider it like an event based trigger, where the event can be anything you can dream of (like webhooks!)
FWIW, Claude Channels[1][2] are probably going to be the solution for that, eventually. While I'm not sure how the WebHook receiver example will work with, say, GitHub and a local Claude, the Chat side of things _would_. So you'd have GH send its web hook to Telegram (for example), and then the Telegram Channel MCP would inject that into Claude, and Claude would start working on the problem. Still experimental, but functional enough to play with.
This sounds... horrible? I mean, it's certainly a solution to the "wake up when this thing happens" problem, but... $SERVICE -> webhook -> $CHAT_APP -> MCP -> remote wakeup sounds both brittle and - as you said - the local code harness route is entirely unserved by something like this.
Am I having a yells-at-cloud moment where a bunch of folks are using cloud hosted LLM harnesses/environments (let's ignore the models, "of course" those are remote) and I just never saw the point?
> The thing with watching CI in an agent loop is that it burns tons of tokens.
Not my experience with Claude Code.
> writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code
This is what Claude Code does, more or less, on the fly. With a short prompt like "I pushed, monitor CI and debug if needed", it writes a monitor script which is responsible for polling CI status (the script is short, so it's not token-heavy), and if CI fails, only then does the agent proceed to pulling out CI logs, grepping them for signs of errors, etc. as continuation to debugging.
I mean, I'm sure it's more token-efficient to have a CLI tool ready-to-go instead of Claude Code dynamically writing its own script each time, but as I'm on a Max sub where it doesn't seem to affect how close I am to the limits, and I only ever hit the limits if I'm running Fable for everything... /shrug
I guess folks' experiences with this stuff will vary wildly by what environment they work in. I use LLMs mostly at work, where I don't have any subscription plans, everything is billed per-token, and there's multiple coding harnesses with different token quotas available (and vastly different functionality). So the sharable CLI that works whether I'm in Claude Code (where tokens cost some outrageous amount) or Devin CLI (a horrible harness that also lacks any sort of scheduling system as far as I've ever figured out, but hey, there's GPT Luna and GLM available, at least) is a huge win.
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
Yeah, I made this point above but LLMs just don't have a good sense of importance. They treat everything at the same level of importance and can spend considerable effort on things that just don't really matter.
I think that's what a future dev team is going to look like.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
You're thinking that the entire economy will collapse down to just 3-4 vendors?
Human power and social structures just don't work that way. No AI company is making my sandwich, operating the bus, or serving soup in the school cafeteria. Real estate, human service, specialized expertise, and have-power influence isn't going away.
Who's making the robots? Who's owning the bus? Who's growing the ingredients in the soup? Who's deciding when the bus needs to stop due to a security or safety concern? Who's flirting with the customers and recognizing the power brokers? Not the AI companies.
Just because robots can do stuff doesn't mean the human power structures or service preferences evaporate.
More tests that aren’t written by you don’t help you understand the system, and I would argue the there’s no confidence without understanding. That was true in the pre-agentic era and is perhaps even more true now.
I want to understand more about how the world around me works. Not less.
Humanity advances in proportion to how well we understand the world. If the machines understand better than us, the world will bend to fit their preferences, and ours only incidentally to the extent they coincide with the machines.
> Another thing to think about is, what would it take for you to care less about the understanding
Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai
The cat's out of the bag already. We can't undo the idea of LLMs or coding agents. If training progress stopped today, we have years and years of harness improvements to extract more performance out of today's models.
We also have open weight models too, and ways to host those at home.
Most people don't look at the assembler output of their C++ code (I used to write win32 programs in asm!). Most people don't look at the opcode instructions or JIT output of their ruby / python code. We're starting to work at a higher level of abstraction using LLMs. It's ok to be sad about it, but just being angry about it isn't going to change that there's a new world out there with a new skill set that's needed for honing.
And where's the results of all this higher level work?
Where's the super awesome 100x turbocharged software that's a result of everyone here having been being a 100x turbocharged programmer for the last 6 months and a 10x supercharged programmer the past 2 years?
I still use the same software I used 2 years ago, but a bit less reliable.
Using an LLM does not require people to stop understanding their codebases, and I find one of the best uses of an LLM is to improve my understanding of my code. I suspect the really good software is going to be written by those doing this, but time will tell.
I'm not angry, but yes I'm being deeply sarcastic to illustrate the extreme case you seem to be advocating, where we relinquish our understanding to the machines.
In a competitive business like software you need an unfair advantage and for very few people that's having near unlimited tokens. Even then I doubt that's going to produce good software.
> what would it take for you to care less about the understanding
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
Yeah, what things can you write to statically eliminate the things that are not acceptable. What would you have to change in your prompting process to get that outcome? What bad patterns is it copying from the code base that maybe you should spend tokens fixing?
I do agree that they're not great at program design by default and that's where we as engineers should spend our time. Data structures and data flow are king. But once you suss that out, they're pretty good at writing the resulting code.
This is also where I disagree with dhh about just using lower level languages. Good abstractions make for excellent program understanding and we should continue to build extremely good building blocks that make program design naturally solid.
I find that we dont need to go all in. I can use LLMs to make tooling custom to my project: linters, skills, rules, some doc etc. Then iteratively improve on that.
For example, write a skill that finds some kind of code smell, say duplication, and generate a report. Give it some supporting scripts.
Then, use this report to file a few tickets. Then make the agent fix those tickets. Then, as you grow confident, automate more of this process.
It does not replace human supervision but it may enhance it. Especially in a team where people start generating PRs faster that anyone can review them.
Continue this improvement process long enough and you may find yourself with an AI Software Factory.
Sol- · · focus · HN ↗
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
maherbeg · · focus · HN ↗
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
datadrivenangel · · focus · HN ↗
xgb84j · · focus · HN ↗
copperx · · focus · HN ↗
mnicky · · focus · HN ↗
nananana9 · · focus · HN ↗
jpease · · focus · HN ↗
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
Hauthorn · · focus · HN ↗
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
[deleted] · · focus · HN ↗
[deleted]
maherbeg · · focus · HN ↗
miki123211 · · focus · HN ↗
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
andrewaylett · · focus · HN ↗
tshaddox · · focus · HN ↗
A lot of old-school software engineering is about how to deal with this reality.
satvikpendem · · focus · HN ↗
massysett · · focus · HN ↗
In the old days even if I knew how the software worked when I wrote it, I’d have no idea how it worked when I looked at it weeks later.
It’s also easy to modify software without knowing how it works. This produces modifications that hopefully appear to work, but that break other things, sometimes unknown things.
tshaddox · · focus · HN ↗
Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.
satvikpendem · · focus · HN ↗
> literally hand-write code they don’t understand even as they write it
I find this literally impossible. How can you even start typing anything without knowing what to type?
klausa · · focus · HN ↗
Have you never "fixed a bug", only to realize that you just papered over a single symptom, while the underlying bug is still intact?
People you're disagreeing with (I think!), would say that during your first attempt, you didn't _really_ understand the part you're modifying.
It is _very easy_ to do this in large codebases, and even more so when working on anything touching UI.
tshaddox · · focus · HN ↗
satvikpendem · · focus · HN ↗
miki123211 · · focus · HN ↗
You don't know what these things do and what their effects really are (examples and syntax illustrative, but this is the kind of code that has disastrous effects when used carelessly), but you know they achieve your particular micro goal of "make things go fast" or "make this fit in packets on these strange industrial networks customer X has" or whatever.
gr_norm · · focus · HN ↗
satvikpendem · · focus · HN ↗
jimbokun · · focus · HN ↗
satvikpendem · · focus · HN ↗
jimbokun · · focus · HN ↗
Etc.
satvikpendem · · focus · HN ↗
smeej · · focus · HN ↗
Eventually we're going to reach a point where they don't have to understand the code themselves. The democratization of software creation is going to be fascinating.
ThrowawayR2 · · focus · HN ↗
smeej · · focus · HN ↗
There are small towns all over the world which could realistically have their own little "hometown app" now, that really does track all the interesting things going on there. People don't need "The (Unofficial) Smallville Happenings" Facebook pages anymore. These don't all have some programmer who cares enough about building such a thing to make it happen, but they probably do have a teenager or even a retiree who would have a lot of fun making something that really works well for the people of that town and how they want to use it. The barrier to entry just became low enough to get over.
And that's just one example. There are SO MANY problems that used to require mega investment to solve at scale to get any traction at all. Now communities can build them for themselves. And they'll never be the kind of data target that a megacorp is, because you'd have to target each little app individually, hoping there was something useful in there.
The democratization of software development is going to have lots of things that go nowhere, and lots of little hobby projects, and a few things people will actually hear about and care about. But a lot of people's lives will become incrementally better and I'm excited to see that.
jimbokun · · focus · HN ↗
So accelerate the vibe coding of shit nobody wants or asked for, just to see some metric go up somewhere.
Are we still getting bonuses for the number of tokens we can burn?
maherbeg · · focus · HN ↗
Frontier models today don't really write incorrect code at the micro level. They do miss edge cases at the high level though, and that's what we want to test, is the scenarios.
klardotsh · · focus · HN ↗
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
unddoch · · focus · HN ↗
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
maherbeg · · focus · HN ↗
ipsi · · focus · HN ↗
[1]: <a href="https://code.claude.com/docs/en/channels" rel="nofollow">https://code.claude.com/docs/en/channels [2]: <a href="https://code.claude.com/docs/en/channels-reference" rel="nofollow">https://code.claude.com/docs/en/channels-reference
klardotsh · · focus · HN ↗
Am I having a yells-at-cloud moment where a bunch of folks are using cloud hosted LLM harnesses/environments (let's ignore the models, "of course" those are remote) and I just never saw the point?
wren6991 · · focus · HN ↗
solatic · · focus · HN ↗
Not my experience with Claude Code.
> writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code
This is what Claude Code does, more or less, on the fly. With a short prompt like "I pushed, monitor CI and debug if needed", it writes a monitor script which is responsible for polling CI status (the script is short, so it's not token-heavy), and if CI fails, only then does the agent proceed to pulling out CI logs, grepping them for signs of errors, etc. as continuation to debugging.
I mean, I'm sure it's more token-efficient to have a CLI tool ready-to-go instead of Claude Code dynamically writing its own script each time, but as I'm on a Max sub where it doesn't seem to affect how close I am to the limits, and I only ever hit the limits if I'm running Fable for everything... /shrug
klardotsh · · focus · HN ↗
crooked-v · · focus · HN ↗
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
maherbeg · · focus · HN ↗
mattm · · focus · HN ↗
miki123211 · · focus · HN ↗
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
emkoemko · · focus · HN ↗
bdamm · · focus · HN ↗
Human power and social structures just don't work that way. No AI company is making my sandwich, operating the bus, or serving soup in the school cafeteria. Real estate, human service, specialized expertise, and have-power influence isn't going away.
xoac · · focus · HN ↗
bdamm · · focus · HN ↗
Just because robots can do stuff doesn't mean the human power structures or service preferences evaporate.
tshaddox · · focus · HN ↗
jimbokun · · focus · HN ↗
I want to understand more about how the world around me works. Not less.
Humanity advances in proportion to how well we understand the world. If the machines understand better than us, the world will bend to fit their preferences, and ours only incidentally to the extent they coincide with the machines.
[deleted] · · focus · HN ↗
[deleted]
willtemperley · · focus · HN ↗
Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai
maherbeg · · focus · HN ↗
We also have open weight models too, and ways to host those at home.
Most people don't look at the assembler output of their C++ code (I used to write win32 programs in asm!). Most people don't look at the opcode instructions or JIT output of their ruby / python code. We're starting to work at a higher level of abstraction using LLMs. It's ok to be sad about it, but just being angry about it isn't going to change that there's a new world out there with a new skill set that's needed for honing.
nananana9 · · focus · HN ↗
Where's the super awesome 100x turbocharged software that's a result of everyone here having been being a 100x turbocharged programmer for the last 6 months and a 10x supercharged programmer the past 2 years?
I still use the same software I used 2 years ago, but a bit less reliable.
willtemperley · · focus · HN ↗
I'm not angry, but yes I'm being deeply sarcastic to illustrate the extreme case you seem to be advocating, where we relinquish our understanding to the machines.
In a competitive business like software you need an unfair advantage and for very few people that's having near unlimited tokens. Even then I doubt that's going to produce good software.
[deleted] · · focus · HN ↗
[deleted]
mattm · · focus · HN ↗
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
ncruces · · focus · HN ↗
maherbeg · · focus · HN ↗
I do agree that they're not great at program design by default and that's where we as engineers should spend our time. Data structures and data flow are king. But once you suss that out, they're pretty good at writing the resulting code.
This is also where I disagree with dhh about just using lower level languages. Good abstractions make for excellent program understanding and we should continue to build extremely good building blocks that make program design naturally solid.
sdeframond · · focus · HN ↗
For example, write a skill that finds some kind of code smell, say duplication, and generate a report. Give it some supporting scripts.
Then, use this report to file a few tickets. Then make the agent fix those tickets. Then, as you grow confident, automate more of this process.
It does not replace human supervision but it may enhance it. Especially in a team where people start generating PRs faster that anyone can review them.
Continue this improvement process long enough and you may find yourself with an AI Software Factory.