Yeah, I have never understood the over reliance on AI. Writing the code is not the challenge. The time it takes to push a new feature and test it out is often trivial, maybe a few hours.
The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product or time spent maintaining and refactoring large code.
Anyways, if you want to continue on the path towards regaining control and take it to the next level I wrote something similar here: <a href="https://blog.sharefile.systems/be-brave-go-low/" rel="nofollow">https://blog.sharefile.systems/be-brave-go-low/
Wriring the code is not the challenge, but it's what was taking up most of the time. Not the typing itself, but also because I had to think of how to implement it.
Now I can just say "add 2FA" and in 5 minutes, while I test something else, it is done.
It also made iterations a lot faster, you can try something out, see how it feels, if it doesn't work, you can just trash all the code and start again.
Haven't typed a line of code or read any code for over 6 months now.
And I used to love coding and be a competitive programmer, but this is how "coding" goes nowdays.
I have a mental model of what it would do, and how it would work, and I ask questions to confirm things and tell it to watch for specific gotchas. Then simply test the feature myself a bit.
Security-wise, I think the latest cyber models are better than me anyway at finding vulnerabilitates and pentesting such features.
Plus 2FA is a very common pattern, so it likely has in the training dataset many really good implementations.
You are abdicating your responsibility, which is fine for toy apps for yourself, but less fine when people expect that 2fa implementation to protect their accounts.
I think it's the opposite, with the right guideance, testing, frameworks, and using the top models today, implementing something with AI is most of the times better than what most developers would do.
I basically only manually test e2e myself, other tests are automated, code review is automated. I can tell the model to test for me too specific things or to add tests for specific potential issues, performance benchmarks, compared different implementations, etc.
The focus is a lot more around the code than on the code.
You don’t, that’s why in agentic world your codebase is only as good as what you can prove. For this reason, you’re going to see more languages evolving feature like capability permission, effect/coeffect types, refinement types, formal verifiers, strict type checkers, static analysis and so forth. Tests are only a small part of the verification. These ideas are old and have sat outside of the mainstream coding world, but the value proposition in the ai age has changed enough their relevance is renewed.
Not arguing for not reading test code. A lot can be achieved by instructing agents to balance out the testing pyramid with the right amount of fast end to end tests, property based tests for the right things, parametrized example based tests. Ensuring the local and CI has the right mix of tests running at right time. On projects where I have less time to review AI generated code I channel my anxiety into setting up guardrails and processes for the agents and then force them to bump up against them. In the end I view it as creating frameworks which allow me to outsource some of my attention to the agents, so that I can claw back some of that time to go set up more guardrails for more agents who are working on something else.
1. If what you say is working, you have a working software factory that should be capable of matching the output of dozens of engineers.
What very impressive externally verifiable results have you had with this?
2. If you are working in software with plenty of customers, my strong suspicion is that there are people on your team who are looking at the code who furiously trying to reign in your output.
> What very impressive externally verifiable results have you had with this?
I notice this weird hostility whenever the topic of AI coding comes up and it's never made much sense to me. If someone told me about their new method for practicing guitar I'd feel like a real tool if I demanded they prove it for me then and there.
People don't owe you their "very impressive externally verifiable results" - /u/XCSme already posted their app in another comment, it looked fine to me.
You've made your ideological position very clear here, you don't need to keep heaping it on.
If someone tells us about their new method for practicing guitar, it is very natural and not at all a toolish behaviour to hand them a guitar and say “go on, play us something”.
In fact it’s actually a very socially agreeable action, as it gives them the opportunity to show off their new skills without them looking like they’re bragging.
Now, if you happen to know for a fact that the guy cannot actually play guitar, then you’re just setting him up to embarrass himself, which is maybe an extreme punishment for the relatively minor crime of spouting some bullshit. I could buy that that is hostile, sure.
More like, if someone told you their workout program that takes 5 minutes a day lets them lift the same weights as an average Olympic weightlifter, you'd ask to see the results.
I don’t care what somebody believes about the pile of code they’ve got sitting on their own computer. The appropriate level of engagement is sort of: well, they don’t owe us any evidence we don’t owe them any credulity, and if we’re all happy to ignore each other that’s fine.
I can see why some folks here want to quibble, though. In other comments in this thread, they’ve compared to code quality favorably to the output of an average developer. That’s slightly insulting to the field in general (although, I guess most programmers have a low opinion of average code quality).
I make my own software, so no team to look over the code or be bothered by it. I don't even know how "vibe-coding" works in a team environment, because for me now it feels like it's "ideas to app" directly, so dumping by brain/ideas directly into a functional product.
I think this one is really cool[0], will be a free piano learning app. I do have other projects, but they are all at around 80% too, because some systems are shared amongst the projects and have to be finalized too (i.e. now I'm implementing my own transactional/marketing email service on top of Amazon SES, I need it before releasing ultimidi so people can register and receive email confirmations).
That is a cool app. But if you’d said you’d made it by hand in a few months I would have believed you.
If I had a working software factory like you describe, I’d expect you to have hundreds of apps of that level of complexity in a year.
> I don't even know how "vibe-coding" works in a team environment, because for me now it feels like it's "ideas to app" directly, so dumping by brain/ideas directly into a functional product.
But you don’t have any users much less paying users, so you have no idea if this system works when you do.
> I’d expect you to have hundreds of apps of that level of complexity in a year.
I am working at around 7-8 projects at the same time. The limit becomes me having to remember what I was doing for each one. AI can implement things nicely, but it really sucks at deciding which features and having its own ideas about novel game mechanics or UI/UX patterns.
I don't consider it a "software factory", just a more robust way to implement features, and it's still a WIP. One system I implemented locally is called "TaskHub", which receives some implementation or testing details from a SoTA mosel like Astra and implements it locally using Qwen 3.8 27b running on a RTX 3090.
I delegate things like app testing, navigating the app and taking screenshots, analyzing screenshots, etc.
And yes, this specific app doesn't have any users yet, I will launch it next week after I finalize the user authentication, payment flows and mobile apps. I will release it mostly as is, see user response, and then change it accordingly.
> I don't even know how "vibe-coding" works in a team environment, because for me now it feels like it's "ideas to app" directly, so dumping by brain/ideas directly into a functional product.
If you have a system that does “ideas to app directly”, that certainly sounds like what people are talking about when they say software factory. In fact my very large company would probably pay you millions if you could come in and make this work for our applications.
Hmm, my phrasing was ambigous, I didn't mean it's one idea fo one fully functional app.
But more of, one feature idea, that the AI can add to an app.
So it's not prompting "a gamefied piano learning app", but adding a single new feature like "add a minigame, where the user can control a character on the musical staff, [... 20 paraphs later ... ], make a good implementation plan, implement and test the app until everything is production-ready and bug free /goal"
That is a single "idea" that also requires a lot of fine-tuning after, but rarely in terms of code, most of the times it's only in terms of functionality and design choices.
I trust a modern model implementing a standard feature like this much more than 99% of the people I've worked with.
People bash LLMs for overengineering but for this it's what you want. Taking extreme edge-cases into account that a human would never bother with and obsessing over security.
The issue with LLM guarding isn't that it's "excessive" in outputting edge case handling, it's that the result often ends up just suppressing an error that actually indicates there is a bug or that should be handled elsewhere in a different way.
While I've definitely experienced it I don't think this is as much of a problem anymore. It's very easy to add an instruction to projects where you want every error to lead to a top-level throw rather than be handled. I also find that when it does try to mitigate it does so gracefully with a path you would actually make if you had infinite time, but your instinct tells you it's overkill.
I mostly vibe coded a queuing system to replace something we’re using at work (last week. Spent about $1500). Then I meticulously went through the code.
It was much harder to review because it was ultra defensive and included guards for tons of edge cases that weren’t possible.
Unnecessary abstractions for possible extension later. Useless indirection. Probably 3x as much code as there would have been if I’d written it by hand.
I didn’t one shot this. I kept a pretty tight leash on the AI. I had probably a dozen markdown files with of plans that I created over hours of back and forth with the AI and reviewed before each implementation round. I had automated reviews and quality gates etc…
What I found in review was that it was full of very subtle bugs that would have bitten hard in prod. Committing offsets asynchronously that would lead to dropped messages. Clock drift bugs that would lead to dropped messages or write amplification storms. Lack of back pressure in some stages of the pipeline that would cause notes to get silently OOM killed. Weird over-insistence on never crashing in most places that would mask systemic errors.
If I’d just shipped it without review, it would have mostly worked. But at the scale it’s going to be used (tens of thousands of messages per second) it would have caused production issues for months while we tracked down each of these issues.
> included guards for tons of edge cases that weren’t possible
It's not possible until it is. This is the justification lazy developers like we all are have been using leading to bugs down the road. This glorification of hand-made code is strange, like we weren't writing dirty code full of shortcuts and hacks all the time.
austin-cheney · · focus · HN ↗
The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product or time spent maintaining and refactoring large code.
Anyways, if you want to continue on the path towards regaining control and take it to the next level I wrote something similar here: <a href="https://blog.sharefile.systems/be-brave-go-low/" rel="nofollow">https://blog.sharefile.systems/be-brave-go-low/
XCSme · · focus · HN ↗
Now I can just say "add 2FA" and in 5 minutes, while I test something else, it is done.
It also made iterations a lot faster, you can try something out, see how it feels, if it doesn't work, you can just trash all the code and start again.
anygivnthursday · · focus · HN ↗
XCSme · · focus · HN ↗
Haven't typed a line of code or read any code for over 6 months now.
And I used to love coding and be a competitive programmer, but this is how "coding" goes nowdays.
I have a mental model of what it would do, and how it would work, and I ask questions to confirm things and tell it to watch for specific gotchas. Then simply test the feature myself a bit.
Security-wise, I think the latest cyber models are better than me anyway at finding vulnerabilitates and pentesting such features.
Plus 2FA is a very common pattern, so it likely has in the training dataset many really good implementations.
OccamsMirror · · focus · HN ↗
XCSme · · focus · HN ↗
I basically only manually test e2e myself, other tests are automated, code review is automated. I can tell the model to test for me too specific things or to add tests for specific potential issues, performance benchmarks, compared different implementations, etc.
The focus is a lot more around the code than on the code.
aprilthird2021 · · focus · HN ↗
But you didn't write or read any of the tests so how do you know they are accurate?
ModernMech · · focus · HN ↗
usewik · · focus · HN ↗
sarchertech · · focus · HN ↗
1. If what you say is working, you have a working software factory that should be capable of matching the output of dozens of engineers.
What very impressive externally verifiable results have you had with this?
2. If you are working in software with plenty of customers, my strong suspicion is that there are people on your team who are looking at the code who furiously trying to reign in your output.
joenot443 · · focus · HN ↗
I notice this weird hostility whenever the topic of AI coding comes up and it's never made much sense to me. If someone told me about their new method for practicing guitar I'd feel like a real tool if I demanded they prove it for me then and there.
People don't owe you their "very impressive externally verifiable results" - /u/XCSme already posted their app in another comment, it looked fine to me.
You've made your ideological position very clear here, you don't need to keep heaping it on.
fwlr · · focus · HN ↗
In fact it’s actually a very socially agreeable action, as it gives them the opportunity to show off their new skills without them looking like they’re bragging.
Now, if you happen to know for a fact that the guy cannot actually play guitar, then you’re just setting him up to embarrass himself, which is maybe an extreme punishment for the relatively minor crime of spouting some bullshit. I could buy that that is hostile, sure.
suttontom · · focus · HN ↗
bee_rider · · focus · HN ↗
I can see why some folks here want to quibble, though. In other comments in this thread, they’ve compared to code quality favorably to the output of an average developer. That’s slightly insulting to the field in general (although, I guess most programmers have a low opinion of average code quality).
XCSme · · focus · HN ↗
I think this one is really cool[0], will be a free piano learning app. I do have other projects, but they are all at around 80% too, because some systems are shared amongst the projects and have to be finalized too (i.e. now I'm implementing my own transactional/marketing email service on top of Amazon SES, I need it before releasing ultimidi so people can register and receive email confirmations).
[0]: <a href="https://game.ultimidi.com" rel="nofollow">https://game.ultimidi.com
sarchertech · · focus · HN ↗
If I had a working software factory like you describe, I’d expect you to have hundreds of apps of that level of complexity in a year.
> I don't even know how "vibe-coding" works in a team environment, because for me now it feels like it's "ideas to app" directly, so dumping by brain/ideas directly into a functional product.
But you don’t have any users much less paying users, so you have no idea if this system works when you do.
XCSme · · focus · HN ↗
> I’d expect you to have hundreds of apps of that level of complexity in a year.
I am working at around 7-8 projects at the same time. The limit becomes me having to remember what I was doing for each one. AI can implement things nicely, but it really sucks at deciding which features and having its own ideas about novel game mechanics or UI/UX patterns.
I don't consider it a "software factory", just a more robust way to implement features, and it's still a WIP. One system I implemented locally is called "TaskHub", which receives some implementation or testing details from a SoTA mosel like Astra and implements it locally using Qwen 3.8 27b running on a RTX 3090.
I delegate things like app testing, navigating the app and taking screenshots, analyzing screenshots, etc.
And yes, this specific app doesn't have any users yet, I will launch it next week after I finalize the user authentication, payment flows and mobile apps. I will release it mostly as is, see user response, and then change it accordingly.
sarchertech · · focus · HN ↗
> I don't even know how "vibe-coding" works in a team environment, because for me now it feels like it's "ideas to app" directly, so dumping by brain/ideas directly into a functional product.
If you have a system that does “ideas to app directly”, that certainly sounds like what people are talking about when they say software factory. In fact my very large company would probably pay you millions if you could come in and make this work for our applications.
XCSme · · focus · HN ↗
But more of, one feature idea, that the AI can add to an app.
So it's not prompting "a gamefied piano learning app", but adding a single new feature like "add a minigame, where the user can control a character on the musical staff, [... 20 paraphs later ... ], make a good implementation plan, implement and test the app until everything is production-ready and bug free /goal"
That is a single "idea" that also requires a lot of fine-tuning after, but rarely in terms of code, most of the times it's only in terms of functionality and design choices.
alasano · · focus · HN ↗
You're acting as if code was incredibly secure before LLMs because humans were reviewing it.
Kiro · · focus · HN ↗
People bash LLMs for overengineering but for this it's what you want. Taking extreme edge-cases into account that a human would never bother with and obsessing over security.
duskdozer · · focus · HN ↗
Kiro · · focus · HN ↗
sarchertech · · focus · HN ↗
I mostly vibe coded a queuing system to replace something we’re using at work (last week. Spent about $1500). Then I meticulously went through the code.
It was much harder to review because it was ultra defensive and included guards for tons of edge cases that weren’t possible.
Unnecessary abstractions for possible extension later. Useless indirection. Probably 3x as much code as there would have been if I’d written it by hand.
I didn’t one shot this. I kept a pretty tight leash on the AI. I had probably a dozen markdown files with of plans that I created over hours of back and forth with the AI and reviewed before each implementation round. I had automated reviews and quality gates etc…
What I found in review was that it was full of very subtle bugs that would have bitten hard in prod. Committing offsets asynchronously that would lead to dropped messages. Clock drift bugs that would lead to dropped messages or write amplification storms. Lack of back pressure in some stages of the pipeline that would cause notes to get silently OOM killed. Weird over-insistence on never crashing in most places that would mask systemic errors.
If I’d just shipped it without review, it would have mostly worked. But at the scale it’s going to be used (tens of thousands of messages per second) it would have caused production issues for months while we tracked down each of these issues.
Kiro · · focus · HN ↗
It's not possible until it is. This is the justification lazy developers like we all are have been using leading to bugs down the road. This glorification of hand-made code is strange, like we weren't writing dirty code full of shortcuts and hacks all the time.
sarchertech · · focus · HN ↗
Overly defensive code is harder to read and change for both humans and LLMs.
And many times it makes debugging harder by moving or suppressing failures.