Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
> What LLMs make possible is for me to say: find out all the ways this thing works.
They are not good at that. The space of possibilities can be massive and LLMs are terrible at exploring such space because they predict from the prior tokens they made. They are inherently bad at exploring new space because it is antithetical to how they work.
To me it's deeply concerning that so many people are getting fooled into thinking that LLMs are actually good at covering their bases like you're describing here. It's one of their weakest qualities.
> Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible.
It usually does quite a bad job at this too, often the tests it wrote feel like that of a lazy student that didn't really want to do the task and just sort of cheats at it or does a really shallow job. It certainly cannot run the software in every scenario possible.
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while
This is something that is said by everyone who contests anyone pointing out the risks in being overly trusting of AI or otherwise points out their flaws. I use the latest and greatest all the time and all the time I'll point out something that it totally overlooked and get hit with the classic "you're absolutely right". This occurs because I actually read the code and can see the myriad of blatant issues that still occur when using LLMs and know better than to trust them. You will find so many issue by delving into the details.
People ask AI to do things either they haven't done, cannot do, or don't want to.
So when the AI produces something that basically works and looks like a plausibly good attempt, they jump to the conclusion that AI must be excellent at that task.
efficax · · focus · HN ↗
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
OGWhales · · focus · HN ↗
They are not good at that. The space of possibilities can be massive and LLMs are terrible at exploring such space because they predict from the prior tokens they made. They are inherently bad at exploring new space because it is antithetical to how they work.
To me it's deeply concerning that so many people are getting fooled into thinking that LLMs are actually good at covering their bases like you're describing here. It's one of their weakest qualities.
> Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible.
It usually does quite a bad job at this too, often the tests it wrote feel like that of a lazy student that didn't really want to do the task and just sort of cheats at it or does a really shallow job. It certainly cannot run the software in every scenario possible.
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while
This is something that is said by everyone who contests anyone pointing out the risks in being overly trusting of AI or otherwise points out their flaws. I use the latest and greatest all the time and all the time I'll point out something that it totally overlooked and get hit with the classic "you're absolutely right". This occurs because I actually read the code and can see the myriad of blatant issues that still occur when using LLMs and know better than to trust them. You will find so many issue by delving into the details.
zero_shift · · focus · HN ↗
People ask AI to do things either they haven't done, cannot do, or don't want to.
So when the AI produces something that basically works and looks like a plausibly good attempt, they jump to the conclusion that AI must be excellent at that task.