Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
I'm rather surprised to hear this. This feels like a post from about 18 months ago. I can't remember the last time I encountered a genuine code hallucination from a frontier model. They have other issues, but rarely this.
C++ and Unreal Engine but it has full source access. If you're actually trying to deep dive on bugs, it's still confidently wrong a lot of the time.
It's a bit better than 18 months ago but it's hard to say by how much. It just seems like the culture has moved to building up fixtures that let the LLMs brute force the problems. To my eyes that's the opposite of solving things logically. It has the added effect of hiding how the sausage is made, though.
I mean, how can they possibly say they haven't written a line of code if they're actually going through it? I can only assume they're just looking at the results. So then how can they judge it's good at logic?
If it was so good at not making mistakes, why even have tests? It's nonsensical on its face.
Fintech, easy to trigger some sort of failure mode or obvious gap once specs get detailed enough and the prompt intents become specific enough. Doesn't require exhausting available context. Opus 5 did feel like a regression, will hold out judgement on 5.5 which seems much more promising.
While possibly sincere, this kind of performative incredulity has grown really tiresome. It's been two straight years of "you must not have used it in the past month." No, there are fundamental flaws with LLMs and diminishing rather than accelerating improvements.
efficax · · focus · HN ↗
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
jayd16 · · focus · HN ↗
It's wild to read this stuff and then also deal with the constant headaches of day to day hallucinations when interacting with Claude et al.
chimprich · · focus · HN ↗
What kind of domain are you working in?
weakfish · · focus · HN ↗
daveguy · · focus · HN ↗
jayd16 · · focus · HN ↗
It's a bit better than 18 months ago but it's hard to say by how much. It just seems like the culture has moved to building up fixtures that let the LLMs brute force the problems. To my eyes that's the opposite of solving things logically. It has the added effect of hiding how the sausage is made, though.
I mean, how can they possibly say they haven't written a line of code if they're actually going through it? I can only assume they're just looking at the results. So then how can they judge it's good at logic?
If it was so good at not making mistakes, why even have tests? It's nonsensical on its face.
smrtinsert · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
pugnacious · · focus · HN ↗