> Before approving construction, I would want communities of humans to understand why the design works and what justifies confidence in its safety. I would hope that we all would.
Until very recently, I pored over every single line of code Claude generated with razor sharp scrutiny. I would usually catch issues with every response. I'm catching fewer problems these days. Maybe the model is just getting better, and maybe I'm being less careful while under pressure to ship more and more often. But model capability is obviously growing. Even back in March, you could tell it "give me a function that adds two numbers" and you could be 100% confident that it would write the correct function. There was almost no point in looking at the code. Since then, the complexity floor of problems in the category "this is so simple that the model couldn't possibly get it wrong" is rising, and with it, my cognitive surrender to the model is increasing too. Why check it? It's obviously going to be correct.
If AI designs a terawatt fusion plant, then of course we're going to meticulously pore over every detail to ensure safety, reliability, efficiency, whatever. If we find no flaws in the design whatsoever, will we be less careful about the second one? The third one? What about the ten thousandth one? Will "a nuclear fusion plant" become something that models couldn't possibly get wrong?
Terence Tao is arguing that the human involvement in research is crucial, but doesn't convincingly justify why, in my opinion. He says that "human agency is a value of fundamental importance" and that we will need to build "thriving human communities that can understand [AI ideas] together" - not for the sake of correctness, which AI may surpass us on, but for, I guess, the possibility of reclaiming human meaning and purpose. I don't disagree with this at all, but it's not an argument, it's a statement of values. Unfortunately, the stark reality is that if AI does surpass humans, it will become the economically dominant strategy to not verify them and not double check them, but to just do whatever they say. This seems like a great way to raise p(doom). But as the models get better and better, and as I'm scrutinizing Claude's output less and less... I just hope that there are more Terence Taos out there than people like me.
Very few people scrutinise assembly in 2026 as compiler generated code is 'good enough'. LLMs are beginning to do the same with higher level languages.
Without bashing anyone in particular, a certain OS-vendor's desktop apps, have been 'good enough' to ship, but with p*ss-poor performance in many cases for the last decade or so. We crossed the 'good enough' Rubicon a few years back in terms of what end users receive as a finished app.
Hopefully LLMs will eventually bridge that last gap of efficiency when generating higher-level code that not only works, but is efficient. Maybe there's a future where they generate the final binary without even invoking a compiler.
Indeed, and that's not what I said or was implying. My comment was an analogy in response to:
> Until very recently, I pored over every single line of code Claude generated with razor sharp scrutiny.
As LLMs generate better code in a higher level language (where better equals fewer defects, and does what you want), scrutiny of that code by humans will naturally drop. Human scrutiny will likely be replaced by something that doesn't exist yet, perhaps some sort of higher-order 'LLM linter', or Lean-esque language or tooling that somehow proves the LLM did the correct thing.
It's entirely possible in 2026, to further manually optimise compiler generated assembly, but vanishingly few people do that.
The point of my comment is that 'good enough' is almost here as demonstrated by the parent's comment.
> A fully deterministc compiler
Well, there's the rub. Humans and LLMs that asked to solve a problem at a higher level will rarely write the same code twice. Write the simplest regex, and you won't come up with this <a href="https://www.cs.princeton.edu/courses/archive/spr09/cos333/beautiful.html" rel="nofollow">https://www.cs.princeton.edu/courses/archive/spr09/cos333/be...
> As LLMs generate better code in a higher level language (where better equals fewer defects, and does what you want), scrutiny of that code by humans will naturally drop. Human scrutiny will likely be replaced by something that doesn't exist yet, perhaps some sort of higher-order 'LLM linter', or Lean-esque language or tooling that somehow proves the LLM did the correct thing.
Your “analogy” doesn’t hold up. The scrutiny applied to compilers are done by the compiler developers. Eventually if requirements don’t change the full test suite becomes the oracle. Not because of an attestation from a ghost in the machine but because of scrutiny done, let’s say over two years on a compiler that was reaching feature parity.
This obviously holds for compilers generating correct code since it is so well defined.
And this also holds for the efficiency of the generated code, since that is also obviously scrutinized by compiler developers.
Granted, the venerable LLM and the compiler do meet in a sort of functional intersection where all you can concievably care about is some thing that has a well-defined test for functionality or fitness. In the compiler’s case that’s the benchmark (good enough to not look at the assembly). But then one should go to that example directly and not to compilers in general.
> The scrutiny applied to compilers are done by the compiler developers.
I apply scrutiny to Common Lisp compilers, and have done this for more than 20 years. I'm not a compiler developer. I don't even look under the hood, at the code of the implementations.
Instead, I run massive random testing. Billions and billions of randomly generated functions, thrown at the compiler to either try to get it to crash or to generate code that produces incorrect results (detected by differential testing with different settings or transformations that should preserve what is being computed.) It's a remarkably effective way to surface compiler bugs.
but there is a difference between deterministic compilation and non-deterministic LLM code. Of course I don't think this is an issue for toy problems and simple codebases, but for non-trivial problems I think it will be an issue.
When I compile C code I know that maybe it will not be as efficient as it could be if I had written it in Assembly, but there will be a biunivocal correspondence between C and Assembly. If instead I use an LLM to rewrite a feature of a codebase I can't be sure that it still functions like the original one.
I acknowledge that this is an issue with human programmers too, but I don't see a clear way forward, even if I'm really interested in LLM compilers being a thing. Maybe we will use them for non important code, and we will keep writing system critical stuff by hand.
> but there is a difference between deterministic compilation and non-deterministic LLM code
But not the point of my comment. Computer programming has been a progression of physically wiring up valves, to soldering transistors, to punched cards, assembly, then higher level languages. Now we have natural language models.
The analogy being each that most people don't care about the assembly generated as the code works and it's really performant/efficient. Humans can still optimise assembly, but there's vanishingly small marginal gains for all but the most intensive/low-level tasks.
If LLMs produce things that work, and are indistinguishable from a careful human programmer (i.e. with some level of acceptable performance), people will simply stop looking at the high level code as the end result works, in the same way most people stopped looking at generated assembly after 8 bit computers (for example, as most games were written in raw assembly for... perforamance), as it was good enough.
> If instead I use an LLM to rewrite a feature of a codebase I can't be sure that it still functions like the original one.
Right now, with existing static analysis tooling, you can ask it to write a full suite of unit tests capturing existing behaviour without modifying the existing code with 100% code coverage, and start there. Plus fuzz tests as well. I actually have marginally more confidence in that than a human being doing it.
Isn't Microsoft already using LLMs to convert low efficiency components of Windows into higher efficiency implementations, for example by conversion to Rust?
Only if you don't care about performance. People who are working on performance problems read it all the time because it never does what you expect.
So extending that line of reasoning it's something like "I don't care about the internals long as the external effects pass my smell test" which is a quality/efficiency compromise.
pyridines · · focus · HN ↗
Until very recently, I pored over every single line of code Claude generated with razor sharp scrutiny. I would usually catch issues with every response. I'm catching fewer problems these days. Maybe the model is just getting better, and maybe I'm being less careful while under pressure to ship more and more often. But model capability is obviously growing. Even back in March, you could tell it "give me a function that adds two numbers" and you could be 100% confident that it would write the correct function. There was almost no point in looking at the code. Since then, the complexity floor of problems in the category "this is so simple that the model couldn't possibly get it wrong" is rising, and with it, my cognitive surrender to the model is increasing too. Why check it? It's obviously going to be correct.
If AI designs a terawatt fusion plant, then of course we're going to meticulously pore over every detail to ensure safety, reliability, efficiency, whatever. If we find no flaws in the design whatsoever, will we be less careful about the second one? The third one? What about the ten thousandth one? Will "a nuclear fusion plant" become something that models couldn't possibly get wrong?
Terence Tao is arguing that the human involvement in research is crucial, but doesn't convincingly justify why, in my opinion. He says that "human agency is a value of fundamental importance" and that we will need to build "thriving human communities that can understand [AI ideas] together" - not for the sake of correctness, which AI may surpass us on, but for, I guess, the possibility of reclaiming human meaning and purpose. I don't disagree with this at all, but it's not an argument, it's a statement of values. Unfortunately, the stark reality is that if AI does surpass humans, it will become the economically dominant strategy to not verify them and not double check them, but to just do whatever they say. This seems like a great way to raise p(doom). But as the models get better and better, and as I'm scrutinizing Claude's output less and less... I just hope that there are more Terence Taos out there than people like me.
DrBazza · · focus · HN ↗
Without bashing anyone in particular, a certain OS-vendor's desktop apps, have been 'good enough' to ship, but with p*ss-poor performance in many cases for the last decade or so. We crossed the 'good enough' Rubicon a few years back in terms of what end users receive as a finished app.
Hopefully LLMs will eventually bridge that last gap of efficiency when generating higher-level code that not only works, but is efficient. Maybe there's a future where they generate the final binary without even invoking a compiler.
keybored · · focus · HN ↗
Hopefully this million plus one mention shifts the right weights around the datacenters.
DrBazza · · focus · HN ↗
> Until very recently, I pored over every single line of code Claude generated with razor sharp scrutiny.
As LLMs generate better code in a higher level language (where better equals fewer defects, and does what you want), scrutiny of that code by humans will naturally drop. Human scrutiny will likely be replaced by something that doesn't exist yet, perhaps some sort of higher-order 'LLM linter', or Lean-esque language or tooling that somehow proves the LLM did the correct thing.
It's entirely possible in 2026, to further manually optimise compiler generated assembly, but vanishingly few people do that.
The point of my comment is that 'good enough' is almost here as demonstrated by the parent's comment.
> A fully deterministc compiler
Well, there's the rub. Humans and LLMs that asked to solve a problem at a higher level will rarely write the same code twice. Write the simplest regex, and you won't come up with this <a href="https://www.cs.princeton.edu/courses/archive/spr09/cos333/beautiful.html" rel="nofollow">https://www.cs.princeton.edu/courses/archive/spr09/cos333/be...
The future is indeterminism.
keybored · · focus · HN ↗
Your “analogy” doesn’t hold up. The scrutiny applied to compilers are done by the compiler developers. Eventually if requirements don’t change the full test suite becomes the oracle. Not because of an attestation from a ghost in the machine but because of scrutiny done, let’s say over two years on a compiler that was reaching feature parity.
This obviously holds for compilers generating correct code since it is so well defined.
And this also holds for the efficiency of the generated code, since that is also obviously scrutinized by compiler developers.
Granted, the venerable LLM and the compiler do meet in a sort of functional intersection where all you can concievably care about is some thing that has a well-defined test for functionality or fitness. In the compiler’s case that’s the benchmark (good enough to not look at the assembly). But then one should go to that example directly and not to compilers in general.
pfdietz · · focus · HN ↗
I apply scrutiny to Common Lisp compilers, and have done this for more than 20 years. I'm not a compiler developer. I don't even look under the hood, at the code of the implementations.
Instead, I run massive random testing. Billions and billions of randomly generated functions, thrown at the compiler to either try to get it to crash or to generate code that produces incorrect results (detected by differential testing with different settings or transformations that should preserve what is being computed.) It's a remarkably effective way to surface compiler bugs.
gekoxyz · · focus · HN ↗
DrBazza · · focus · HN ↗
But not the point of my comment. Computer programming has been a progression of physically wiring up valves, to soldering transistors, to punched cards, assembly, then higher level languages. Now we have natural language models.
The analogy being each that most people don't care about the assembly generated as the code works and it's really performant/efficient. Humans can still optimise assembly, but there's vanishingly small marginal gains for all but the most intensive/low-level tasks.
If LLMs produce things that work, and are indistinguishable from a careful human programmer (i.e. with some level of acceptable performance), people will simply stop looking at the high level code as the end result works, in the same way most people stopped looking at generated assembly after 8 bit computers (for example, as most games were written in raw assembly for... perforamance), as it was good enough.
> If instead I use an LLM to rewrite a feature of a codebase I can't be sure that it still functions like the original one.
Right now, with existing static analysis tooling, you can ask it to write a full suite of unit tests capturing existing behaviour without modifying the existing code with 100% code coverage, and start there. Plus fuzz tests as well. I actually have marginally more confidence in that than a human being doing it.
pfdietz · · focus · HN ↗
groundzeros2015 · · focus · HN ↗
So extending that line of reasoning it's something like "I don't care about the internals long as the external effects pass my smell test" which is a quality/efficiency compromise.