‹ BackHN Continuity

Thread

Coding is not solved

584 points · 544 comments · firstSpeaker

  1. efficax · · focus · HN ↗
    Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.

    What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

    If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.

    Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.

    1. pu_pe · · focus · HN ↗
      I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.

      Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.

      TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.

      1. westurner · · focus · HN ↗
        > I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop

        I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage?

        So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example).

        Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering?

        > I think a codebase generated by AI is actually more understandable than one generated by humans at this point,

        From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage, this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare.

        I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review.

        Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat.

        Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue.

        One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria;

        From &quot;Groundtruth – checks your AI coding agent&#x27;s claims against the Git diff&quot; <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48838209">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48838209 :

        &gt; &quot;Follow up to verify that the work was actually satisfactorily completed&quot;

        &gt; Are there other sound management practices that aren&#x27;t yet effectively implemented in current gen agents?

        Oh, and always write tests, docs, commit messages, and changelog entries; but don&#x27;t waste tokens on documenting something that doesn&#x27;t verifiably pass sufficient tests.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.