‹ BackHN Continuity

Thread

Coding is not solved

584 points · 544 comments · firstSpeaker

  1. efficax · · focus · HN ↗
    Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.

    What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

    If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.

    Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.

    1. pu_pe · · focus · HN ↗
      I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.

      Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.

      TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.

      1. suddenlybananas · · focus · HN ↗
        >TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.

        If coding were solved, then this would be true no?

        1. pu_pe · · focus · HN ↗
          There are far more moving pieces in deploying an application than coding. My impression is TFA is arguing that replacing humans with LLMs for coding would make other things like quality assurance more difficult, in which case I disagree.
      2. palmotea · · focus · HN ↗
        > I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.

        At least some places are abolishing formal QA because LLMs. There's a cult of speed uber alles that has a big intersection with LLM enthusiasm.

        1. itsalwaysgood · · focus · HN ↗
          That's an old saying: cheaper, better, faster. Pick 2.
        2. bdcravens · · focus · HN ↗
          > There's a cult of speed uber alles

          That cult was well established prior to LLMs

          1. ModernMech · · focus · HN ↗
            It's very amusing how for a decade, The Way (TM) has been "move fast and break things", but now with LLMs it's suddenly "woah, AI is moving so fast, it's going to break things!"
            1. IAmBroom · · focus · HN ↗
              Yes, but MFABT has been mocked for the entire time by many.
      3. eithed · · focus · HN ↗
        If you can quantify quality/reliability/understandability, you can tell LLM what kind of code do you expect. If not, you get whatever.

        At my current place we not only have automated tests, static analysis and static rector (linting, but also automatic pattern matcher for problematic code) but also: - architecture tests that define relationships between application layers - ADRs that guide developers (and agents as well) that communicate how new code should be written and how existing code should be treated

        I find that "how code should look like"/"what code should do" is an ambiguous idea that always is preached, but never defined = everyone's idea of quality is slightly different and only looking at existing code you tend to align. Everyone's idea of what the product does/should do is kept within their heads. If we define this knowledge in writing LLMs can not only write code according to the patterns that are thus defined, review existing code based on these documents, but also actually read acceptance criteria documents to check if the code does what it's intended to do (gherkin)

        Same goes for understandability - if LLM applies one pattern this time, another pattern another time, if you have multiple coding patterns then that hurts clarity. Sometimes LLMs work as common denominator thus achieving clarity, but I find that actually giving LLMs reference works.

      4. westurner · · focus · HN ↗
        > I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop

        I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage?

        So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example).

        Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering?

        > I think a codebase generated by AI is actually more understandable than one generated by humans at this point,

        From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage, this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare.

        I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review.

        Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat.

        Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue.

        One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria;

        From &quot;Groundtruth – checks your AI coding agent&#x27;s claims against the Git diff&quot; <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48838209">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48838209 :

        &gt; &quot;Follow up to verify that the work was actually satisfactorily completed&quot;

        &gt; Are there other sound management practices that aren&#x27;t yet effectively implemented in current gen agents?

        Oh, and always write tests, docs, commit messages, and changelog entries; but don&#x27;t waste tokens on documenting something that doesn&#x27;t verifiably pass sufficient tests.

      5. bdcravens · · focus · HN ↗
        &gt; I agree, I think this false dichotomy between using LLMs and caring about quality&#x2F;reliability needs to stop.

        It&#x27;s not just a false dichotomy, it&#x27;s intellectual dishonesty. It wasn&#x27;t that long that conversations about code quality, technical debt, etc were on the front page of HN on the regular. Whether it was coding bootcamp grads who had just enough confidence to be dangerous, &quot;just ship it!&quot; cargo culters, or the product of management breathing down the necks of otherwise good developers, there&#x27;s plenty of &quot;human slop&quot; running in production across servers worldwide.

        1. pona-a · · focus · HN ↗
          But we shouldn&#x27;t use one evil to justify another, or your argument regresses to whataboutism.
          1. bdcravens · · focus · HN ↗
            I don&#x27;t think it&#x27;s a matter of justifying any evil, so long as the argument is made in good faith. Pushing the idea of bad code emerging from AI isn&#x27;t a good faith argument.
      6. munksbeer · · focus · HN ↗
        &gt; I agree, I think this false dichotomy between using LLMs and caring about quality&#x2F;reliability needs to stop

        You&#x27;re not going to get people to stop doing that by arguing on the internet, but in the end it won&#x27;t matter, because it will stop, naturally.

        In the future, you&#x27;ll just get left behind and not hired if you&#x27;re building code by hand, it&#x27;s that simple. Even traditional code reviews are going to go away. It&#x27;ll be more about the scope and then verifying correctness.

        1. brazukadev · · focus · HN ↗
          &gt; In the future, you&#x27;ll just get left behind and not hired if you&#x27;re building code by hand, it&#x27;s that simple.

          I expect the exactly opposite to happen. These are going to be the most requested developers as the last ones that understand how it work.

          They would then be convinced to use AI for speed, but vibecoders that just prompt AI are the ones that won&#x27;t find jobs.

          1. munksbeer · · focus · HN ↗
            I&#x27;m not sure I understand what you&#x27;re saying? To keep it simple, how much code do you expect will be written by humans in say, two years time?
            1. bigstrat2003 · · focus · HN ↗
              I don&#x27;t know about two years&#x27; time, but in ten years&#x27; time the majority of code will be written by hand. LLMs are like outsourcing to cheap labor was 20 years ago: it&#x27;s trendy, but when the abysmal quality becomes apparent the pendulum will swing back.
              1. munksbeer · · focus · HN ↗
                Ok. I completely disagree.
            2. brazukadev · · focus · HN ↗
              99% of the code will be generated by LLMs but the 1% created&#x2F;curated by software engineers will still be the most used apps and systems.
          2. sally_glance · · focus · HN ↗
            Agreed for pure vibecoders, but I would expect vibecoding to just become a must have skill for other roles like product owners etc. Still expect some amount of developers to be retained for grooming the vibecoding environment, reviewing and incident response.
        2. slopinthebag · · focus · HN ↗
          but you think the prompt kiddies will have jobs? for what?

          if anything the safer bet is on skill and knowledge, or else you might as well go into a different field altogether.

        3. latentsea · · focus · HN ↗
          &gt;In the future, you&#x27;ll just get left behind and not hired if you&#x27;re building code by hand, it&#x27;s that simple. Even traditional code reviews are going to go away. It&#x27;ll be more about the scope and then verifying correctness

          I don&#x27;t really buy it actually. There isn&#x27;t really anything meaningful you can learn with how to use LLMs&#x2F;agents that has a half-life greater than a few months at this point, so you can just start doing it at any point in the future and not be meaningfully left behind. On the other hand years of letting your actual engineering skills atrophy will have a negative effect on you. I&#x27;ve been witnessing the effects of this. Going back to more coding by hand with AI-assistance circa the 2023 era as a happy medium. I think this is the sweet spot. Full agentic engineering has nasty failure modes and in the long-term is kind of a bad option for basically everyone. I say this after having done it for almost a year at this point, and transitioning away from it now.

          1. Terr_ · · focus · HN ↗
            &gt; I don&#x27;t really buy it actually. There isn&#x27;t really anything meaningful you can learn with how to use LLMs&#x2F;agents that has a half-life greater than a few months at this point, so you can just start doing it at any point in the future and not be meaningfully left behind.

            Right, remember how in ~2023 &quot;prompt engineering&quot; was ostensibly the critical thing everybody &quot;needs to learn now or get left behind&quot;?

            I feel these tools—or at least the business-model behind them—have a pattern: Basic adoption is easy, while protecting yourself against their flaws is a shifting target where expertise goes stale.

          2. munksbeer · · focus · HN ↗
            I use a combination of prompting and hand crafting my code. Sometimes I&#x27;ll sketch out the design in code, without details, other times I&#x27;ll delete and change what the LLM has done. I&#x27;m still reviewing everything it does, because:

            1. Other humans also review my code and I know they won&#x27;t approve until they&#x27;re ok with it anyway.

            2. I don&#x27;t trust that the verifications we have at the moment are good enough to catch the errors LLMs still make.

            So trust me, I&#x27;m not a &quot;vibe coder&quot;

            But I can see the writing on the wall. We&#x27;re not really maximising the potential of these things by doing it this way. In my particular industry, we&#x27;re in an arms race for clients who can easily move to a competitor. So we need to be continually improving. How I predict this unfolding is that we&#x27;ll need to get much better at verification of correctness (functional and non-functional) and once that is satisfied, humans will stop manually writing and reviewing code, otherwise you&#x27;ll just lose to your competitors. At least in my industry, maybe not all.

            1. latentsea · · focus · HN ↗
              &gt;How I predict this unfolding is that we&#x27;ll need to get much better at verification of correctness (functional and non-functional) and once that is satisfied, humans will stop manually writing and reviewing code,

              This also used to be my line of thinking too. I started going down this road. Then I stopped because I realised for this to really work you need the whole team to buy into it and then be competent at it, including product managers. I think it&#x27;s more or less above the skill floor of the average engineer, so trying to actually implement it in a team is going to be painful. For sure I think some people will go down this road, but I think the vast majority actually will not.

      7. beej71 · · focus · HN ↗
        &gt; I think this false dichotomy between using LLMs and caring about quality&#x2F;reliability needs to stop.

        Given the quantity of shit software before LLMs, they do indeed appear to be independent variables. :)

        1. supersour · · focus · HN ↗
          100%. I am really getting tired of the narrative that code before LLMs was optimally performant, perfectly architected, completely understood, and bug free...

          Does the author not have the experience of working in a legacy codebase that nobody really &quot;understood&quot;? Something sufficiently complex where even the senior SW devs needed to scope out project work and research the codebase for dependencies or potential issues?

          I fail to remember a time at LARGE_CORP where even the most experienced developers were able to scope out or design a feature without studying the existing documentation, timing diagrams, etc....

          1. beej71 · · focus · HN ↗
            What concerns me is that now AI has enabled us to make that shit software at Internet speed.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.