‹ BackHN Continuity

Thread

Coding is not solved

584 points · 544 comments · firstSpeaker

  1. efficax · · focus · HN ↗
    Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.

    What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

    If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.

    Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.

    1. layer8 · · focus · HN ↗
      > I never understood the code. You think it works a certain way, until you find out that it doesn't. What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

      Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.

      “Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.

      1. emtel · · focus · HN ↗
        You’re technically correct, but the vast majority of software has never been built to the kinds of standards you are describing. LLMs are not displacing that kind of work!
        1. ffsm8 · · focus · HN ↗
          yeah, the llm approach is incredibly wasteful wrt pretty much everything. Performance, RAM, Development (Tokens).

          But it does give you surprisingly stasble rube-goldberg machines.

          And thats basically what 95-99% of enterprises want from their software.

          It annoyed me to no end when i began my career, but at this point ive accepted it and can definitely still have fun developing software with llms. As a matter of fact, as my perfectionism approach to software in my earlier years was never really appreciated... So i dont really mind the new MO.

          I still occasionally hand write though, esp. at the dayjob where ive got super small token budgets while continuously being told to use more AI. But that's normal, employers usually give off bipolar vibes with multiple stakeholders wanting to advance each of their bonus package KPI of any given quarter

          1. jonahx · · focus · HN ↗
            > yeah, the llm approach is incredibly wasteful wrt pretty much everything.

            Everything except what matters most: human time.

            1. tcfhgj · · focus · HN ↗
              well, depends on how rich you are
            2. danaris · · focus · HN ↗
              Oh, it wastes so, so much of that...
        2. boomlinde · · focus · HN ↗
          The person they responded to refers specifically to "healthcare, finance, automotive, defense, power plans, aviation, manufacturing", areas where I'd at least hope that we aspire to understand what the code does.
          1. kabes · · focus · HN ↗
            Having worked in software for healthcare, defense and finance I guarantee you we don't
            1. boomlinde · · focus · HN ↗
              Do you agree that this isn't a desirable state of affairs?
            2. robertlagrant · · focus · HN ↗
              I have also done those three and we definitely did!
              1. fragmede · · focus · HN ↗
                TLA+
                1. akkad33 · · focus · HN ↗
                  TLA+ ensures your model is correct, not that the code is correct
          2. [deleted] · · focus · HN ↗

            [deleted]

          3. Maxatar · · focus · HN ↗
            I worked briefly in heath care and the code is so brittle and so poorly understood that almost everyone is afraid to touch anything and instead it's just layers and layers of stuff trying to patch around existing code.
            1. mwwaters · · focus · HN ↗
              I know this is true. But I don’t think “Some parts of important codebases are black boxes. Therefore it’s fine if all of that code base becomes a far bigger black box” sounds like a good argument.

              Also, there was probably some human at some point that had some understanding of what they were trying to do and why. The black boxes generally get programmed around after they long left but at the time they had bugs ironed out over decades. (Yes I know sometimes true slop is done over a short period of time and the programmer leaves. But I’ve generally seen the black box built over decades instead).

            2. yetihehe · · focus · HN ↗
              I worked briefly in aviation and the part I've seen was very understandable and easy to extend and modify in understandable way. Some parts were hard, but by necessity. We also had very good tests. But maybe that one software was just a good exception.
            3. amoss · · focus · HN ↗
              In many codebases when you see this happening it is because someone reported a bug and the dev fixed the narrowest instance they could. Typically if(some specific case) do_fix() else do_normal(). Over time this fragments meaning in the source and makes it harder to reason about until it is a giant pile of slop.

              The thing the dev should do that often doesn't get done is to take the information about the undesired behavior back up to the design/spec level and rework the definition to account for it. This doesn't normally get done because it takes longer and it is conceptually more difficult than the patch.

              This is something that LLMs can accelerate significantly and the result is a shallower gradient of technical debt over the long term.

              1. cvak · · focus · HN ↗
                > back up to the design/spec level and rework the definition to account for it.

                > his doesn't normally get done because it takes longer and it is conceptually more difficult than the patch.

                This usually isn't done, because design/spec no longer exists.

                1. amoss · · focus · HN ↗
                  Sometimes this also happens because the design/spec never existed, but hopefully not in the industries under discussion.
          4. ponector · · focus · HN ↗
            I worked for a reinsurance company and their main pricing tool is a huge brittle excel file full of spaghetti VBS code. And yet, they manage to underwrite billions.
            1. tstrimple · · focus · HN ↗
              When the derecho came through Iowa and many businesses were out of power for up to multiple weeks, the very large organization I was working for at the time in the insurance space got to discover just how many of their processes relied on machines sitting under people’s desks. Business critical servers and processing just chilling on a PC under someone’s desk. Also tens of billions of dollars in revenue a year.
          5. leonidasrup · · focus · HN ↗
            Also software like SQLite:

            " SQLite is built using a DO-178B-inspired process. The testing standards for SQLite are among the highest for commercial software.

            SQLite is open-source but it is not open-contribution. All the code in SQLite is written by a small team of experts. The project does not accept "pull requests" or patches from anonymous passers-by on the internet. "

            <a href="https:&#x2F;&#x2F;sqlite.org&#x2F;hirely.html" rel="nofollow">https:&#x2F;&#x2F;sqlite.org&#x2F;hirely.html

            <a href="https:&#x2F;&#x2F;sqlite.org&#x2F;testing.html" rel="nofollow">https:&#x2F;&#x2F;sqlite.org&#x2F;testing.html

          6. Sophira · · focus · HN ↗
            Healthcare has featured many a time on thedailywtf, because a lot is (or at least was) written in MUMPS[0]. An example: <a href="https:&#x2F;&#x2F;thedailywtf.com&#x2F;articles&#x2F;A_Case_of_the_MUMPS" rel="nofollow">https:&#x2F;&#x2F;thedailywtf.com&#x2F;articles&#x2F;A_Case_of_the_MUMPS

            [0] <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;MUMPS" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;MUMPS

        3. skydhash · · focus · HN ↗
          The vast majority of software is not that important. I don’t really care about easytag (which I use for flac metadata), but I do care about xterm and tmux.
      2. coldtea · · focus · HN ↗
        &gt;“Finding out that it doesn&#x27;t” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.

        We&#x27;re not writing theorems, dude.

        Except in the equally pedantic sense that every program is a proof to a theorem...

        We&#x27;re writing plain enterprise and web software, closer to CRUD than NASA.

        If you said that even before LLMs 0.1% of teams &quot;checked all assumptions against what the code and underlying systems are actually guaranteeing&quot; in any kind of formal way, you&#x27;d be overestimating it.

        1. layer8 · · focus · HN ↗
          I’m not talking about formal verification, but about diligent informal or semi-formal reasoning through the code, so that you can rightfully claim that you understand the code and will be unlikely to be surprised by its behavior. Having learned formal verification does train that form of exhaustive reasoning about properties of the program. This practice also has you structure the code such (and select your dependencies such) that you can reason about all relevant properties. This is perfectly applicable to what you’d call CRUD and enterprise applications (that’s half the projects I earn my living with). Testing and fuzzing are complementary, but not a substitute by any stretch.
          1. jonahx · · focus · HN ↗
            GP is correct. Very few people were capable of even the informal analysis you are describing, and fewer did it. I&#x27;m not saying it&#x27;s not valuable... just stating that, empirically, it rarely happened.
            1. IAmBroom · · focus · HN ↗
              And layer8 is saying (two responses upwards by them) that this is a novel benefit of AI: it can do a particularly thorough and repetitive kind of fault analysis that is a real PITA for humans to do (by their nature, versus the nature of computers).
        2. Silamoth · · focus · HN ↗
          Who’s “we” here? Formal verification isn’t common, sure. But you don’t speak for all programmers. You might work on “plain enterprise and web software”. But there’s still plenty of other software out there that many of us work on. And lots of code being written for internal use (e.g., data analysis code) that needs to be correct.

          Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years.

          1. cat-snatcher · · focus · HN ↗
            You were really looking for reasons to get offended huh
          2. [deleted] · · focus · HN ↗

            [deleted]

      3. ModernMech · · focus · HN ↗
        Your code is only as good as what you can prove. Understanding the code is not the goal, it’s only important insofar as it helps you evolve the codebase predictably and without bugs or regressions, and understanding is not easily measurable or transferable.

        Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.

        1. layer8 · · focus · HN ↗
          Regarding the last part, the strategy is to not have everything depend on everything, to instead modularize with succinct interfaces, so that you can reason locally. Of course beyond a certain project size, there is no single person who understands every part in detail. But for every part you can have someone who understands it, and can reason about it in terms of the interface contracts with the other parts. It’s also not essential that every detail is still understood at every point in time, as long as it’s sufficiently documented. What is essential is that for every part someone did reason through it with the necessary rigor at some point.
          1. ModernMech · · focus · HN ↗
            &gt; to instead modularize with succinct interfaces, so that you can reason locally

            Okay but how does AI change any of that? You can still do that with AI.

            &gt; as long as it’s sufficiently documented.

            AI definitely helps with that.

            &gt; What is essential is that for every part someone did reason through it with the necessary rigor at some point.

            Why is that essential though? What if the person who reasoned about it dies or leaves? Moreover, why is it imperative the reasoning happens at the source code level?

            1. discreteevent · · focus · HN ↗
              &gt;&gt; to instead modularize with succinct interfaces, so that you can reason locally

              &gt; Okay but how does AI change any of that? You can still do that with AI

              With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc.

              With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.

              1. ModernMech · · focus · HN ↗
                &gt; With your own code you reasoned about it which contributed to its stability.

                Okay but to what extent? People say this but there&#x27;s no way to measure it really. Did you live through the 90s? People reasoned through all that code and it was very often quite unstable. I&#x27;m sure everyone involved with Windows ME reasoned about it quite a lot.

                What fixed that situation wasn&#x27;t that engineers today are reasoning better than engineers in the 90s, but IMO better tooling.

                &gt; With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.

                And? You haven&#x27;t established reasoning about it is actually necessary and it certainly isn&#x27;t sufficient.

            2. jimbokun · · focus · HN ↗
              You’re right, you need TWO people who have read and reasoned through the code.

              In reality, you need people who understand the design of the code well enough that they can quickly find, read, and understand the relevant parts when something goes wrong.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.