‹ BackHN Continuity

Thread

If AI coding is lowering your code quality, you're not managing quality right

129 points · 169 comments · bucket2015

  1. axegon_ · · focus · HN ↗
    Ah, the "skill issue" argument again. Same crap aswhen everyonewas worshiping Musk 5-6 years ago, this time it's dario and altman with a claude/chatgpt mask. Crash can't come soon enough.
    1. ModernMech · · focus · HN ↗
      I think the point is just it doesn’t have to get worse, so there are things you can do to prevent / change it if it is deteriorating.
      1. Sharlin · · focus · HN ↗
        Yes, but it doesn't matter if nobody actually does that. Either because

        1. they don't care

        2. the rest of the team doesn't care

        3. the powers that be actively discourage it because velocity.

        1. bucket2015 · · focus · HN ↗
          That's a fair point. I guess step 0 is that you have to care about code/product quality and prioritize it.
        2. rgoulter · · focus · HN ↗
          Without LLMs, you can still have bad development processes which lead to increasing technical debt with no plan for paying it off.

          LLMs let you move faster.

          But it's not as if introducing them is the only reason your codebase isn't high quality.

          1. hajile · · focus · HN ↗
            When you’re required to approve thousands of lines a day (code you can’t possible understand), it certainly IS causing issues that didn’t exist before.

            Every study I’ve seen correlates the use of AI with large increases in the number of bugs. Look at Amazon dialing back AI after massive outages. Microsoft patch Tuesday releases are bricking computers (they even managed to break notepad somehow). The rash of Facebook bugs also coincided with their move to AI. Leaks from Google have engineers saying AI either doesn’t save any time because it takes so to remote stuff or it causes breakages if they speed up.

            These companies can afford to get the best devs. They have access to essentially unlimited token budgets. They have STILL fallen off a cliff in quality.

            What more proof could there be that this isn’t sustainable?

            1. bitwize · · focus · HN ↗
              A correlation-causation link between AI use and these bugs has not been established. Until it has proven to come from the region of AI-pagne, it's just sparkling enshittification.
              1. hajile · · focus · HN ↗
                This isn’t the human body or some other thing with billions of unknown parameters. It’s a single (relatively simple) math equation that you are ascribing tons of non-existent complexity to.

                Left to its own devices, that system will result in AI autophagy (model collapse) and iterative degradation.

                The only area with serious uncertainty is how humans interact, but we now have research showing humans suffer cognitive issues very quickly using AI (some studies indicate effects happen in as little as 10 minutes) with cognitive surrender being a particularly big issue.

                In a lot of systems, the only new data seems to be a few brainstorming sentences (you can read slop as entropy decaying things). The AI slops that into requirements. That slop feeds into an agent which generates a bunch of “reasoning” slop, maybe compacts everything (more slop), and spins up agents that get handed slop. They then open files with who knows how many generations of slop (maybe never even touched by a human) and write out a bunch more slop (it’s ironic that humans get better the more they edit a file, but AI gets worse). That slop gets “tested” by another agent reading all the other slop and maybe all this recurses a few generations.

                At the end of this AI equivalent to “the human centipede”, you get a developer who’s handed 10x or maybe even 100x more code than their brain would possible process. They are suffering complete cognitive surrender (not to mention often reaching mental and maybe physical collapse from the workload and stress). They don’t understand the system and rubber stamp it so they can move on to the next 50 PRs of the day.

                From start to finish, it’s 100% entropy outside a handful of lines worth of human input.

                Many people predicted bugs and even discussed entropy issues before AI coding was popular. The buggy mess timing aligns not only with AI adoption, but happens to each company ramping up as they ramp up AI usage.

                This is like seeing Einstein’s predictions happen, but arguing he can’t prove correlation/causation. What evidence would you actually accept that is feasible to study?

          2. Sharlin · · focus · HN ↗
            So you're saying that LLMs let you accrue technical debt faster? I suppose that's like the fact that living on payday loans let you accrue monetary debt faster.
            1. rgoulter · · focus · HN ↗
              Yes.

              I think if you're on a team that cares about quality, LLMs can help you write quality code faster.

              If you're on a team that's mindful about technical debt, you can have make practical trade-offs for velocity now at the expense of paying off technical debt later.

              And if you're on a team that's unable to care about code quality ("I gotta merge this code now!"), then you can write mountains more code than you can understand.

        3. axegon_ · · focus · HN ↗
          Exactly. Most restaurants you go to don't care about the food they serve you nor do they care about the products they use, as long as it doesn't harm their business. Same with groceries - manufacturers don't care, as long as what they sell you is acceptable and passes regulations. But once you go to a restaurant with standards, you can immediately tell the difference. I do not come from a wealthy family and even as such, I certainly prefer paying the higher price now that I can afford it.
    2. skybrian · · focus · HN ↗
      And why wouldn’t writing software be a skill issue? Yes, it’s an annoying meme, but we should expect that there are better and worse ways to write software. It would be weird if everyone got the same results regardless of experience.

      I’m doubtful that the author’s recommendation always work, but I do some similar things and they do seem to help.

      1. malfist · · focus · HN ↗
        Anyone who thinks they can produce high quality code from an LLM is mistaken about how to judge code. Trust me, I've seen enough PRs to last a life time. A lot of professionals wouldn't know good code if it slapped them in the face.
        1. asutekku · · focus · HN ↗
          I'd argue for most people LLM produces much better code than they would be able to write themselves.
          1. malfist · · focus · HN ↗
            That is not an argument that LLMs produce good code
            1. bluGill · · focus · HN ↗
              They produce good code when I'm personally reviewing them. There are a few other people who work with who likewise know how to review code and thus can get good code out of an LLM. There are, however, a lot of people who just accept the first slop that they get out of it and that's not good code.

              The larger issue of good code isn't the actual individual lines, it's the overall architecture. And that's what I'm going to be reviewing first is, is this a good approach? Then the interfaces to other code is this a good interface. Get those two right and we can go back for the details. In a lot of cases, the LLM is plenty good at those details.

              In some cases, an LLM is better than what I could do. Well, I suppose I can trace down all the locks in all the different special cases, and I have done that, but that was a huge amount of effort that I really don't want to repeat.

              Note that I'm talking about recent models. If you're asking about the models of just one year ago, I would give a very different answer about the type of code an LLM produces.

          2. abalashov · · focus · HN ↗
            I'm not sure how literally you mean "most people". This might be true in a purely volumetric sense, but that's not really the bar around these parts...
        2. preg_match · · focus · HN ↗
          You can most definitely produce high-quality code via an LLM, particularly if you test your code aggressively and review it. Yes there are a lot of shoddy PRs, but that's nothing too new. The problem with LLMs is the amount of code they produce. More code = more garbage. But, that code velocity can be leveraged to increase quality. Through careful design and testing.

          Ultimately, I would take LLM code + high-quality multi-strategy testing over human code with little to no tests. And some would say "well that's a false dichotomy". I disagree, before LLMs engineers didn't have the time or incentives to aggressively test. The tests either would not exist, or would be shitty unit tests intended to get an arbitrary coverage percentage. Now, we can write high-quality tests, differential testing, fuzzing, and more, in much less time.

          1. rented_mule · · focus · HN ↗
            Going much deeper on tests has been transformative for me. In a solo project started from scratch, I'm 6-7 weeks in, and it's up to ~90K lines. ~60K of those are tests. Those tests have now found (and then the agent has correctly diagnosed) multiple bugs in broadly used libraries that I'm using in my project. That's because those bugs surfaced as occasional issues in my project. Especially powerful are all the property-based tests (perhaps what you are calling fuzzing? I'm using the Python package called Hypothesis for this).

            Another spectrum that I've found useful to explore is the scope of what I ask the coding agent to do in one turn. I see some people trying to do one massive prompt that the coding agent works on for a day or more. I find a large boost in overall quality if I do 10-20 prompts per day (not counting the prompts where I'm just trying to understand things). It's still much less of my time than hand-coding, but the resulting architecture looks like my own. The quality of the overall system is great. There are certainly issues here and there in the code, but it's always that way once a project gets large enough. Now it's easier to address any particular issue throughout the code base in one go.

          2. maksDanylenko · · focus · HN ↗

            [dead]

        3. deterministic · · focus · HN ↗
          > Anyone who thinks they can produce high quality code from an LLM is mistaken about how to judge code

          Sorry, but you are 100% wrong.

          I have 30+ years of professional development experience working on complex, very large-scale C++ code used by companies around the world.

          I care deeply about code quality and always have. More than any other developer I've worked with in my 30+ year career. And I'm now using Claude Code to push the quality bar much higher.

          But you have to learn how to use it properly. It's a tool. Quality doesn't happen automatically.

          It's a big mistake, and frankly quite arrogant, to assume that because it doesn't work for you, it can't work for anyone else. Or that the rest of us must either be lying or incompetent.

          1. player1234 · · focus · HN ↗

            [dead]

          2. blub · · focus · HN ↗
            It’s fair to ask people bragging about their amazing AI skills to show their code or GTFO. Hope it becomes established.
            1. deterministic · · focus · HN ↗
              I work on proprietary software, so no, that’s not really a fair question to ask.

              However, if you’re willing to share how you’re prompting the AI and some examples of where you think the results are poor, we might be able to help identify what’s causing the difference. I’d be genuinely interested in understanding why we’re getting such different results.

              Rather than assuming one of us must be wrong, it would be more useful to compare approaches and see what we can learn from each other.

      2. tyleo · · focus · HN ↗
        Not only that, but you really want it to be a skill. The book _Making Software_ describes skills as things you can get better at through practice, and talents as things you're born with.

        I'd like to think the time and practice I've put into software engineering has made me better at it. If that's not true, then there's no reason to prefer senior or principal engineers with years of experience over newcomers.

      3. pydry · · focus · HN ↗
        >And why wouldn’t writing software be a skill issue

        You've missed the point. Nobody doubts writing code well or badly is indeed a skill issue.

        The question is that "once you account for all of the things you need to do to make the code very high quality, did vibe coding actually provide any real value?"

        I'm certain there are guardrails that help bolster vibe coding but I'm equally certain that when ive prompted something important I usually have to redo it enough times that just writing it manually myself usually would have been quicker.

        Then I watch other people who code who dump on that opinion and I see total slop. They just can't tell the difference.

        1. AndrewKemendo · · focus · HN ↗
          I hate typing and like reading

          That seems to be the primary difference I’ve found between people who embrace gen code and those who dont

          The ones who dont, seem to like the physical act of typing, and that tends to cluster with people who write software all day

          1. pydry · · focus · HN ↗
            high quality code means boiling down code to its bare essentials, not spewing boilerplate. that means deleting code where you can and crafting good abstractions while leaving functionality intact.

            if you find the ratio between typing and thinking to be very high then you're probably producing a lot of slop.

            This is a common theme I find when I hear about people's AI coding success stories. Where they say "its good at X" where X might be "backfilling unit tests" or "writing boilerplate" I usually think "if you find you need to do X a lot youre definitely doing programming wrong.

            Ive actually yet to hear an X applied to production code that doesnt make me think that.

            1. Daishiman · · focus · HN ↗
              > Ive actually yet to hear an X applied to production code that doesnt make me think that.

              This is one of those things we value in theory in engineering but not in practice. Reducing code as an artifact might mean coming up with clever ways or compressing data, like making code that generalizes and abstracts. This is fine if you're experienced and clever. But a lot of organizations don't have that many clever or experienced engineers and those tools cause more harm in the hands of those people. Hence compromises must be reached and verbosity is valued because it is explicit.

              I used to believe otherwise but then I worked in larger orgs with a lot of mediocre people who still provided value but needed to be given the means to add value.

          2. Sharlin · · focus · HN ↗
            My experience is that most people don't enjoy code review, which is why it must usually be actively encouraged and not just something that happens naturally.

            Mechanically writing boilerplate is not enjoyable, and unfortunately in some languages and domains most of the coding is writing boilerplate. Machines can help with that no problem.

            What is presumably enjoyable to most programmers is writing the parts where the actual magic happens. The translation of informal ideas into formal representation has beauty, like mathematics has beauty. Designing and implementing structures of code and data that are as simple as possible, but not simpler, is rewarded with a feeling of artisanal satisfaction and pride. Few things in life are as satisfactory as figuring out an elegant solution to a challenging problem.

            None of the above are necessarily bound to the actual typing of words and symbols. AIs can help with all of them, and act as a genuine force multiplier. I would describe that as "responsible use of AI". Unfortunately, it seems that incentives are often against such use.

    3. Sharlin · · focus · HN ↗
      In a way it reminds me of the good old "if agile doesn't work for you, you're not doing agile right".
      1. osigurdson · · focus · HN ↗
        Agree. These days, if you think you have a methodology that works better than others, you can actually try it / compare it and publish it so that others can replicate and critique your work. Articles like this one, that merely claim they've found the secret sauce, therefore should not carry much weight.

        That wasn't the case with 00s agile / Uncle Bob stuff since proving that any of it was helpful was impossible - you just had to believe (and if you didn't believe there was something wrong with you!).

    4. CrimsonRain · · focus · HN ↗
      It is indeed skill issue.

      You don't think crash will happen because XYZ. You _wish_ for the crash because you are hateful of progress that you are not part of.

      1. mitxela · · focus · HN ↗
        Studies (e.g. METR) show AI programmers think they're better but they're worse.
        1. Madmallard · · focus · HN ↗
          assuredly true
        2. vasko · · focus · HN ↗
          Anthropic's own study showed the same, yet people ignore it.
        3. bitwize · · focus · HN ↗
          METR has admitted that that study was flawed, and when they tried to rerun it they hit a snag: nobody wants to not use AI to code anymore!
      2. axegon_ · · focus · HN ↗
        > You _wish_ for the crash because you are hateful of progress that you are not part of.

        Microsoft 2000

        > You _wish_ for the crash because you are hateful of progress that you are not part of.

        Facebook 2008

        > You _wish_ for the crash because you are hateful of progress that you are not part of.

        Cryptobros 2013

        > You _wish_ for the crash because you are hateful of progress that you are not part of.

        Altman/Dario/Musk 2020-onwards.

        There might be a trend here...

        1. SaucyWrong · · focus · HN ↗
          Microsoft is a monopolist that makes some of the most hated products on earth.

          Facebook has been shown to derange young minds and has been a nonstop firehose of disinformation into global public discourse.

          Crypto moved an insane amount of wealth from the poor to the rich through and uncountable number of scams and empty promises.

          These? These are what you call progress? Yes, the created wealth for a few individuals, but progress? No, my friend.

          1. deterministic · · focus · HN ↗
            > Microsoft is a monopolist that makes some of the most hated products on earth

            Really? That is quite a claim. You are talking about one of the worlds most successful software companies. What fact based research or large scale survey do you base that on?

            1. SaucyWrong · · focus · HN ↗
              I’ll cop to not having facts to back up the perception of its software, but its status as an illegal monopoly is cemented by its successful prosecution by the DOJ.

              But Facebook is extremely successful. Some crypto companies are very successful. I don’t hold any of these up as exemplars of the progress of humanity.

              EDIT: My response to the GP, whose claim was that the only reason for the haters is that in each case they wanted the business to fail because they weren’t part of it, should been, no, actually there were at the time other and valid reasons detractors of those companies thought the way they did, and the same is true this time.

            2. player1234 · · focus · HN ↗

              [dead]

          2. axegon_ · · focus · HN ↗
            I was being sarcastic. I hate all of those with a passion.
            1. SaucyWrong · · focus · HN ↗
              Bravo, the sarcasm went right over my head :)
        2. vanuatu · · focus · HN ↗
          Famously, Microsoft stop progressing after 2000. Facebook disappeared after 2008. BTC stopped appreciating since 2013. All the investors in these trends regretted it.
          1. SaucyWrong · · focus · HN ↗
            Yeah again, you’re confusing wealth accumulation for progress.
      3. mcmcmc · · focus · HN ↗
        Why presume hate?
    5. hypfer · · focus · HN ↗
      It actually is though?

      Though arguably more of a process and judgement issue than skill.

      What makes LLM-generated code a bit special there is that misjudging how to deal with it seems to be what most people do. So the default is broken.

      Whereas in prior iterations of "skill issue", the default was working.

    6. post-it · · focus · HN ↗
      What's a crash going to do? The internet didn't disappear after the dot com bubble popped.
    7. rgoulter · · focus · HN ↗
      LLMs are not magical tools which take slop as input, and produce well thought out documentation and tests and code as a result.

      Over the last year, LLM coding agents gotten pretty good. It's no longer "if your results suck, you gotta try the latest and greatest model". You can get capable results on a wide variety of tasks, with a wide variety of models, used in a wide variety of ways.

    8. mitxela · · focus · HN ↗
      What were people worshipping Musk for 5-6 years ago?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.