‹ BackHN Continuity

Thread

Coding is not solved

584 points · 544 comments · firstSpeaker

  1. askonomm · · focus · HN ↗
    What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster. As a result of the sheer amount of code now being pushed out, code reviews, a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code, is effectively dead in the water since no human can actually review such amounts of code realistically anymore. Some companies have adopted AI to review code, which, well ... you have AI make code, AI review code ... I hope you can see the stupidity here if you expect to see any deterministic results at all.

    I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.

    Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.

    1. whatever1 · · focus · HN ↗
      Even if you are competent I cannot review your 5,000 lines of code you produce per day vs the 100 you were producing before the LLM apocalypse.
      1. rfgplk · · focus · HN ↗
        5,000 is the output velocity of someone not fully immersed in agentic coding. I've seen repos do ~100k to ~250k loc changes per week.
        1. whatever1 · · focus · HN ↗
          I mean these guys are not even pretending to be reviewing the code.

          It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.

          1. bunderbunder · · focus · HN ↗
            This is exactly what I've seen.

            For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:

            1. The promised scalability improvements did not materialize. Instead, it got worse.

            2. Observability had been lost. The telemetry was no longer trustworthy.

            3. Users stopped trusting it because it was producing incorrect outputs.

            What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.

            The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.

            But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.

            1. CamperBob2 · · focus · HN ↗
              TL,DR: skill issue.
        2. 0c3ca83 · · focus · HN ↗
          Yes, they're certainly squeezing 500 lines of functionality into 250,000 lines of code. Agents are great at this.
          1. bitwize · · focus · HN ↗
            Tell me you're not using a frontier model without telling me you're not using a frontier model
            1. paganel · · focus · HN ↗
              Where's the great software, then? I'm genuinely asking: where is it? Because I can't find it, and it's been close to a year since AI for programming has started to take off.
              1. dawnerd · · focus · HN ↗
                In fact, a lot of the once great software that's switched to AI driven development has gotten worse.
                1. cat-snatcher · · focus · HN ↗
                  For example?
                  1. 0c3ca83 · · focus · HN ↗
                    Google.
                    1. cat-snatcher · · focus · HN ↗
                      It was going downhill way before AI.
                      1. intrikate · · focus · HN ↗
                        That's true, but also doesn't prevent the accelerated decay that it seems to be undergoing since AI hit the scene en masse a few years ago.
                  2. dawnerd · · focus · HN ↗
                    Clickup. Windows. Github, VSCode...

                    Could be argued they were going downhill before but it's a much faster decline since 2023-ish

                    1. necovek · · focus · HN ↗
                      Somebody posted that GitHub status summary: in ten years, they averaged 10 incidents per month — this includes the last 12 months. In the last six months, average is 22.
            2. 0c3ca83 · · focus · HN ↗
              Tell me you never once bothered to look at the generated code without telling me you don't look at the generated code.

              Luajit is under 80,000 lines of code.

            3. whateveracct · · focus · HN ↗
              i have unlimited tokens and i throw Fable / Astra at everything. They suck ass still for anything nontrivial. I could commit that garbage but if I kept doing it, I will end up with a ball of mud only Fable / Astra can grok..convenient for Dario and SamA..
            4. necovek · · focus · HN ↗
              I've asked Codex with GPT 6 Astra to *review" a one time benchmarking script for any mistakes (built by Claude Code using Opus 5.5) and it refactored the shit out of it claiming all sorts of stuff without even asking about the context in which it was developed.

              If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight.

              More and more frequently, I instruct frontier models to stop and go back to the task at hand.

              1. vandopereira · · focus · HN ↗

                [dead]

            5. perrygeo · · focus · HN ↗
              Frontier models are subjectively worse at this, in my experience. Very intelligent but very prone to expanding scope. I need to spend more time prompting to get good results. Maybe I'm just working on boring stuff that doesn't require "frontier" intelligence?
        3. Ambolia · · focus · HN ↗
          Can the users of the software even keep up at that point? We may have reached diminishing returns on software production, and not enough impact on the rest of the process.
        4. teiferer · · focus · HN ↗
          So where is all that new software? My laptop and phone run essentially the same software as 2 or 3 years ago. Yes there were some minor updates to some apps, but nothing faster than in the years prior.

          Where does all that supposed productivity go?

          1. al_borland · · focus · HN ↗
            The only real updates I've seen to anything have been AI features... so all this AI is only being used to add AI to stuff. Most of which the average person doesn't seem to want or use.
            1. Kotlopou · · focus · HN ↗
              If I understand it correctly, Firefox has some new-ish features that are marketed as AI-made, like tab grouping and tab comments... which also just get in the way for me.

              The problem is that as a lay user of Firefox, I don't know how you could even make it better in terms of features (I could see it being faster etc.).

              1. al_borland · · focus · HN ↗
                Firefox has added a lot of anti-features, like constant pop-ups that randomly trying to explain the UI to me. It’s almost as bad as the needless browser onboarding process in Edge (on servers this a huge pain).
          2. Powdering7082 · · focus · HN ↗
            Number of git pushes in GH is way up. 320M in Q1 this year vs 80M in Q1 or 2020.

            <a href="https:&#x2F;&#x2F;innovationgraph.github.com&#x2F;global-metrics&#x2F;git-pushes" rel="nofollow">https:&#x2F;&#x2F;innovationgraph.github.com&#x2F;global-metrics&#x2F;git-pushes

            Just because you haven&#x27;t installed new software doesn&#x27;t mean that new software doesn&#x27;t exist.

            1. californical · · focus · HN ↗
              But also, having significantly more code churn doesn’t necessarily mean there is more or better software
            2. nemetroid · · focus · HN ↗
              The question was &quot;where is all that new software?&quot;.
              1. necovek · · focus · HN ↗
                I see somebody asking &quot;where is that great software&quot;, but I think it&#x27;s obvious to everyone that there is more (mostly bad) software.
                1. teiferer · · focus · HN ↗
                  There are a few one-shotted personal single page apps, yes. Mostly useless beyond an initial &quot;that&#x27;s cool&quot; factor. Pre-LLM that was occupied by the &quot;I hacked this in a weekend&quot; niche.
            3. legulere · · focus · HN ↗
              To quote the article: &quot;Don&#x27;t confuse motion with progress.&quot;
              1. tripledry · · focus · HN ↗
                I remember fighting with business about this in a startup way before LLMs, more code, velocity or motion doesn&#x27;t mean more value.

                If I put on my optimist hat for a moment, I hope LLM&#x27;s would finally make this obvious for people, more stuff !== more value. 250k lines of code a week isn&#x27;t a flex IMO, it doesn&#x27;t really mean anything out of context.

            4. perrygeo · · focus · HN ↗
              There&#x27;s a difference between code checked into github vs production software.

              So to answer the original question, &quot;Where does all this supposed productivity go?&quot;: Sitting in a github repo somewhere, undeployed.

        5. pretendscholar · · focus · HN ↗
          What kind of applications are people building that involve 250k loc a week? Genuinely trying to understand this.
          1. mstaoru · · focus · HN ↗
            So far I mostly see a metric sh.. ton of meta- and meta-meta-projects re-wrapping AI wrapper tools, with sloppy slogans like &quot;One Model, Five Harnesses. Combined.&quot; or &quot;You run in the park. Rrrunnnrr.ai runs your AI.&quot; Looking inside, out of 250k it&#x27;s often 200k of verbally incontinent self-explaining comments; or &quot;smart&quot; redesign of builtins.

            No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel.

      2. bitwize · · focus · HN ↗
        That&#x27;s okay. Reviewing the code will become the agents&#x27; job as well.

        A couple more step functions in model capability of the type we&#x27;ve seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

        1. xpct · · focus · HN ↗
          &quot;A couple more step functions&quot; is doing a lot of heavy lifting here
          1. pydry · · focus · HN ↗
            It&#x27;s the AI bro&#x27;s mantra.

            It didnt go wrong

            And if it did, it was because you werent using the latest model.

            And if you were, it was because you didnt have the appropriate guardrails.

            And if you did, it&#x27;s because you didnt have AGENTS.MD.

            And if you did, it&#x27;s because you didnt prompt it properly.

            And if you did, it you&#x27;re still going to be redundant soon because I&#x27;m sure the next model released will fix whatever went wrong.

            1. mywittyname · · focus · HN ↗
              I&#x27;ve stopped calling out Claude mistakes on team meetings because this is so true.

              I mean, sure, I could have predicted in what ways an LLM would fuck up, but there&#x27;s just so many ways I can&#x27;t keep up.

              We just had a major production issue because someone&#x27;s LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn&#x27;t catch it, neither did the human reviewer for that matter.

              1. ysajang · · focus · HN ↗
                someone&#x27;s LLM wrote queries against dev databases... AI reviewer didn&#x27;t catch it, neither did the human reviewer
        2. weatherlite · · focus · HN ↗
          &gt; All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

          Kinda what i&#x27;m doing already, but for the young startup I&#x27;m at that&#x27;s surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays.

        3. necovek · · focus · HN ↗
          &gt; All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

          Sounds like the easiest thing in the world: I wonder why did we not think of it earlier?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.