‹ BackHN Continuity

Thread

Pop!_OS bans AI-generated code from much of its codebase

116 points · 167 comments · bundie

  1. brink · · focus · HN ↗
    I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
    1. lucianmarin · · focus · HN ↗
      Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
      1. dnautics · · focus · HN ↗
        I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.

        Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance

      2. fasterik · · focus · HN ↗
        I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.
        1. globular-toast · · focus · HN ↗
          You could have also copied an SVG renderer that implements the whole spec from whatever open source project the model copied it from.
          1. fasterik · · focus · HN ↗
            It didn't copy any source code from any external projects. I had it write a stratified sampling renderer for ground truth, then had it implement feature by feature by matching the pixels. Unless you mean it "copied" it in the sense of third-party code being part of the training data. I don't think that definition of "copy" makes any sense given how these models represent embeddings. It would also imply that humans are "copying" the things they've learned from.
            1. sashank_1509 · · focus · HN ↗
              I don’t think we should hold humans and models to the same standards. Humans have a very small working memory. Most humans cannot reproduce code they wrote even a year back exactly.

              A model can reproduce large swaths of its training data exactly. It’s a different algorithm that powers its learning process (it’s why it needs trillions of tokens to even learn basics of language).

              If there was a spectrum from copying on one end to creative production inspired from something else on the other end, the human generally lies heavily on the right end, while the model is much more on the left, that gap is large enough, that yes the model is in some sense “copying”.

              1. XajniN · · focus · HN ↗

                [dead]

              2. fasterik · · focus · HN ↗
                Humans are also capable of memorizing long texts and sequences of numbers. It's more a question of scale than a fundamental difference. The point is that LLMs are like human brains in the sense that they're not databases that store text, even though they can and do memorize things.

                Just because a model can reproduce parts of its training set doesn't mean that's what it's doing when it solves a programming problem. It can also reason about the problem, draw on its knowledge of algorithms and data structures, write tests targeting APIs it's never seen before, generate synthetic data and run experiments, etc. etc. Also, the claim that it can reproduce large swaths of its training data verbatim is an empirical one. I would be surprised if it could even reproduce 0.1% of the books it's ingested, for example.

                Saying that AI can't do anything but copy or steal from humans seems to be a rhetorical technique used by people who are still unaware or in denial about the capabilities of the agentic systems released in the past few months. They can now one-shot theorems and programming problems in a few minutes that would be difficult and time-intensive for even the 99.9th percentile human expert.

                1. sashank_1509 · · focus · HN ↗
                  Empirically I just asked Astra without using the internet to exactly reproduce the first paragraph from Hamlet, Othello, Pride and Prejudice, Sartor Resartus by Thomas Carlyle), Sybil by Benjamin Disraeli (1845). GPT-6 got everything correct: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ac2a72c-e8bc-83e8-8f42-da8c7542decc?ogimg=plain" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ac2a72c-e8bc-83e8-8f42-da8c7542de...

                  No human on earth can copy at this scale. Yes AI is beyond a database lookup, it does have reasoning on top of this knowledge, but my original claim that if there is a spectrum between copying with minimal changes and creative inspiration with minimal copying, AI is one the copying end while humans are on the creative end. Humans are very bad at reproducing anything verbatim, even if they wrote it like a week ago.

                  1. fasterik · · focus · HN ↗
                    Obviously the scale of retrieval from memory is different, I&#x27;m not disputing that. The empirical question I was referring to is what percent of its training set the model is capable of reproducing. What looks like &quot;large swaths&quot; of text to us could be less than one in a million for all we know.

                    I&#x27;m skeptical that there is a single spectrum like you&#x27;re describing. It&#x27;s not well defined. Say a human and an LLM prove a new theorem independently (without external help, i.e. from their own neural weights and reasoning). How do we measure how much each of them copied from previous work, as opposed to having learned from or been influenced by it?

        2. bitexploder · · focus · HN ↗
          It is also easier than ever to build specs and have nice easy to maintain projects. It just doesn&#x27;t happen magically via few shot prompts :)
      3. fignews · · focus · HN ↗
        Have you considered that maybe this is a reflection of your skills rather than that of the LLM?
        1. runarberg · · focus · HN ↗
          Meaning OP is a better programmer then a statistical model which produces the most probable results?
          1. fignews · · focus · HN ↗
            Meaning a bad workman blames his tools.
            1. nvme0n1p1 · · focus · HN ↗
              A bad workman blames his tools.

              A good workman shuts up and finds better tools without complaining.

              1. runarberg · · focus · HN ↗
                An expert workman knows which tools to avoid, and explains it to their peers why to avoid said tools.
        2. sirsinsalot · · focus · HN ↗
          Have you considered it isn&#x27;t?

          Save your &quot;you&#x27;re holding it wrong&quot; if you&#x27;re not going to suggest how to hold it.

          Cult speak escape hatches are intellectually lazy.

          1. williamcotton · · focus · HN ↗
            I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.

            Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!

            What was your process?

            1. hirvi74 · · focus · HN ↗
              Do you mind sharing any code from what you have produced? People talk about LLM successes and failures, but what&#x27;s there to really talk about when the code can speak for itself?

              In case it is unclear, I am genuinely curious. I have great success with chatbots, but vibing coding has never gotten me further than a proof-of-concept.

              1. [deleted] · · focus · HN ↗

                [deleted]

              2. williamcotton · · focus · HN ↗
                Here’s an example:

                <a href="https:&#x2F;&#x2F;williamcotton.github.io&#x2F;datafarm-studio" rel="nofollow">https:&#x2F;&#x2F;williamcotton.github.io&#x2F;datafarm-studio

                Some demos of the above charting language:

                <a href="https:&#x2F;&#x2F;williamcotton.github.io&#x2F;algraf&#x2F;demos" rel="nofollow">https:&#x2F;&#x2F;williamcotton.github.io&#x2F;algraf&#x2F;demos

                WASM, in browser editor, LSP, and more.

          2. fignews · · focus · HN ↗
            Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D

            <a href="https:&#x2F;&#x2F;github.com&#x2F;NousResearch&#x2F;hermes-agent" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;NousResearch&#x2F;hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn&#x27;t load for me presumably because it can&#x27;t handle this scale of commits. But I estimate ~1K commits per day on average.

            There&#x27;s a blog entry <a href="https:&#x2F;&#x2F;nousresearch.com&#x2F;refactoring-hermes-with-1393-agents" rel="nofollow">https:&#x2F;&#x2F;nousresearch.com&#x2F;refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.

            I guess they know how to hold it?

            1. chmod775 · · focus · HN ↗
              You&#x27;re using quantity metrics to answer a quality question.

              I had a look at the kind of issues that are reported at that project (there&#x27;s 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.

              If you really want to check some quantity metrics to try to reason about code quality, look at whether &quot;fix&quot; PRs are overall LOC neutral or negative (not counting tests). In this project, almost every &quot;fix&quot; is an addition. Worse, almost every fix is more branching.

              If almost every PR is some sort of fix, and most of them add branching, and there&#x27;s thousands of them weekly... That leads to only one place and I want to be nowhere near it.

              1. nvme0n1p1 · · focus · HN ↗
                Agreed. As they say, quantity has a quality all of its own.

                Show me an AI that adds features by deleting code (<a href="https:&#x2F;&#x2F;www.folklore.org&#x2F;Negative_2000_Lines_Of_Code.html" rel="nofollow">https:&#x2F;&#x2F;www.folklore.org&#x2F;Negative_2000_Lines_Of_Code.html) and I&#x27;ll pay attention.

            2. northstar702 · · focus · HN ↗

              [dead]

            3. diek · · focus · HN ↗
              I find it funny that the original complaint was: &quot;AI made a mess of the codebase&quot;.

              Your response was: &quot;Well you&#x27;re not doing it right, but these hermes devs know what they&#x27;re doing&quot;.

              But the blog post you linked to shows their prompt, which is:

              &gt; I want god files broken up. I want simplification across the board. I want unification of helpers and methods that can be reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want interpretability of the codebase and how things connect to each other up.

              So it sounds like AI made their code a mess too. They then tried to make the point of how much money they saved cleaning up the code with AI, that AI made a mess of to begin with.

              And if you look at the merged PRs on that project, a ton of them are bug fixes... to the code the AI wrote. And that&#x27;s been my personal experience too: AI creates a huge amount of churn in a codebase. Just vast amounts of PRs fixing code that the AI itself wrote.

        3. catlifeonmars · · focus · HN ↗
          It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.

          I’m saying it’s probably multiple factors and both you and GP are right.

      4. satvikpendem · · focus · HN ↗
        When and what model?
      5. Keyframe · · focus · HN ↗
        I&#x27;m reading y&#x27;all comments and it seems we&#x27;re still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I&#x27;ve been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you&#x27;ve guys mentioned. From good, to bad, to ugly. It&#x27;s like there&#x27;s no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
      6. rickydroll · · focus · HN ↗
        I&#x27;ve had a different experience. My code is better. The ease of refactoring and practicing very defensive programming without getting burned out from the repetitive boilerplate. It all comes together better than it did when I was programming professionally too many moons ago. IFF I don&#x27;t let the LLM toddler run amok. :-)

        Seriously, I find I need to slow down the rate of change. I don&#x27;t move forward until I understand the change proposed and have updated the docs. At the same time, I find that keeping up with the LLM&#x2F;agent is exhausting. 3 hours with an LLM leaves me as tired as 6 hours with a keyboard had previously. I find that coding when tired or fuzzy yields code that shouldn&#x27;t have been written in the first place. Sadly, once it&#x27;s been written and debugged, the temporary fix becomes permanent.

      7. cyanydeez · · focus · HN ↗
        I&#x27;ve adopted open source projects and have been working exclusively in languages I have very little experience with: typescript and go.

        I&#x27;m using local models, and they go slow enough that I have no trouble following along with what they&#x27;re doing; but visually, both go and typescript, along with react native, make me puke. So I wouldn&#x27;t be able to do this without AI.

        I describe how to do it in my comment history, but it&#x27;s basically a Super-TDD along with some custom engineering harness.

        I don&#x27;t want to say skill issue, but the same way you can give a chain saw to a teenager and one to a skill craftsman, well, AI can obviously create whatever you want it to do.

        I think some of the variety is simply how fast SOTA models pump out garbage that you simply have to close your eyes because it&#x27;s not sensible to just watch characters flow across the screen.

        Almost all the coding I&#x27;m doing via AI is just faster than readable. But I can see the thinking traces and I stop to model when it&#x27;s obvious it doesn&#x27;t understand my intent, etc.

        So I&#x27;m not doubting you created garbage. I&#x27;m just doubting that it&#x27;s a product of soley AI use.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.