‹ BackHN Continuity

Thread

Pop!_OS bans AI-generated code from much of its codebase

116 points · 167 comments · bundie

  1. brink · · focus · HN ↗
    I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
    1. lucianmarin · · focus · HN ↗
      Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
      1. fignews · · focus · HN ↗
        Have you considered that maybe this is a reflection of your skills rather than that of the LLM?
        1. runarberg · · focus · HN ↗
          Meaning OP is a better programmer then a statistical model which produces the most probable results?
          1. fignews · · focus · HN ↗
            Meaning a bad workman blames his tools.
            1. nvme0n1p1 · · focus · HN ↗
              A bad workman blames his tools.

              A good workman shuts up and finds better tools without complaining.

              1. runarberg · · focus · HN ↗
                An expert workman knows which tools to avoid, and explains it to their peers why to avoid said tools.
        2. sirsinsalot · · focus · HN ↗
          Have you considered it isn't?

          Save your "you're holding it wrong" if you're not going to suggest how to hold it.

          Cult speak escape hatches are intellectually lazy.

          1. williamcotton · · focus · HN ↗
            I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.

            Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!

            What was your process?

            1. hirvi74 · · focus · HN ↗
              Do you mind sharing any code from what you have produced? People talk about LLM successes and failures, but what's there to really talk about when the code can speak for itself?

              In case it is unclear, I am genuinely curious. I have great success with chatbots, but vibing coding has never gotten me further than a proof-of-concept.

              1. [deleted] · · focus · HN ↗

                [deleted]

              2. williamcotton · · focus · HN ↗
                Here’s an example:

                <a href="https:&#x2F;&#x2F;williamcotton.github.io&#x2F;datafarm-studio" rel="nofollow">https:&#x2F;&#x2F;williamcotton.github.io&#x2F;datafarm-studio

                Some demos of the above charting language:

                <a href="https:&#x2F;&#x2F;williamcotton.github.io&#x2F;algraf&#x2F;demos" rel="nofollow">https:&#x2F;&#x2F;williamcotton.github.io&#x2F;algraf&#x2F;demos

                WASM, in browser editor, LSP, and more.

          2. fignews · · focus · HN ↗
            Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D

            <a href="https:&#x2F;&#x2F;github.com&#x2F;NousResearch&#x2F;hermes-agent" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;NousResearch&#x2F;hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn&#x27;t load for me presumably because it can&#x27;t handle this scale of commits. But I estimate ~1K commits per day on average.

            There&#x27;s a blog entry <a href="https:&#x2F;&#x2F;nousresearch.com&#x2F;refactoring-hermes-with-1393-agents" rel="nofollow">https:&#x2F;&#x2F;nousresearch.com&#x2F;refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.

            I guess they know how to hold it?

            1. chmod775 · · focus · HN ↗
              You&#x27;re using quantity metrics to answer a quality question.

              I had a look at the kind of issues that are reported at that project (there&#x27;s 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.

              If you really want to check some quantity metrics to try to reason about code quality, look at whether &quot;fix&quot; PRs are overall LOC neutral or negative (not counting tests). In this project, almost every &quot;fix&quot; is an addition. Worse, almost every fix is more branching.

              If almost every PR is some sort of fix, and most of them add branching, and there&#x27;s thousands of them weekly... That leads to only one place and I want to be nowhere near it.

              1. nvme0n1p1 · · focus · HN ↗
                Agreed. As they say, quantity has a quality all of its own.

                Show me an AI that adds features by deleting code (<a href="https:&#x2F;&#x2F;www.folklore.org&#x2F;Negative_2000_Lines_Of_Code.html" rel="nofollow">https:&#x2F;&#x2F;www.folklore.org&#x2F;Negative_2000_Lines_Of_Code.html) and I&#x27;ll pay attention.

            2. northstar702 · · focus · HN ↗

              [dead]

            3. diek · · focus · HN ↗
              I find it funny that the original complaint was: &quot;AI made a mess of the codebase&quot;.

              Your response was: &quot;Well you&#x27;re not doing it right, but these hermes devs know what they&#x27;re doing&quot;.

              But the blog post you linked to shows their prompt, which is:

              &gt; I want god files broken up. I want simplification across the board. I want unification of helpers and methods that can be reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want interpretability of the codebase and how things connect to each other up.

              So it sounds like AI made their code a mess too. They then tried to make the point of how much money they saved cleaning up the code with AI, that AI made a mess of to begin with.

              And if you look at the merged PRs on that project, a ton of them are bug fixes... to the code the AI wrote. And that&#x27;s been my personal experience too: AI creates a huge amount of churn in a codebase. Just vast amounts of PRs fixing code that the AI itself wrote.

        3. catlifeonmars · · focus · HN ↗
          It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.

          I’m saying it’s probably multiple factors and both you and GP are right.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.