‹ BackHN Continuity

Thread

Pop!_OS bans AI-generated code from much of its codebase

116 points · 167 comments · bundie

  1. brink · · focus · HN ↗
    I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
    1. lucianmarin · · focus · HN ↗
      Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
      1. fasterik · · focus · HN ↗
        I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.
        1. globular-toast · · focus · HN ↗
          You could have also copied an SVG renderer that implements the whole spec from whatever open source project the model copied it from.
          1. fasterik · · focus · HN ↗
            It didn't copy any source code from any external projects. I had it write a stratified sampling renderer for ground truth, then had it implement feature by feature by matching the pixels. Unless you mean it "copied" it in the sense of third-party code being part of the training data. I don't think that definition of "copy" makes any sense given how these models represent embeddings. It would also imply that humans are "copying" the things they've learned from.
            1. sashank_1509 · · focus · HN ↗
              I don’t think we should hold humans and models to the same standards. Humans have a very small working memory. Most humans cannot reproduce code they wrote even a year back exactly.

              A model can reproduce large swaths of its training data exactly. It’s a different algorithm that powers its learning process (it’s why it needs trillions of tokens to even learn basics of language).

              If there was a spectrum from copying on one end to creative production inspired from something else on the other end, the human generally lies heavily on the right end, while the model is much more on the left, that gap is large enough, that yes the model is in some sense “copying”.

              1. XajniN · · focus · HN ↗

                [dead]

              2. fasterik · · focus · HN ↗
                Humans are also capable of memorizing long texts and sequences of numbers. It's more a question of scale than a fundamental difference. The point is that LLMs are like human brains in the sense that they're not databases that store text, even though they can and do memorize things.

                Just because a model can reproduce parts of its training set doesn't mean that's what it's doing when it solves a programming problem. It can also reason about the problem, draw on its knowledge of algorithms and data structures, write tests targeting APIs it's never seen before, generate synthetic data and run experiments, etc. etc. Also, the claim that it can reproduce large swaths of its training data verbatim is an empirical one. I would be surprised if it could even reproduce 0.1% of the books it's ingested, for example.

                Saying that AI can't do anything but copy or steal from humans seems to be a rhetorical technique used by people who are still unaware or in denial about the capabilities of the agentic systems released in the past few months. They can now one-shot theorems and programming problems in a few minutes that would be difficult and time-intensive for even the 99.9th percentile human expert.

                1. sashank_1509 · · focus · HN ↗
                  Empirically I just asked Astra without using the internet to exactly reproduce the first paragraph from Hamlet, Othello, Pride and Prejudice, Sartor Resartus by Thomas Carlyle), Sybil by Benjamin Disraeli (1845). GPT-6 got everything correct: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ac2a72c-e8bc-83e8-8f42-da8c7542decc?ogimg=plain" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ac2a72c-e8bc-83e8-8f42-da8c7542de...

                  No human on earth can copy at this scale. Yes AI is beyond a database lookup, it does have reasoning on top of this knowledge, but my original claim that if there is a spectrum between copying with minimal changes and creative inspiration with minimal copying, AI is one the copying end while humans are on the creative end. Humans are very bad at reproducing anything verbatim, even if they wrote it like a week ago.

                  1. fasterik · · focus · HN ↗
                    Obviously the scale of retrieval from memory is different, I&#x27;m not disputing that. The empirical question I was referring to is what percent of its training set the model is capable of reproducing. What looks like &quot;large swaths&quot; of text to us could be less than one in a million for all we know.

                    I&#x27;m skeptical that there is a single spectrum like you&#x27;re describing. It&#x27;s not well defined. Say a human and an LLM prove a new theorem independently (without external help, i.e. from their own neural weights and reasoning). How do we measure how much each of them copied from previous work, as opposed to having learned from or been influenced by it?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.