‹ BackHN Continuity

Thread

If AI coding is lowering your code quality, you're not managing quality right

129 points · 169 comments · bucket2015

  1. fishfasell · · focus · HN ↗
    I think there's a lot of setup and context required for an AI agent to consistently write good code. Once the agent has these guard rails in place I usually get great quality- far better than what I would write in most cases.

    I think where things get dicey is being able to write in any language. I write and review code in many languages and frameworks I'm not fluent in, so it's hard for me to distinguish between working code and great code. I can spot when the fundamental logic is wrong, but when it comes to "best fit" choices I'm clueless.

    1. this_user · · focus · HN ↗
      The issue is that in order to have the agent write good code, you need to implement standard SWE best practices. But that also means a lot of manual intervention in terms of writing specs, checking acceptance criteria, and reviewing code. So you end up spending a lot of time on managing your agent, which means you won't get a 1000% productivity gain, you get maybe 50 or 100, possible less in some areas and with some issues.
      1. beezlewax · · focus · HN ↗
        50 or 100 seems unlikely. Even with all these improvements, custom setups and guardrails it just isn't that much faster for me.
      2. user43928 · · focus · HN ↗
        A 1000% productivity gain is quite possible on solo greenfield projects.

        At work, with a team and code reviews, the 50%-100% figure seems much more likely.

        This can probably move towards the more spectacular productivity gains as the AI's output becomes more reliable, people realize this, and less time is spend on code review and cleaning up the output.

      3. lolakutty · · focus · HN ↗
        > implement standard SWE best practices

        The thing is, if you follow SWE best practices indiscriminately, then you ll have a shit code base in no time.

        There is no silver bullet, and no replacement for experience and mindfulness.

      4. bigstrat2003 · · focus · HN ↗
        You get 0% productivity gains if you are careful and actually reviewing the code the LLM produces. The only way to actually get the massive productivity gains that AI bros claim is to throw quality out the window.
    2. kuczmama · · focus · HN ↗
      I'm curious as to what guardrails you've tried.

      This is something I have been trying to get right as well. I've attempted to use lots of linting and things like strong typing, duplicate checks, cyclomatic complexity, and robust tests. However, I still happen to find issues, which requires me to look at the code (at least at a high level)

      For example, I can say "Don't repeat yourself, and don't re-write helper functions" and I will even have a duplicate linter check, but inevitably the LLM will always want to re-write a similar yet slightly different helper function. Like it will always want to re-write something small like a trim() or a toString() function in every file.

      1. esprehn · · focus · HN ↗
        Have you tried something like "Always consult the utils/ package before writing helper functions. When adding a new generic helper function justify it in your design or PR description."

        I have better luck telling it positive things rather than lots of "never do X" style things.

        1. kuczmama · · focus · HN ↗
          That's a good idea to give more positive instructions as opposed to negative instructions. I think you've stated it well, I suppose the problem with negative instructions is that the LLM doesn't know what to do instead.

          "Never re-write a helper function" vs "Always search for helper functions before writing one" the "never... " one doesn't tell the LLM what to do, so it would have to make the logical leap from not re-writing to knowing that it should search. While it's a minor leap to make in isolation, I suppose stacking many negative rules in an AGENTS.md would assume that every time it will always make that logical conclusion on what to do.

      2. bucket2015 · · focus · HN ↗
        I find that if I leave an instruction in AGENTS.md to "do not do X", there's a good chance the agent will forget it.

        But if I add a separate post-implementation pass to "find and fix X" by the agent, it'll usually find and fix the issues.

        So I've started doing it for everything from naming conventions to duplicate code to other problems. It does cost more tokens, but now I get less frustrated at having to fix basic issues in the PRs.

      3. ytoawwhra92 · · focus · HN ↗
        > Don't repeat yourself, and don't re-write helper functions

        It's worth reflecting on why these things are important to you and whether they remain important in an agent-developed codebase.

        1. sevenseacat · · focus · HN ↗
          Yes, they are both still very important for consistency throughout your codebase and any user interface for it.
          1. ytoawwhra92 · · focus · HN ↗
            Consistency of behaviour and UI can be tested.
    3. nicce · · focus · HN ↗
      I would say that it is like gardening. If you let them go havoc from the start, the weed will take over. If you keep focusing on removing the weed and enforce specific standards and practices over the code base and it keeps growing, over time LLMs start to suddenly follow that and they don't make so much slop anymore. At least that is my experience. But I force specific audit agent after every added feature which says them to force compliance with AGENTS.md and check the consistency with the code base.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.