‹ BackHN Continuity

Thread

AI coding has made CI a bottleneck, so we reworked ours to keep up

317 points · 408 comments · julian_digital

  1. aliclark · · focus · HN ↗
    In my case it's not the CI that's the bottleneck. It's the human testing side. Does it work, sure. But does it actually do the thing we want (and more importantly) does it do it in a way our customers will understand and actually like?
    1. shykes · · focus · HN ↗
      Think of it as a layered problem. If the bottom layer (CI) cannot keep up with the output of agents, then solving problems at a higher layer - like user experience checks - will be exponentially slower and less reliable. Kind of like how optimizing tight inner loops makes your whole program faster.
      1. Sharlin · · focus · HN ↗
        Uh, no. If stage X is the bottleneck, it's immaterial how much you speed up stage Y.
        1. shykes · · focus · HN ↗
          You're assuming QA reviews won't be fully automated, and triggered from a CI pipeline.
          1. bluGill · · focus · HN ↗
            QA is not automatable. QA is about finding the things that you didn't think of and so didn't write a test for it. That many people think qa is not important shows in the bad software we have.
            1. shykes · · focus · HN ↗
              > QA is about finding the things that you didn't think of and so didn't write a test for it

              QA is about quality. That it has been increasingly used to mean "repetitive manual testing only" is part of the very trend you decry of qa being undervalued.

              Fundamentally QA is about two groups of people collaborating in an adversarial process ("you build it we break it") to achieve a common goal: a better product. QA by definition requires a relationship of equals, or the process ceases to be adversarial and becomes useless ceremony. If you're not able to assemble two distinct groups of people, one person can wear both hats (ie you test your own software) but you'll have more blind spots.

              How much of QA is done by a human or automated, is and has always been an implementation detail.

              If you're in charge of QA, you're in charge not just of finding new problems, but preventing regressions as well. And you're responsible for embedding as much of that work into the devloop itself, so that builders can find and remediate problems earlier and with less pressure on your limited time. How do you achieve that? automation.

              I started my career doing QA and 99% of what I was doing could be automated by an AI today. Today I still do QA on my own product- I just spend all my time on the remaining 1% and my product is better as a result.

              1. bluGill · · focus · HN ↗
                Only is a bit too strong, but it was a lot of repetitive. There has long been a trend to try to automate it - I sat in sales pitches for automation tools in the 1990s, and we were already had long been writing scripts to automate a lot of testing. While the tools have got better, the only advance since the 1990s that I have heard of is automated unit tests (which were being developed then, but hadn't yet spread). We are slightly better at using the tools, but automation of testing is a very old idea.

                Despite all that, it was understood if you want a high quality product you will be spending more than half of your development budget on testing.

                I have long ago concluded that it is not possible to test my own code - and nobody else can test their own code either. I know how to test code - like you I started in QA - but I can't test my own. I have too many blind spots because I'm too close to how it works.

                1. shykes · · focus · HN ↗
                  Yes the blind spots are a hard psychological barrier...

                  I actually think that for a lot of software, AI can really help with that blind spot, even beyond regressions, in ways that don't replace human testers but are complementary.

                  For starters, not all software is used directly by humans. A good chunk was already consumed by APIs, and now an even larger chunk will be consumed by AIs. In those cases, it's very possible that AIs will be better at QA than humans, even for original problem discovery. After all, they are the users.

                  Even for human-facing CLIs (the kind of software I develop these days), it's trivially easy for AIs to interact with the software, and in my experience the newer models are shockingly good at understanding the principles and conventions of CLI usability, at least in the Unix world. I routinely run CLI changes through a gauntlet of AI agents, and their feedback is genuinely good.

                  TLDR I think in these debates over how automatable QA really is, we often tend to forget that there is a lot of software out there, and not all software is tested the same way.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.