‹ BackHN Continuity

Thread

AI coding has made CI a bottleneck, so we reworked ours to keep up

317 points · 408 comments · julian_digital

  1. torben-friis · · focus · HN ↗
    Here's my constant question:

    Everyone's going so fast that they keep hitting walls. Review, CI, product asking for things, whatever.

    Why have we not seen an improvements in products?

    While every post and thread feels like a 90's wall street office, the new android and iphone ship with fewer features than usual. No indie guys come up with a linux-sized alternative OS. Switch 2 remains unhacked. Windows takes 3 seconds to show the right click menu.

    Is everyone just running full speed in circles or something?

    1. saithound · · focus · HN ↗
      > Why have we not seen an improvements in products? [..] Is everyone just running full speed in circles or something?

      The simplest explanation is that they don't give a flying flamingo about what you or I consider "improvements to products".

      This report is an example.

      There are several changes that modify CI behaviour, where the article gives no corresponding quality measurement.

      They replaced type aware custom lint rules with AST-only static analysis. They don't say anything about what those new rules detect, didn't do old-vs-new rule comparison. They switched the TypeScript check from tsc to tsgo. Again, they are very proud of the performance improvement, but don't seem to care about diagnostic equivalence. The list goes on. They don't even report pass/fail agreement between the old and the new CI. They have 4x more tests, but no idea whether this big test suite works any better than the smaller old one, or even whether it works at all.

      1. plant-ian · · focus · HN ↗
        Are they really adding 2,000 tests a week to their codebase?
        1. Aeolun · · focus · HN ↗
          Anecdotally, codex is very fond of checking strings are equal between UI and test.
          1. Rapzid · · focus · HN ↗
            It's really fond of testing external libraries too lol.

            And dumping tests it needs for intermediate work steps in your suite to run for all of eternity.. Sometimes it'll create these in tmp, but not always!

            1. plagiarist · · focus · HN ↗
              Are we holding it wrong? Mine does the same godawful external library tests. When it is having a particularly stupid day it also tests language features. For example, checking that using a callback executes the code supplied in the callback (it does).
        2. maccard · · focus · HN ↗
          They said all their tests are written by agents so probably, yeah.
        3. nottorp · · focus · HN ↗
          Why not? If you let the LLM run amok, you get 4 one line functions calling each other instead of one 4 line function. Run that for 24 hours and you'll get more than 2000 tests.
        4. Rapzid · · focus · HN ↗
          Yeah, that's what not so many are talking about but a few have pointed out; it's the AI writing wayyyyy too many tests causing the new CI load.

          Unless you prompt them otherwise, the models tend to write WAY too many useless tests and in the most inefficient ways imaginable. This balloons test counts and lines of code to insane levels and frankly, likely, slowly makes it more and more costly for the AI to make future changes.. To the point it can't wrap its context around the code base and effectively make necessary changes.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.