‹ BackHN Continuity

Thread

AI coding has made CI a bottleneck, so we reworked ours to keep up

317 points · 408 comments · julian_digital

  1. classictraffic · · focus · HN ↗
    > Moving our workloads off GitHub Actions to third-party runners with faster CPUs, higher-performance storage, and better cache infrastructure gave us faster machines to run the same pipeline on

    Yeah, was not surprised to read this. Actions is convenient if you already use GitHub, but it can also be pretty slow. Given reliability is also a major issue with GitHub these days I expect to see more orgs moving to different pipelines

    1. speedgoose · · focus · HN ↗
      I wonder how much faster GitHub actions could be if they weren’t running on Azure. Because Azure is either slow or very expensive.
      1. chris_money202 · · focus · HN ↗
        I think its actually because GitHub isn't solely running on Azure. They are currently using a mix of on prem, AWS, and Azure. The Azure migration has been a challenge with Azure running into capacity issues due to AI load.
      2. stackskipton · · focus · HN ↗
        It's not Azure as much as self-hosted runners are basement bin servers on clearly massively oversubscribed machines.
      3. rtpg · · focus · HN ↗
        to be fair all CI providers I've had the pleasure of working with have gnarly performance profiles for the boxes they provide.

        I don't think it's out of malice, but I do feel uncomfortable with the fact that the people who sell me the CI coordination software also sell me the minutes for the boxes that run the CI software. There's _some_ alignment of interests but not as much as I would want!

    2. dneri · · focus · HN ↗
      We switched our Actions workload to blacksmith.sh (not affiliated) and have been pretty happy with how fast and inexpensive they are. I wouldn't be surprised to see this trend continue.
    3. wereHamster · · focus · HN ↗
      I wonder why larger companies don't use self hosted github runners. You can buy a pretty beefy machine (TB of RAM, 256 cores, fast NVMe disk) and tests will run faster than on any hosted platform. Plus you don't have to shard as aggressively because more fits into one machine, benefit of shared page cache, shared persistent disk, can easily preserve working directory (for example my pnpm install takes 0 seconds, because the node_modules folder is already present from the previous run).
      1. oblio · · focus · HN ↗
        > I wonder why larger companies don't use self hosted github runners.

        Because they're running away from on-premise and the associated Ops teams.

      2. Lucasoato · · focus · HN ↗
        Microsoft is pushing so hard to get companies far away from the on-premises world. Once they're in the Cloud, there's too much vendor lock-in for big enterprises to go away.
        1. deaux · · focus · HN ↗
          It's sad that even so many greenfield businesses still default to it. There is no longer a need to. The barrier to alternatives used to be technical, now that's gone. 99.9% of web apps' CI needs are solved by a $10/month box.
    4. DanielHB · · focus · HN ↗
      My company did this a couple of years ago to some machines we host ourselves. It is a bit of trouble to set up but it is not that bad and can easily save a 1000+ dollars per month since GHA runners are very expensive.

      It is that grey zone where it is kinda worth it to pay someone to do it, but also might not be worth it the headaches of managing that person and the infra (like what happens if they go on vacation). The improved speed is the thing that tilts the balance.

    5. idkasam · · focus · HN ↗
      Agree, and it's already happening. Mostly smaller orgs so far though, they can just swap the runner label and move on. Larger orgs are starting to look into it but it's slow, lots of red tape (security review, procurement, "we're already paying GitHub").
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.