‹ BackHN Continuity

Thread

AI coding has made CI a bottleneck, so we reworked ours to keep up

317 points · 408 comments · julian_digital

  1. azkalam · · focus · HN ↗
    If you can pay the setup cost, Bazel will get you build times ~10s with a warm cache for even massive projects.
    1. criemen · · focus · HN ↗
      I lead a bazel conversion for a pretty complex piece of software written in 5+ programming languages and shipping native binaries to all 3 major OSes a few years ago, and it took multiple years to get it done.

      For a less complex project (1 programming language, still shipping to all 3 major OSes), with my knowledge and agents I got the bazel conversion done in 2 weeks.

      The setup cost for bazel just went down by a lot, and I don't think the industry as a whole is aware of that yet.

      1. rokob · · focus · HN ↗
        Were any of those using JS/TS? I've seen Bazel perform wonders with many compiled languages, but I'm not impressed with the JS ecosystem.
        1. criemen · · focus · HN ↗
          Nothing substantial, no. I've never personally experienced builds to be so slow to begin investigating bazel as an option for JS/TS.

          Tsgo, oxlint, caching dependencies etc. what linear outlined in their blog post would be more impactful for the average TS project I've worked on.

      2. teaearlgraycold · · focus · HN ↗
        The entire industry, including its outputs that LLMs are trained on, hasn’t reconsidered what’s easy vs. hard or fast vs. slow. LLMs consistently recommend against code changes because they will take “a weekend”. No, Claude. You will do the work and it will take 20 minutes.
    2. manquer · · focus · HN ↗
      Can't comment on Bazel specifically, but having worked with both nx and turbo, the bottleneck was usually network and disk IOPS rarely compute.

      Even fully cached outputs needs to fetched and read from a remote server[1]. A step n-1 outout fetched from remote cache server need to written to disk and then again read by step n[3] - all disk I/O and network bound operations.

      10s may be achievable/realistic goal in the Java/C++ world where Bazel normally seen. In TS eco-system most people would be over the moon to get into ballpark of 1-2m for a decently large monorepo.

      We should define Build more clearly here, if you mean running just transpile/compile steps or the full series of steps that includes tests (as the linear post here is talking about). It is hard to see even a small sub-set of a large suite of test that require a virtual DOM or a real browser can run in 10s or less.

      [1] Typical for say managed CI setup .

      [3] Common run-of-the-mill frontend + backend stacks in different languages etc.

      1. criemen · · focus · HN ↗
        If you don't need to rebuild anything, bazel can fetch only the final artifact (not the intermediates) from the remote cache.

        Also, if you have persistent CI workers with a persistent bazel instance, you save on some network roundtrips, but that's obviously harder to set up and make bulletproof.

        1. rienbdj · · focus · HN ↗
          Bazel can even skip the final artefact!
          1. criemen · · focus · HN ↗
            well for test runs you kinda need the binary to run, but you're correct, if the job is just "does this build" no download necessary.
        2. manquer · · focus · HN ↗
          Typically you need the intermediates to compute if you need the next one , so you cannot skip to final until you have the intermediaries.

          The final asset/artifact is rarely small either. even best optimized artifacts can be few hundred MB docker image or more commonly multiple image layers running GBs in size .

          each step is a network pull then recompute cache if stale and keep going till end .

          1. sluongng · · focus · HN ↗
            We solve it 2 ways in the Bazel ecosystem: for the intermediate artifacts, we only fetch the digest (hash + size) of the blobs to calculate the merkle tree forward. The blob itself can stay on the remote cache server.

            For the bigger final artifacts, we support using Content Defined Chunking (rolling gear hashing) to only fetch the missing chunks between incremental builds. Binaries executable with stable layout benefits from this quite a lot.

            We are definitely not done with all of the improvements here. But since all the major AI labs are using Bazel, we know that the tools can support “Agent Scale”. <a href="https:&#x2F;&#x2F;webazel.dev&#x2F;" rel="nofollow">https:&#x2F;&#x2F;webazel.dev&#x2F;

            1. manquer · · focus · HN ↗
              &gt; calculate the merkle tree forward. he blob itself can stay on the remote cache server.

              Not sure how that would work with building say a docker image, reproducible builds are pretty hard problem to solve, and caching intermediate layers is not always simple or even doable, we typically still need to publish to a registry which is not the cache server.

              1. sluongng · · focus · HN ↗
                The Bazel ecosystem builds container images not by using Dockerfile, which contains non-reproducible primitives such as RUN and others. We do it by actually constructing the file trees and tarballs manually, then using them to compose the JSON manifest and indices. This is done via smaller hermetic and reproducible Bazel actions and thus enables the ecosystem to scale way beyond what alternative BuildKit-based solutions can.

                <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=biYXmAv4Ppk&amp;t=314s" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=biYXmAv4Ppk&amp;t=314s should be a good talk to study up on the matter. The speaker is now working at Apple.

    3. seniorsassycat · · focus · HN ↗
      Bazel seems to have a lot of tradeoffs, from setup time of the sandbox for each task, to ergonomics that lead folks to maintain parallel &#x27;normal&#x27; tooling.

      Plus, &#x27;with a warm cache&#x27; is doing heavy lifting, what&#x27;s the real cache hit rate for a week of development? Investing in improving the cold build and frequent actions is still important with bazel or any incremental builder.

      I&#x27;m not sure it&#x27;s useful to talk about bazel broadly, it&#x27;s actual performance and behavior comes down to the rules you use. You can configure bazel like turbo&#x2F;nx and cache tsc&#x2F;vitest&#x2F;eslint on each package.json module, and get course cached units that are evicted on every change, or you can use gazelle and target per-file actions which are only invalidated when their dependencies change. But that trades off batching unless you use workers.

      1. rienbdj · · focus · HN ↗
        Most PRs only touch a handful of targets so cache hits are extremely high in practice.
        1. seniorsassycat · · focus · HN ↗
          Depends on which targets and how granular the caching is. - if you touch package.json that might invalidate everything - touch core and you&#x27;ll invalidate everything - touch one file in api and you may run all api tests (see gazelle)

          I ran an experiment where I migrated a package to bazel then replayed a weeks worth of changes and it saved 20%. That&#x27;s nothing to scoff at, but not the headline numbers you see after a full hot build.

      2. sluongng · · focus · HN ↗
        What i have seen on my end is that it’s pretty easy to setup a cloud coding agent with a warm Bazel cache. “Warm” here can means multiple layers: same disk to CoW&#x2F;hardlink, different disk, network disk&#x2F;block devices, same rack&#x2F;datacenter, same AZ, etc…

        And yes, there are a ton of investments going toward Bazel recently to unlock these newer use cases.

    4. hamolton · · focus · HN ↗
      Is this really easy in the AI era? It seems like the repetitive, verifiable work that the agents should be good at.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.