‹ BackHN Continuity

Thread

The state of SIMD in Rust in 2026

209 points · 65 comments · verdagon

  1. Archit3ch · · focus · HN ↗
    Hot take: there is no portable SIMD.

    You can either have performance (=write manual ASM for each platform), or portability, but not both.

    What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)

    1. raphlinus · · focus · HN ↗
      You've got a point but are overstating it considerably. There is a big gap between just autovectorization and the portable primitives a library like Highway or Fearless SIMD will give you. For example, I haven't seen autovectorization do select or swizzle.

      But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.

      So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.

      Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.

      As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.

      1. Asmod4n · · focus · HN ↗
        The issue with libraries which offer you portable simd is that the auto vectorizer of the compiler will likely generate faster code.
        1. vlovich123 · · focus · HN ↗
          The autovectorizer afaik rarely emits optimizations for the different vector units to support + efficiently caches the CPUid check to happen once on program start. It’s a good baseline but the continuum (today) is scalar -> auto vectorized -> portable SIMD -> hand rolled explicit. That portable SIMD lets you bridge into hand rolled explicit ergonomically is a power auto-vectorization doesn’t have. Either the compiler does it or doesn’t but you have no way to even have a check that says “fail to build the program if this function isn’t vectorized”. This is important if you’re relying on that property and someone accidentally adds a data dependency and breaks the optimization without you realizing. Portable and explicit SIMD don’t have this problem by definition.
          1. Asmod4n · · focus · HN ↗
            Well, let’s hope rust is better than std::simd from c++ here which lacks the capability to emit some intrinsics.

            Gcc and llvm can tell you if they can’t Auto vec a function, maybe rust could turn this into an error at comptime.

        2. horseloverthin · · focus · HN ↗

          [dead]

        3. simonask · · focus · HN ↗
          This isn't true, often even in trivial cases. Auto-vectorization is actually quite fragile in 2026. The reason is that it's subject to (a) scalar float semantics (i.e. the resulting code must not produce different results from the scalar version), and (b) a number of opaque compiler heuristics that sometimes work out, sometimes don't.

          For example, consider you want to compute the average of a list of floats. The compiler cannot autovectorize this, because float addition is not commutative. However, it's much faster to do component-wise addition in groups, then a horizontal sum at the end, and then divide. Whether it matters depends on your use case, and the compiler unfortunately can't read your mind, so it has to be conservative.

          1. astrange · · focus · HN ↗
            If you try writing SIMD by hand autovectorization can mess it up, eg if you have to write a scalar trailing loop then it might try to autovectorize it.
        4. throawayonthe · · focus · HN ↗
          but like, this isn't the case
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.