‹ BackHN Continuity

Thread

The state of SIMD in Rust in 2026

209 points · 65 comments · verdagon

  1. ack_complete · · focus · HN ↗
    AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.

    The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.

    1. officialchicken · · focus · HN ↗
      I really hope this is my last x86-64 CPU, Intel has become an incompetent steward. The "killer feature" for AVX-2(56) at the time of original release was basically lag/jitter-free video playback. IMO, 512 should have never been released for desktop CPUs and restricted to servers. One day I will to switch to a mainstream Neoverse dev box running linux. I also target Cortex-M in rust, so it's got a lot of the typical issues related to missing docs (e.g. bringup of non-heterogenous cores, meaning that M3/M4 still can't be used in a big.little chip)
      1. nixon_why69 · · focus · HN ↗
        > IMO, 512 should have never been released for desktop CPUs and restricted to servers.

        Why not? To save die space?

        1. i80and · · focus · HN ↗
          I think they meant it should have been deployed to all SKUs instead of the very silly ecosystem split Intel did for a long time.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.