‹ BackHN Continuity

Thread

The state of SIMD in Rust in 2026

209 points · 65 comments · verdagon

  1. ack_complete · · focus · HN ↗
    AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.

    The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.

    1. brohee · · focus · HN ↗
      "The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs."

      Are those not based on standard ARM cores or ARM documentation on its cores not deep enough for your purpose?

      1. ack_complete · · focus · HN ↗
        Both. The specific documentation I'm referring to is their optimization guides, such as the Cortex-A72 optimization guide:

        <a href="https:&#x2F;&#x2F;support.arm.com&#x2F;documentation&#x2F;uan0016&#x2F;a&#x2F;" rel="nofollow">https:&#x2F;&#x2F;support.arm.com&#x2F;documentation&#x2F;uan0016&#x2F;a&#x2F;

        This has detailed information on latencies and throughput, which are important when optimizing SIMD code. But ARM doesn&#x27;t publish optimization guides for all their cores.

        On top of that, the cores are often modified in significant ways. Snapdragon CPUs, for instance, have used modified Cortex cores in the past and can have performance differences from the original core.

        To be fair, Intel&#x27;s been slacking a lot on this too, not even bothering to update their own optimization guide for their latest cores. But that&#x27;s made up for the community mining this information in great detail on sites like uops.info, and also being a lot less x86 cores to deal with. On the ARM side, there&#x27;s practically not much analogous other than the Apple M1 microarchitecture analysis.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.