You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
Starts to get a bit philosophical on what constitutes "portable" but JIT compilers would emit an opcode based off of whatever the frontend/IR is saying to do surely?
That assumes no per-platform optimisation, which most JIT compilers will do. I can only speak to the dotnet platform as that's what I am using the most at the moment, but its JIT will spot certain functions, like Vector512.LoadAligned and replace it with the the SIMD equivalent. So, it's not just IR op-code to CPU op-code, it's looking for patterns-of-code that can be made more efficient at JIT compile-time.
That still just gets you autovectorization, and generally locks you out of the performance you could have with direct SIMD intrinsics.
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
C# definitely does [1] and has intrinsics that allow auto-fallback (in the JIT compilation phase) to the most supported SIMD instructions (and in-software implementation if non are); it has a lot of low-level coding primitives that are picked up by the JIT compiler and optimised. I would assume JVM based languages too.
How does the JVM spot that the implementations should be converted to direct SIMD instructions? On dotnet, MS have the [Intrinsic] method-attribute that allows them to mark certain methods that the JIT knows about and do a compile-time replacement. With the JVM being a bit more 'general', how is that achieved? Or, do you just mean there's more implementations of the JVM itself and because of that there's choice?
I mean you get Hotspot on regular OpenJDK, Amazon and Microsoft OpenJDK forks have their own JIT downstream changes, Falcon on Azul, OpenJ9, PTC, Aicas,...
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.
The libraries, yes, but the code you write is portable (at least until you get into squeezing the last few percent and switch to Triton / Helion in case of GPU, and even those are decently portable).
There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.
Getting 2x or 4x performance in your inner loops using a reasonable SIMD library is infinitely better than theoretically getting 8x performance with hand-coded nonportable intrinsics, because the latter is never going to happen in most programs, so the actual point of comparison is scalar code, or autovectorized code at best.
No it's not because it sucks the air out from the actual solution. ISPC more than a decade ago managed to demonstrate close-to-linear speedups for increasing vector sizes, even for branchy code.
Nowadays you can even get AI to write intristics and it works just fine, the portable libraries/autovec aren't really a serious player here.
Portability is also overstated - see the recent shift where Spotify decided to make native Android/iOS apps again instead of React Native. Usually, the number of relevant platforms is somewhere between 2 and 3, so portability concerns are more theoretical than real.
It's a continuum. Some things basically all SIMD implementations support. Want to add 2 4xf32 vectors together? That's pretty easy to do portably.
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
I agree with that historically auto-vertorization does not seem to work reliably. I'm not sure about your broad claim.
Thoughts on an abstraction over ARM and x86, at 128, 256, and 512-bit widths which, either in a manual or automatic way (The latter more challenging) makes your floating point computations 4-16x faster with minimal restructuring? I think that's doable, and a nice goal of SIMD.
You've got a point but are overstating it considerably. There is a big gap between just autovectorization and the portable primitives a library like Highway or Fearless SIMD will give you. For example, I haven't seen autovectorization do select or swizzle.
But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.
So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.
Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.
As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.
The autovectorizer afaik rarely emits optimizations for the different vector units to support + efficiently caches the CPUid check to happen once on program start. It’s a good baseline but the continuum (today) is scalar -> auto vectorized -> portable SIMD -> hand rolled explicit. That portable SIMD lets you bridge into hand rolled explicit ergonomically is a power auto-vectorization doesn’t have. Either the compiler does it or doesn’t but you have no way to even have a check that says “fail to build the program if this function isn’t vectorized”. This is important if you’re relying on that property and someone accidentally adds a data dependency and breaks the optimization without you realizing. Portable and explicit SIMD don’t have this problem by definition.
This isn't true, often even in trivial cases. Auto-vectorization is actually quite fragile in 2026. The reason is that it's subject to (a) scalar float semantics (i.e. the resulting code must not produce different results from the scalar version), and (b) a number of opaque compiler heuristics that sometimes work out, sometimes don't.
For example, consider you want to compute the average of a list of floats. The compiler cannot autovectorize this, because float addition is not commutative. However, it's much faster to do component-wise addition in groups, then a horizontal sum at the end, and then divide. Whether it matters depends on your use case, and the compiler unfortunately can't read your mind, so it has to be conservative.
If you try writing SIMD by hand autovectorization can mess it up, eg if you have to write a scalar trailing loop then it might try to autovectorize it.
Sure, in the same way that there is no such thing as portable code at all. The result will be suboptimal, but it will still be better than not having it.
Said it before and will say it again if binaries were distributed using a bytecode then the host o/s would and should be able to produce optimal binary when loading into memory.
> Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice
IMO this is just people over-indexing on 10 year old GCC. Modern LLVM versions (and even GCC) mostly do good things out of the box. The hard part for the compiler is the vectorization strategy, so using portable intrinsics gives the compiler the shape and it generally does a very good job from there.
Have you used the highway library? It’s portable SIMD done right. It does not rely on auto-vectorization. Instead it gives you a nicer API than using intrinsics, plus machinery to do dynamic dispatch.
Archit3ch · · focus · HN ↗
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
louthy · · focus · HN ↗
Except in languages with a JIT compiler
Tanjreeve · · focus · HN ↗
louthy · · focus · HN ↗
marginalia_nu · · focus · HN ↗
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
louthy · · focus · HN ↗
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
[1] <a href="https://learn.microsoft.com/en-us/dotnet/api/system.runtime.intrinsics?view=net-10.0" rel="nofollow">https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
pletnes · · focus · HN ↗
louthy · · focus · HN ↗
[1] <a href="https://learn.microsoft.com/en-us/dotnet/api/system.runtime.intrinsics?view=net-10.0" rel="nofollow">https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
pjmlp · · focus · HN ↗
louthy · · focus · HN ↗
pjmlp · · focus · HN ↗
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.
louthy · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
Scene_Cast2 · · focus · HN ↗
izacus · · focus · HN ↗
Scene_Cast2 · · focus · HN ↗
There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.
Sharlin · · focus · HN ↗
Pannoniae · · focus · HN ↗
Nowadays you can even get AI to write intristics and it works just fine, the portable libraries/autovec aren't really a serious player here.
Portability is also overstated - see the recent shift where Spotify decided to make native Android/iOS apps again instead of React Native. Usually, the number of relevant platforms is somewhere between 2 and 3, so portability concerns are more theoretical than real.
Archit3ch · · focus · HN ↗
I'm getting 50x faster code with manual ASM. That's the difference between audio code that runs in realtime and code that does not.
The competing implementations use SIMD and native code, autovec works nicely there. I symbolically invert the LinAlg system at compile time.
IshKebab · · focus · HN ↗
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
the__alchemist · · focus · HN ↗
Thoughts on an abstraction over ARM and x86, at 128, 256, and 512-bit widths which, either in a manual or automatic way (The latter more challenging) makes your floating point computations 4-16x faster with minimal restructuring? I think that's doable, and a nice goal of SIMD.
raphlinus · · focus · HN ↗
But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.
So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.
Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.
As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.
Asmod4n · · focus · HN ↗
vlovich123 · · focus · HN ↗
Asmod4n · · focus · HN ↗
Gcc and llvm can tell you if they can’t Auto vec a function, maybe rust could turn this into an error at comptime.
horseloverthin · · focus · HN ↗
[dead]
simonask · · focus · HN ↗
For example, consider you want to compute the average of a list of floats. The compiler cannot autovectorize this, because float addition is not commutative. However, it's much faster to do component-wise addition in groups, then a horizontal sum at the end, and then divide. Whether it matters depends on your use case, and the compiler unfortunately can't read your mind, so it has to be conservative.
astrange · · focus · HN ↗
throawayonthe · · focus · HN ↗
xboxnolifes · · focus · HN ↗
MiroslavPokorny · · focus · HN ↗
MiroslavPokorny · · focus · HN ↗
nnevatie · · focus · HN ↗
adgjlsfhk1 · · focus · HN ↗
IMO this is just people over-indexing on 10 year old GCC. Modern LLVM versions (and even GCC) mostly do good things out of the box. The hard part for the compiler is the vectorization strategy, so using portable intrinsics gives the compiler the shape and it generally does a very good job from there.
kccqzy · · focus · HN ↗