You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
Starts to get a bit philosophical on what constitutes "portable" but JIT compilers would emit an opcode based off of whatever the frontend/IR is saying to do surely?
That assumes no per-platform optimisation, which most JIT compilers will do. I can only speak to the dotnet platform as that's what I am using the most at the moment, but its JIT will spot certain functions, like Vector512.LoadAligned and replace it with the the SIMD equivalent. So, it's not just IR op-code to CPU op-code, it's looking for patterns-of-code that can be made more efficient at JIT compile-time.
Archit3ch · · focus · HN ↗
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
louthy · · focus · HN ↗
Except in languages with a JIT compiler
Tanjreeve · · focus · HN ↗
louthy · · focus · HN ↗