You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
Starts to get a bit philosophical on what constitutes "portable" but JIT compilers would emit an opcode based off of whatever the frontend/IR is saying to do surely?
That assumes no per-platform optimisation, which most JIT compilers will do. I can only speak to the dotnet platform as that's what I am using the most at the moment, but its JIT will spot certain functions, like Vector512.LoadAligned and replace it with the the SIMD equivalent. So, it's not just IR op-code to CPU op-code, it's looking for patterns-of-code that can be made more efficient at JIT compile-time.
That still just gets you autovectorization, and generally locks you out of the performance you could have with direct SIMD intrinsics.
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
C# definitely does [1] and has intrinsics that allow auto-fallback (in the JIT compilation phase) to the most supported SIMD instructions (and in-software implementation if non are); it has a lot of low-level coding primitives that are picked up by the JIT compiler and optimised. I would assume JVM based languages too.
How does the JVM spot that the implementations should be converted to direct SIMD instructions? On dotnet, MS have the [Intrinsic] method-attribute that allows them to mark certain methods that the JIT knows about and do a compile-time replacement. With the JVM being a bit more 'general', how is that achieved? Or, do you just mean there's more implementations of the JVM itself and because of that there's choice?
I mean you get Hotspot on regular OpenJDK, Amazon and Microsoft OpenJDK forks have their own JIT downstream changes, Falcon on Azul, OpenJ9, PTC, Aicas,...
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.
Archit3ch · · focus · HN ↗
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
louthy · · focus · HN ↗
Except in languages with a JIT compiler
Tanjreeve · · focus · HN ↗
louthy · · focus · HN ↗
marginalia_nu · · focus · HN ↗
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
louthy · · focus · HN ↗
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
[1] <a href="https://learn.microsoft.com/en-us/dotnet/api/system.runtime.intrinsics?view=net-10.0" rel="nofollow">https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
pletnes · · focus · HN ↗
louthy · · focus · HN ↗
[1] <a href="https://learn.microsoft.com/en-us/dotnet/api/system.runtime.intrinsics?view=net-10.0" rel="nofollow">https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
pjmlp · · focus · HN ↗
louthy · · focus · HN ↗
pjmlp · · focus · HN ↗
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.
louthy · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]