You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
> Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice
IMO this is just people over-indexing on 10 year old GCC. Modern LLVM versions (and even GCC) mostly do good things out of the box. The hard part for the compiler is the vectorization strategy, so using portable intrinsics gives the compiler the shape and it generally does a very good job from there.
Archit3ch · · focus · HN ↗
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
adgjlsfhk1 · · focus · HN ↗
IMO this is just people over-indexing on 10 year old GCC. Modern LLVM versions (and even GCC) mostly do good things out of the box. The hard part for the compiler is the vectorization strategy, so using portable intrinsics gives the compiler the shape and it generally does a very good job from there.