You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
Sure, in the same way that there is no such thing as portable code at all. The result will be suboptimal, but it will still be better than not having it.
Said it before and will say it again if binaries were distributed using a bytecode then the host o/s would and should be able to produce optimal binary when loading into memory.
Archit3ch · · focus · HN ↗
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
xboxnolifes · · focus · HN ↗
MiroslavPokorny · · focus · HN ↗