I’m currently working on adding better SIMD support to Dart (<a href="https://github.com/dart-lang/sdk/issues/64170" rel="nofollow">https://github.com/dart-lang/sdk/issues/64170) and I have a question for the author or others here.
Does Rust or any other language support customizing the compiler so that interprocedural analyses can track custom subsets of, for example, doubles so that the compiler can choose the most efficient instruction sequence for example for min/max? If we know a double is never NaN then we can emit only one instruction on x86, but have to emit one more on arm64. If we know a double is never zero and never NaN, we can emit a single instruction on both.
This whole conversation between relaxed SIMD and deterministic SIMD seems to only exist because our compilers are not smart enough and/or their whole program analyses don’t support any plugin-like capabilities.
There are other examples where if we know a SIMD bitmask is canonical (all 1s per lane) then we can implement horizontal reductions more efficiently. This is very niche and I doubt that any language supports interprocedural analyses with such a rich domain, so it feels like a hole in the programming language space.
But I suspect you're overvaluing the potential savings. Knowing when a float is 0.0 or NaN beforehand is almost entirely impossible, except for the most trivial of cases - like when you first initialize a variable or first enter a loop. Everything after that is very hard or impossible with floats as they are.
Those cases can be const folded at compile time.
Those cases are never a measurable bottleneck.
The closest thing I know of in the realm of the optimization you're curious about is Rust NonZero* variants, but they're used for enum compression afaik.
I’m not sure I agree on the impossible part, I feel like a sufficiently smart interprocedural analysis that also implements range analysis interprocedurally could prove a lot to where it becomes useful.
I guess what I would like to see is SIMD libraries being able to confidently say nobody needs to use intrinsics (or differentiate between relaxed/normal SIMD on the user API level) because the language + high level SIMD APIs are smart enough to choose the right implementation.
IIRC IEEE min/max with proper NaN handling needs 8 instructions on x86 vs 1 on arm64 I find it very sad that we apparently haven’t really solved that yet without forcing the user to use different APIs.
That vminpd+vminpd+vorpd actually does handle signed zero properly! Screws up NaN payloads to the max though. Can be easily extended to canonicalize the NaN with 2 instrs + constant though of course. (which ends up at the same number of instrs as your proper impl (albeit with worse port distribution and latency), but you get to have a canonical NaN!)
Hit upon <a href="https://github.com/llvm/llvm-project/issues/217376" rel="nofollow">https://github.com/llvm/llvm-project/issues/217376 while playing around with proper minimumnum, failing to SMT-verify whatever version of LLVM I had; did find a funky working 6-instr (+ constant) version though:
vpandn ymm2, ymm1, ymm0
vpcmpeqd ymm2, ymm2, 0x80000000 # whether ymm0 is -0 and ymm1 is +0 (or other cases that magically don't cause issues)
vcmpltpd ymm3, ymm0, ymm1
vcmpunordpd ymm3, ymm3, ymm1 # regular NaN-is-larger ymm0<ymm1
vpor ymm2, ymm2, ymm3
vpblendvb ymm0, ymm1, ymm0, ymm2
NonZerof32::from_bits(1) multiplied with itself is zero.
Doing range analysis needs the language to support it at compile time, and the dev to specify what range it is.
The only 'stable' thing i can think of is a type for 'greater-eq-one' using only addition and multiplication. Practically every other operation breaks most of the type knowledge up to that point.
Funnily enough, intrinsics have become almost trivial to write, with agentic tooling. I used to be an ISPC advocate but now I'm finding myself gravitate to direct use of intrinsics with bespoke dynamic/runtime dispatching more than ever.
Yes, Highway is pretty nice but also quite elaborate when it comes to dealing with multiple vectorized versions and dispatch. The macros burn my eyes still.
modulovalue · · focus · HN ↗
Does Rust or any other language support customizing the compiler so that interprocedural analyses can track custom subsets of, for example, doubles so that the compiler can choose the most efficient instruction sequence for example for min/max? If we know a double is never NaN then we can emit only one instruction on x86, but have to emit one more on arm64. If we know a double is never zero and never NaN, we can emit a single instruction on both.
This whole conversation between relaxed SIMD and deterministic SIMD seems to only exist because our compilers are not smart enough and/or their whole program analyses don’t support any plugin-like capabilities.
There are other examples where if we know a SIMD bitmask is canonical (all 1s per lane) then we can implement horizontal reductions more efficiently. This is very niche and I doubt that any language supports interprocedural analyses with such a rich domain, so it feels like a hole in the programming language space.
athrowaway3z · · focus · HN ↗
But I suspect you're overvaluing the potential savings. Knowing when a float is 0.0 or NaN beforehand is almost entirely impossible, except for the most trivial of cases - like when you first initialize a variable or first enter a loop. Everything after that is very hard or impossible with floats as they are.
Those cases can be const folded at compile time.
Those cases are never a measurable bottleneck.
The closest thing I know of in the realm of the optimization you're curious about is Rust NonZero* variants, but they're used for enum compression afaik.
modulovalue · · focus · HN ↗
I guess what I would like to see is SIMD libraries being able to confidently say nobody needs to use intrinsics (or differentiate between relaxed/normal SIMD on the user API level) because the language + high level SIMD APIs are smart enough to choose the right implementation.
IIRC IEEE min/max with proper NaN handling needs 8 instructions on x86 vs 1 on arm64 I find it very sad that we apparently haven’t really solved that yet without forcing the user to use different APIs.
orlp · · focus · HN ↗
If you want propagating NaNs but don't care about signed zero or NaN payload/sign, you can use
What I do in Polars is a bit different, there for propagating NaNs I do this isn't fully optimal on x86-64 but it's fairly simple and autovectorizes decently on various platforms, here's AVX2:dzaima · · focus · HN ↗
Hit upon <a href="https://github.com/llvm/llvm-project/issues/217376" rel="nofollow">https://github.com/llvm/llvm-project/issues/217376 while playing around with proper minimumnum, failing to SMT-verify whatever version of LLVM I had; did find a funky working 6-instr (+ constant) version though:
[deleted] · · focus · HN ↗
[deleted]
dzaima · · focus · HN ↗
athrowaway3z · · focus · HN ↗
Note that:
NonZerof32 * NonZerof32 -> NonNanf32
NonZerof32::from_bits(1) multiplied with itself is zero.
Doing range analysis needs the language to support it at compile time, and the dev to specify what range it is.
The only 'stable' thing i can think of is a type for 'greater-eq-one' using only addition and multiplication. Practically every other operation breaks most of the type knowledge up to that point.
nnevatie · · focus · HN ↗
Funnily enough, intrinsics have become almost trivial to write, with agentic tooling. I used to be an ISPC advocate but now I'm finding myself gravitate to direct use of intrinsics with bespoke dynamic/runtime dispatching more than ever.
janwas · · focus · HN ↗
In C++, one can also have portable intrinsics plus simpler runtime dispatching using our Highway library :)
nnevatie · · focus · HN ↗
Yes, Highway is pretty nice but also quite elaborate when it comes to dealing with multiple vectorized versions and dispatch. The macros burn my eyes still.
janwas · · focus · HN ↗
I do think we've converged on the best that's possible in C++.