> The interface conversion and type switch look like they should be inefficient, but the compiler-side implementation of simd specializes code and optimizes away the type switch.
I don’t understand this - how is it able to if the same go binary might run on unknown types? I’m assuming what it means is that the switch is implemented efficiently due to CPU branch prediction? I know fearless SIMD is doing cool stuff with static dispatch so that the feature set is checked just once at program start - is that what it means it’s doing under the hood? Very unclear.
It creates multiple versions of functions referencing SIMD and lifts the dispatch switching cost to their callers.
> The AST rewrite creates multiple specialized copies of functions, variables, and types that mention simd types, where simd types are replaced with references to size-specialized types in simd/internal/bridge. Each of these bridge types is defined as an archsimd type, but with a restricted set of methods. The specialized functions, variables, and types acquire a suffix of the form @simdNNN, where NNN is either a vector length (128, 256, or 512) or 0, indicating emulation. Functions that mention simd internally, but not in their signature, are converted to wrappers that switch on the SIMD level detected at program start, and call the appropriate specialized version of that function. Specialized functions call other specialized functions directly without dispatch overhead (and perhaps with inlining). This rewrite strategy was chosen as a compromise between code duplication and SIMD performance; the overhead is hoisted as high as necessary to avoid dispatch within SIMD computations, but not higher. If SIMD dispatch appears “too low” in a computation, a gratuitous mention of a simd type will move it upwards, as in this example:
Question on the multi-versioning approach: if the AST rewrite creates N specialized copies of every function that mentions the SIMD types, doesn't binary size scale with the number of SIMD-touching functions? For generic hot paths (helpers parameterized over vector widths, instantiated across many call sites) that could multiply code size noticeably. Is there dedup when two specialized copies would be identical, or has anyone measured the binary-size cost on a real codebase?
Binary size only scales for the simd-touching platforms, with scaling depending on platform. Each platform gets emulation+number-of-variants. Emulation is there as an option, primarily for testing (do you trust us? I don't trust us). So wasm is 2x, amd64 is 4x, arm64 is currently 2x but should be 4x in the near future. ppc64 ought to be 2x, maybe 3x, don't know their variants story. Loong64 will be 3x. But only the SIMD-touching stuff is replicated, and N-1 copies of the replicated code are never executed by a given program execution.
SVE and RVV, because they are "scalable", may end up with a larger amount of replication.
Go uses fixed-size stack frames, so spill space is fixed size, etc. We could "just" fix this limitation, or we could generate multiple sizes for what actually exists in the wild, with a masked implementation for longer. So, SVE would end up with 0 (emulation), 128, 256, and 512-masked, plus (as arm64) would also have NEON and NEON-noclmul (Raspberry Pi). I'm not sure what sizes exist for RVV.
vlovich123 · · focus · HN ↗
I don’t understand this - how is it able to if the same go binary might run on unknown types? I’m assuming what it means is that the switch is implemented efficiently due to CPU branch prediction? I know fearless SIMD is doing cool stuff with static dispatch so that the feature set is checked just once at program start - is that what it means it’s doing under the hood? Very unclear.
Scaevolus · · focus · HN ↗
> The AST rewrite creates multiple specialized copies of functions, variables, and types that mention simd types, where simd types are replaced with references to size-specialized types in simd/internal/bridge. Each of these bridge types is defined as an archsimd type, but with a restricted set of methods. The specialized functions, variables, and types acquire a suffix of the form @simdNNN, where NNN is either a vector length (128, 256, or 512) or 0, indicating emulation. Functions that mention simd internally, but not in their signature, are converted to wrappers that switch on the SIMD level detected at program start, and call the appropriate specialized version of that function. Specialized functions call other specialized functions directly without dispatch overhead (and perhaps with inlining). This rewrite strategy was chosen as a compromise between code duplication and SIMD performance; the overhead is hoisted as high as necessary to avoid dispatch within SIMD computations, but not higher. If SIMD dispatch appears “too low” in a computation, a gratuitous mention of a simd type will move it upwards, as in this example:
keel_dev · · focus · HN ↗
dr2chase · · focus · HN ↗
SVE and RVV, because they are "scalable", may end up with a larger amount of replication.
Go uses fixed-size stack frames, so spill space is fixed size, etc. We could "just" fix this limitation, or we could generate multiple sizes for what actually exists in the wild, with a masked implementation for longer. So, SVE would end up with 0 (emulation), 128, 256, and 512-masked, plus (as arm64) would also have NEON and NEON-noclmul (Raspberry Pi). I'm not sure what sizes exist for RVV.
dr2chase · · focus · HN ↗