I write in C++ almost every day but never have the need to optimize for speed. Even when you write straightforward code it's already blazingly fast.
True, but moving from a list of unique polymorphic pointers to a std::variant gains you at least a 2-3x speed up in terms of TLB and cacheline locality. From there, swapping to SOA will net you another 4-8x, so you're looking at nearly 25x improvement by going data first. That may not matter in the unique case of say, games, where rendering a million entities will dwarf the cost of SIMD processing a million entities, but in something like numerical simulations (fluids) or quant it will be warmly welcomed
That, and it frees the compiler from reasoning about virtual inlining, and that the std::variant approach can pack potentially more than one object into a single cacheline. TLBs also work with 4096 byte pages, so 32 polymorphic 128 byte entities may (at the absolute worst case) use 32 distinct pages which requires 32 TLB virtual translations, while the std::variant one uses 1.
The next step of going SOA benefits from all of the above, it just further unlocks you packed quad and oct instructions (AVX256 and 512 depending if you buy AMD or not).
If you take the linked benchmark and use the latest compiler version, the std::variant version is faster. The annoying thing about std::variant (and with some other features of modern C++) is that it generates a bunch of code that the compiler has to optimize away.
Somehow and for some reason rust enums don’t have this problem and are far more ergonomic and easier to work with (not to mention compile times are amazing). I don’t know exactly why it’s better to have it as a first class language primitive and why the compiler has such a problem with std::visit, but clearly c++ meta programming slows down things in a super linear way such that the compiler has problems both from code gen and then optimization.
You don't render a million entities, only the ones that are actually on-screen. But when you do render them, you also want to render them in SoA style, culling them by comparing all the bounding boxes against the view frustum, generating a single big list of draw commands and passing it to draw-indirect one time.
hn_submit · · focus · HN ↗
glouwbug · · focus · HN ↗
cenamus · · focus · HN ↗
glouwbug · · focus · HN ↗
The next step of going SOA benefits from all of the above, it just further unlocks you packed quad and oct instructions (AVX256 and 512 depending if you buy AMD or not).
tcfhgj · · focus · HN ↗
<a href="https://stackoverflow.com/questions/69444641/c17-stdvariant-is-slower-than-dynamic-polymorphism" rel="nofollow">https://stackoverflow.com/questions/69444641/c17-stdvariant-...
creata · · focus · HN ↗
vlovich123 · · focus · HN ↗
jasode · · focus · HN ↗
GCC 11 (2021) std::visit was slower than virtual dispatch.
GCC 12 (2022) optimized std::visit so it can be faster than virtual dispatch.
<a href="https://shubhankar-gambhir.github.io/posts/your-stdlib-implementation-matters-more-than-the-dispatch-pattern/" rel="nofollow">https://shubhankar-gambhir.github.io/posts/your-stdlib-imple...
someonebaggy · · focus · HN ↗
someonebaggy · · focus · HN ↗
jplusequalt · · focus · HN ↗
Modern graphics APIs allow you to render as many objects as you want with a single draw indirect call.