I was recently able to get a workload of mine optimized enough on Zen 5 to hit a sustained 6.0 IPC/core (3.0 / thread) at 5.1 GHz. Seeing the a > 99.8% branch prediction rate and a > 99.99% L2 cache hit rate retiring > 1T instructions every 3 seconds feels amazing.
Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.
How do you explicitly tell the CPU to fetch the next batch of data before you start processing the current batch so that the cache is hot when you need it? I imagine there are instructions for that.
There is a prefetch instruction but you also don't need to in obvious loops like that because the CPU has hardware to detect sequential access patterns. IIRC one test showed a CPU could prefetch up to 11 interleaved sequential access patterns with different strides.
bmenrigh · · focus · HN ↗
Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.
ultrahax · · focus · HN ↗
mitxela · · focus · HN ↗
rbanffy · · focus · HN ↗
mitxela · · focus · HN ↗