‹ BackHN Continuity

Thread

How did AMD Ryzen get 50% faster in two years?

487 points · 206 comments · ibobev

  1. bmenrigh · · focus · HN ↗
    I was recently able to get a workload of mine optimized enough on Zen 5 to hit a sustained 6.0 IPC/core (3.0 / thread) at 5.1 GHz. Seeing the a > 99.8% branch prediction rate and a > 99.99% L2 cache hit rate retiring > 1T instructions every 3 seconds feels amazing.

    Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.

    1. rithdmc · · focus · HN ↗
      For someone who has never touched workload optimization, can you share a little about how you measure this? Does IPC mean inter-process communication in this context?

      I know nothing about workload optimization, and I'd like to know more - right now I feel like the Good Burger gif. "Yeah, I know some of these words".

      1. j4k0bfr · · focus · HN ↗
        I've never done workload optimisation myself, but have looked into it just-in-case. It's a really interesting field and unfortunately, you need to know about both your compiler and target CPU/GPU to get the big gains. I believe most modern CPUs have registers for cache use. You have to do some guesswork to get IPC numbers, since core frequency can vary.

        Increasing instructions-per-clock is all about minimising program branching (essentially 'if' statements). Because a CPU core can execute instructions faster than main memory can fetch em. It's a fun game to look at an 'if' statement and figure out how you could instead make it an arithmetic operation :).

        Maximising cache hits is all about how you structure and access data. For example, if a cache entry is n bytes long, you want to ensure your struct is smaller than n bytes. Having very consistent access patterns can also help (e.g. arrays-of-structs vs structs-of-arrays).

        This field is super deep, it's very fun to learn about!

        1. rithdmc · · focus · HN ↗
          Thank you. Does workload optimisation usually refer to the CPU/GPU alone, or would it include disk or network IO? Or would that be more accurately called something else, like profiling?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.