‹ BackHN Continuity

Thread

How did AMD Ryzen get 50% faster in two years?

487 points · 206 comments · ibobev

  1. bmenrigh · · focus · HN ↗
    I was recently able to get a workload of mine optimized enough on Zen 5 to hit a sustained 6.0 IPC/core (3.0 / thread) at 5.1 GHz. Seeing the a > 99.8% branch prediction rate and a > 99.99% L2 cache hit rate retiring > 1T instructions every 3 seconds feels amazing.

    Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.

    1. rithdmc · · focus · HN ↗
      For someone who has never touched workload optimization, can you share a little about how you measure this? Does IPC mean inter-process communication in this context?

      I know nothing about workload optimization, and I'd like to know more - right now I feel like the Good Burger gif. "Yeah, I know some of these words".

      1. j4k0bfr · · focus · HN ↗
        I've never done workload optimisation myself, but have looked into it just-in-case. It's a really interesting field and unfortunately, you need to know about both your compiler and target CPU/GPU to get the big gains. I believe most modern CPUs have registers for cache use. You have to do some guesswork to get IPC numbers, since core frequency can vary.

        Increasing instructions-per-clock is all about minimising program branching (essentially 'if' statements). Because a CPU core can execute instructions faster than main memory can fetch em. It's a fun game to look at an 'if' statement and figure out how you could instead make it an arithmetic operation :).

        Maximising cache hits is all about how you structure and access data. For example, if a cache entry is n bytes long, you want to ensure your struct is smaller than n bytes. Having very consistent access patterns can also help (e.g. arrays-of-structs vs structs-of-arrays).

        This field is super deep, it's very fun to learn about!

        1. bluGill · · focus · HN ↗
          Though you should always ask what your compiler/optimizer is doing for you. Research into how to turn your ifs into arithmetic is ongoing. If the compiler can optimize your simple ifs into the complex arithmetic then you should go with the simple if and let the compiler do that. (unfortunately in many cases the compiler cannot be sure because of some special case - even though odds are you don't care about it)
          1. j4k0bfr · · focus · HN ↗
            Agreed! Although I like the idea of considering each logical branch added to my programs. It kind of forces me to think through how variance is being treated in my call stack.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.