‹ BackHN Continuity

Thread

How did AMD Ryzen get 50% faster in two years?

487 points · 206 comments · ibobev

  1. bmenrigh · · focus · HN ↗
    I was recently able to get a workload of mine optimized enough on Zen 5 to hit a sustained 6.0 IPC/core (3.0 / thread) at 5.1 GHz. Seeing the a > 99.8% branch prediction rate and a > 99.99% L2 cache hit rate retiring > 1T instructions every 3 seconds feels amazing.

    Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.

    1. stinkbeetle · · focus · HN ↗
      Can you share any more details about the workload? Always interesting to hear of something like that which isn't a useless microbenchmark.

      Does it use vector? What can you hit with SMT disabled?

      1. bmenrigh · · focus · HN ↗
        It’s a backtracking search program looking for integer solutions to a specific problem. I’ve tuned a series of bloom filters to fill my L2 cache so that I rarely have to touch main memory (this alone took my IPC from 0.1-0.3 to 3.0 per thread). Without SMT it’s 4.6 IPC/core.

        I think it’s only able to exceed 4.0/thread with SMT off because of a uops cache? From what I’ve read the Zen5 front end only had a 4-wide instruction decode per thread.

        1. explodingwaffle · · focus · HN ↗
          coded entirely or partially in assembly, i presume?
          1. bmenrigh · · focus · HN ↗
            Entirely in C. Compilers are pretty good these days. Almost every time I try to outsmart them I make things worse.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.