‹ BackHN Continuity

Thread

How did AMD Ryzen get 50% faster in two years?

487 points · 206 comments · ibobev

  1. Yokolos · · focus · HN ↗
    I suspect the improvements are even more dramatic going from Zen 1 through to Zen 5. AMD has really hit the jackpot with how scalable the Ryzen CPU is considering how they're able to improve the performance from year to year. This is a stark difference to the FX series during the 2010s, which saw very small YoY performance increases by comparison. Ryzen really is AMD's equivalent to what Nehalem/Core was for Intel back in the mid 2000s.
    1. FlowingRiver · · focus · HN ↗
      That is all true but I will defend the FX series a little. Mostly now that there is a lot of software that scales across cores better now, they haven't aged as terribly as others have. They aren't great but not terrible considering.
      1. PorciiVorbesc · · focus · HN ↗
        >Mostly now that there is a lot of software that scales across cores better now

        That's pretty much irrelevant since the AMD's FX arch's issues weren't that SW at the time wasn't using all the 8 cores. Intel dropped the Core 2 Duo and Quad into the era where most SW was still stuck in single threaded for a long time and those CPUs still ripped single-threaded SW tasks regardless.

        Here's the big reasons why the FX sucked back then and why they still suck today in the multi-thread SW era:

          Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each module contained two integer execution units, but they had to share critical resources like one FPU, the instruction fetch/decode pipeline, and the L2 cache so when both "cores" inside a module were heavily taxed especially with math or physics-heavy calculations (like in videogames), they choked fighting over shared hardware.
        
          AMD designed Bulldozer with a very long pipeline, betting they could sacrifice efficiency per clock cycle in exchange for extraordinarily high clock speeds(a-la Intel Pentium 4) but the IPC was so bad that an FX core was often slower clock-for-clock than AMD’s previous-generation Phenom II chips and also their power consumption exploded. 
        
          FX processors were plagued by high cache latencies and an inefficient memory subsystem as another bottleneck.
        
        
        So unless you're into collecting vintage CPUs as display pieces, this one definitely belongs in the e-waste pile instead of burning electricity, because it did not age like wine with the adoption of SW multi threading like people were hoping.
        1. mort96 · · focus · HN ↗
          Hyperthreading/SMT is a significant boon for heavily threaded workloads. What makes that such a win while Bulldozer's implementation of "two integer units sharing a front-end, cache and FPU" is supposedly so bad? Because that description makes the 8 core Bulldozers sound exactly like a 4 core with SMT.
          1. PorciiVorbesc · · focus · HN ↗
            >Hyperthreading/SMT is a significant boon for heavily threaded workloads.

            That's hugely debatable and depends on SW workloads and the SMT implementation + CPU pipeline design.

            In SMT the execution engines, ALUs, FPUs, and caches are completely shared. When one thread stalls waiting for RAM, the second thread sneaks into the idle execution units. At best, SMT yields a ~10% to 20% throughput boost over a single thread.

            >Because that description makes the 8 core Bulldozers sound exactly like a 4 core with SMT.

            It's not the same thing. Bulldozer arch sits between a true 8-core and 4-core + SMT implementation.

            1. mort96 · · focus · HN ↗
              > At best, SMT yields a ~10% to 20% throughput boost over a single thread.

              Exactly, which is a significant benefit for how marginal the costs are.

              > It's not the same thing. Bulldozer arch sits between a true 8-core and 4-core + SMT implementation.

              Then surely it should be even better than 4 cores with SMT?

              If you're going to argue that the problem with Bulldozer was it's weird semi-SMT solution, you need to explain how it would've been better without it (aka as a regular quad core). Because even if it just gets the 10-20% performance improvements from being a form of SMT it would be better to have it than to not. And if you have lots of integer unit-bound threads, it should be even better than that.

              1. PorciiVorbesc · · focus · HN ↗
                >Then surely it should be even better than 4 cores with SMT? If you're going to argue that the problem with Bulldozer was it's weird semi-SMT solution, you need to explain how it would've been better without it

                I explained all the bottlenecks of the architecture in a comment above, that the issue was more than 4-core +SMT instead of true 8 cores. Please read it.

                1. mort96 · · focus · HN ↗
                  I did read it. It's not clear from it why you think having 2 integer units per core in a SMT-like configuration makes it worse. If you have <=4 threads it doesn't matter, just schedule the threads on different proper cores. If you have >4 memory/FPU/front-end heavy threads you should see the same benefit as SMT. If you have >4 integer arithmetic-bound cores, you should see a significant benefit beyond what SMT would give.

                  Now the very long pipeline and high memory latency are obviously significant issues with the architecture but those seem disconnected from the 4-core+SMT issue? I'm not questioning those issues at all, it's just not the part of your comment which interested me

                  EDIT: okay so in this comment: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49809017">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49809017, you explain that there&#x27;s actually a fairly large part of a core that&#x27;s duplicated, not just two integer units. If each &quot;core&quot; gets its own integer unit, register file and L1 cache, you&#x27;re actually paying a ton of die space for it, unlike SMT which is &quot;free&quot;. I can totally get how that can be a terrible trade-off for most workloads if it all ends up mostly starved due to front-end&#x2F;FPU&#x2F;memory throughput.

                  1. to11mtm · · focus · HN ↗
                    Well, &#x27;core&#x27; gets weird when we talk about Dozer. And where everything else had problems making it work.

                    AFAIR, a bulldozer &#x27;module&#x27; has what is exposed to a core as two CPUs, but, per everything above, is two integer cores, one shared FPU core, and depending on the version of the arch, possibly shared fetch&#x2F;decode&#x2F;other resources between all of that. Also AFAIR the decoder sucked as far as being able to feed both the integer cores, and the integer cores were more anemic compared to what was in, say, a K10H Phenom.

                    1. mort96 · · focus · HN ↗
                      Sorry, I edited my comment while you were writing. I had missed that it&#x27;s more than just &quot;one core with two ALUs&quot;. The more silicon you dedicate to this almost-but-not-quite-SMT solution, the worse of a trade-off it becomes in situations which don&#x27;t benefit from it, and it sounds like quite a lot of silicon was dedicated.
                  2. PorciiVorbesc · · focus · HN ↗
                    Your theory that &quot;1 shared FPU per module should equal 1 shared FPU per SMT core&quot; makes sense on paper, but Bulldozer lost to Intel’s Sandy Bridge 4C&#x2F;8T in floating-point and memory-heavy workloads because Intel&#x27;s individual FPU, cache hierarchy, and front-end pipelines were vastly wider and faster than Bulldozer&#x27;s shared components.

                    Having the same count of units (4 FPUs on the chip) did not mean having the same throughput. It&#x27;s a HW bottleneck, not something AMD could fix via the OS&#x27;s kernel allocation and scheduling of resources to the CPU to be able match Intel.

                    In strictly integer 4-8 thread benchmarks, yeah, AMD was often tied to Intel&#x27;s 4C+SMT.

                    Bulldozer’s design didn&#x27;t lose because the concept of sharing an FPU between two threads is worse than SMT. It lost because:

                      Intel&#x27;s FPU was natively twice as wide (256-bit vs. split 128-bit).
                    
                      AMD&#x27;s write-through L1 cache caused catastrophic write contention in L2.
                    
                      AMD&#x27;s L2 and L3 caches had double to triple the access latency of Intel&#x27;s.
                    
                      A single shared 4-wide decoder couldn&#x27;t feed an FPU and two integer units simultaneously.
                    1. mort96 · · focus · HN ↗
                      Wait but this new list of things you claim made Bulldozer lose is different from the previous list of things you claim made Bulldozer lose, and the &quot;two integer pipelines per core&quot; thing is no longer a part of it. You&#x27;re also responding argumentatively to my conciliatory comment.

                      This is exactly the kind of thing I&#x27;ve seen ChatGPT do way too much FWIW, arguments silently change after push-back without acknowledgement. I think you&#x27;re either a bot, or (more embarrassingly) using ChatGPT to formulate your arguments.

          2. wmf · · focus · HN ↗
            Bulldozer was twice the die size of Intel 4C&#x2F;8T but with lower performance.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.