‹ BackHN Continuity

Thread

How did AMD Ryzen get 50% faster in two years?

487 points · 206 comments · ibobev

  1. Yokolos · · focus · HN ↗
    I suspect the improvements are even more dramatic going from Zen 1 through to Zen 5. AMD has really hit the jackpot with how scalable the Ryzen CPU is considering how they're able to improve the performance from year to year. This is a stark difference to the FX series during the 2010s, which saw very small YoY performance increases by comparison. Ryzen really is AMD's equivalent to what Nehalem/Core was for Intel back in the mid 2000s.
    1. FlowingRiver · · focus · HN ↗
      That is all true but I will defend the FX series a little. Mostly now that there is a lot of software that scales across cores better now, they haven't aged as terribly as others have. They aren't great but not terrible considering.
      1. PorciiVorbesc · · focus · HN ↗
        >Mostly now that there is a lot of software that scales across cores better now

        That's pretty much irrelevant since the AMD's FX arch's issues weren't that SW at the time wasn't using all the 8 cores. Intel dropped the Core 2 Duo and Quad into the era where most SW was still stuck in single threaded for a long time and those CPUs still ripped single-threaded SW tasks regardless.

        Here's the big reasons why the FX sucked back then and why they still suck today in the multi-thread SW era:

          Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each module contained two integer execution units, but they had to share critical resources like one FPU, the instruction fetch/decode pipeline, and the L2 cache so when both "cores" inside a module were heavily taxed especially with math or physics-heavy calculations (like in videogames), they choked fighting over shared hardware.
        
          AMD designed Bulldozer with a very long pipeline, betting they could sacrifice efficiency per clock cycle in exchange for extraordinarily high clock speeds(a-la Intel Pentium 4) but the IPC was so bad that an FX core was often slower clock-for-clock than AMD’s previous-generation Phenom II chips and also their power consumption exploded. 
        
          FX processors were plagued by high cache latencies and an inefficient memory subsystem as another bottleneck.
        
        
        So unless you're into collecting vintage CPUs as display pieces, this one definitely belongs in the e-waste pile instead of burning electricity, because it did not age like wine with the adoption of SW multi threading like people were hoping.
        1. adfghopmnoi · · focus · HN ↗
          >Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each module contained two integer execution units, but they had to share critical resources like one FPU, the instruction fetch/decode pipeline, and the L2 cache so when both "cores" inside a module were heavily taxed especially with math or physics-heavy calculations (like in videogames), they choked fighting over shared hardware.

          They did just fine in parallel workloads, so I think this is not accurate. The design scaled just fine. The problem was that each core was weak.

          1. PorciiVorbesc · · focus · HN ↗
            >They did just fine in parallel workloads, so I think this is not accurate

            Depends how you define "doing just fine in parallel workloads". The contemporary competition from Intel that was 4-core + SMT was beating AMD's 8-core FX CPUs in most real-world tasks and benchmarks at the time. The 8-core AMD broke even and rarely won only in >4-thread strictly integer benchmarks and some >4-thread media encoding tasks/benchmarks. So if you wanted a prosumer media encoding workstation a budget then yeah, the AMD was better, but for most real world task, it really wasn't.

            >The design scaled just fine. The problem was that each core was weak.

            Can you elaborate and be more exact? What you wrote is technically vague and doesn't mean anything in technical dissection/terms.

            1. adfghopmnoi · · focus · HN ↗
              When we say a design "scales", that means that increasing the size of the workload does not incur a lot of overhead. If contention between shared resources meant that the design was not able to achieve an ~8x speedup when run with eight parallel threads, that would mean the design was not scalable. But we did in fact see a roughly 8x speedup with eight threads, so the design scaled just fine. The problem with the design was that each core was individually crummy, so even eight cores running in parallel had lackluster performance.

              The myth that each two-core module functioned more like one core with hyperthreading would suggest that these CPUs would have much higher per-core performance when lightly loaded than when fully loaded. That is not what happened. Each core was crummy even when lightly loaded, but under full load you would have eight crummy cores, which would beat four Intel cores on a lot of workloads.

              The only time contention was a serious problem was with workloads that were dominated by floating point, which were relatively rare.

              1. PorciiVorbesc · · focus · HN ↗
                > If contention between shared resources meant that the design was not able to achieve an ~8x speedup when run with eight parallel threads, that would mean the design was not scalable.

                By that definition it definitely was not scalable.

                >But we did in fact see a roughly 8x speedup with eight threads,

                Care to share a source? Because AFAIR there definitely was no 8x linear speedup with 8 threads even in benchmarks, let alone in real world use cases. The only benchmarks where those 8 threads would scale best and beat Intel were archival compression/decompression and media encoding. At everything else Intel wiped the floor with it.

                >The only time contention was a serious problem was with workloads that were dominated by floating point, which were relatively rare.

                Many real-world compute workloads, especially gaming related, are floating point.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.