‹ BackHN Continuity

Thread

How did AMD Ryzen get 50% faster in two years?

487 points · 206 comments · ibobev

  1. Yokolos · · focus · HN ↗
    I suspect the improvements are even more dramatic going from Zen 1 through to Zen 5. AMD has really hit the jackpot with how scalable the Ryzen CPU is considering how they're able to improve the performance from year to year. This is a stark difference to the FX series during the 2010s, which saw very small YoY performance increases by comparison. Ryzen really is AMD's equivalent to what Nehalem/Core was for Intel back in the mid 2000s.
    1. FlowingRiver · · focus · HN ↗
      That is all true but I will defend the FX series a little. Mostly now that there is a lot of software that scales across cores better now, they haven't aged as terribly as others have. They aren't great but not terrible considering.
      1. PorciiVorbesc · · focus · HN ↗
        >Mostly now that there is a lot of software that scales across cores better now

        That's pretty much irrelevant since the AMD's FX arch's issues weren't that SW at the time wasn't using all the 8 cores. Intel dropped the Core 2 Duo and Quad into the era where most SW was still stuck in single threaded for a long time and those CPUs still ripped single-threaded SW tasks regardless.

        Here's the big reasons why the FX sucked back then and why they still suck today in the multi-thread SW era:

          Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each module contained two integer execution units, but they had to share critical resources like one FPU, the instruction fetch/decode pipeline, and the L2 cache so when both "cores" inside a module were heavily taxed especially with math or physics-heavy calculations (like in videogames), they choked fighting over shared hardware.
        
          AMD designed Bulldozer with a very long pipeline, betting they could sacrifice efficiency per clock cycle in exchange for extraordinarily high clock speeds(a-la Intel Pentium 4) but the IPC was so bad that an FX core was often slower clock-for-clock than AMD’s previous-generation Phenom II chips and also their power consumption exploded. 
        
          FX processors were plagued by high cache latencies and an inefficient memory subsystem as another bottleneck.
        
        
        So unless you're into collecting vintage CPUs as display pieces, this one definitely belongs in the e-waste pile instead of burning electricity, because it did not age like wine with the adoption of SW multi threading like people were hoping.
        1. throwawayffffas · · focus · HN ↗
          > Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each...

          Yeah that's hyper-threading intel was doing it as well and all modern CPUs do it as well. Where AMD dropped the ball, was they did not disclose that in their marketing as clearly as they should.

          All CPUs today are marketed as x cores 2x threads, back then some AMD marketing genius in their infinite wisdom put 8 cores on the box, instead of the honest 4 cores with hyperthreading.

          1. PorciiVorbesc · · focus · HN ↗
            No, AMD's FX "fake" 8-core was more than just 4-cores + hyperthreading. In SMT(hyperthreading) the execution engines, ALUs, FPUs, and caches are completely shared, whereas on FX design, they built two completely separate integer pipelines (schedulers, register files, ALUs, and L1 data caches) inside one module. Only the instruction fetch/decode front-end, the FPU, and the L2 cache were shared. So the FX design would be an in-between a 4-core + SMT and a true 8-core.
            1. throwawayffffas · · focus · HN ↗
              Fair, I had not delved into the details, but still they were not full cores and the marketing did not make a real distinction.

              I had a pilledriver one, it was a perfectly good cpu, I would buy it again. If I remember back then it was the best overall performance per dollar, the alternatives if I remember correctly were i7-39.. and i7-38.. and were at best 50% more expensive for 10-15% more performance.

              1. PorciiVorbesc · · focus · HN ↗
                Depends what you were doing with it. The Piledriver only beat the Intels in heavily multi threaded (preferably integer) workloads like media encoding, which is why it was popular with media creator workstations on a budget, but for most consumer real world tasks at the time, like video games, Intel was way ahead in performance even though it was more expensive.

                The Piledriver would win the consumer bang/buck mindset back then because of the 6-core part was reasonably priced and unlocked for overclocking, so people would overclock them to beat the more expensive (locked?) 4-c/8-t Intels at a lower price, but that ignored the costs of massive extra power draw(100+ W) over the Intel, the need for beefier more expensive coolers and power supplies, more expensive AMD motherboards with beefier MOSFET power delivery stages built to withstand the higher power draws of the Piledriver, so in the end the actual bang/buck gain of the AMD system wasn't remotely as big as people were making it out to be, they were just happy to get a "6-core" AMD cheaper than a 4-core Intel thinking more cores = more "better", same how having more mega-herz was also more "better" a decade before that.

                The Team AMD VS Team Intel wars on forums on these topics were wild back then.

                1. phire · · focus · HN ↗
                  > (preferably integer) workloads like media encoding

                  Media encoding is actually an FPU workload, and a pretty brutal one at that.

                  Media encoding might not use much floating point arithmetic, but it does use massive amounts of packed integer SIMD. And all SIMD instructions (both integer and floating) execute on the shared FPU, not the integer unit. It's only scalar integer instructions that execute on the integer unit.

                  Which leads me to believe that Bulldozer's shared FPU is not a bottleneck at all. Most evidence seems to point to the shared frontend being the primary bottleneck (which is why steamroller puts some effort into duplicating the instruction decoding, for some pretty large IPC wins)

          2. Zardoz84 · · focus · HN ↗
            No. it's just the opposite of Hyper-thereading. SMT it's about maximize the resource usage of a CPU, running 2 or more threads at the same time. AMD used the opposite technique of SMT. It had a technical name that I can't remember now, and wasn't invented by AMD.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.