‹ BackHN Continuity

Thread

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

562 points · 458 comments · giuliomagnifico

  1. fooker · · focus · HN ↗
    What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?
    1. nutjob2 · · focus · HN ↗
      Nothing except CPU manufacturer choices. Mac laptops use it and they're consumer products.

      People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.

      1. fooker · · focus · HN ↗
        Apple's "unified memory" marketing is so strong that even tech literate people seem to have this misconception!

        They have managed to pull this sort of thing off many many times. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Reality_distortion_field" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Reality_distortion_field

        1. Rohansi · · focus · HN ↗
          Yup, the only reason Macs have higher memory bandwidth is because they use more memory channels, which gives them a wider bus. Both Intel and AMD only allow more than dual channel memory on server class processors these days.
          1. kstrauser · · focus · HN ↗
            The “Apple only does X better because they do Y” thing has been a meme for ages. I remember dismissals like “PowerPC is only faster at math because it has more integer units” or something along those lines, and thinking, uh, isn’t that a good thing?
            1. pixl97 · · focus · HN ↗
              Depends on the expense trade off.

              If I get 10% more performance for 50% more cost it really depends on one&#x27;s needs, for example.

            2. Rohansi · · focus · HN ↗
              I am just saying it&#x27;s not magic and x86 is capable of doing the same. Quad channel memory used to be more common in consumer hardware but now it looks like you don&#x27;t even have the option anymore for desktops. AMD&#x27;s Strix Halo was the first sign to reversing that (it has quad channel memory!) and hopefully we see more of that in the future.
              1. colejohnson66 · · focus · HN ↗
                AMD does market segmentation and limits consumer chips to dual-channel. You need to cough up the dough and get Threadripper for quad-channel. Or even more for Threadripper Pro to get octa-channel.
                1. Rohansi · · focus · HN ↗
                  As I mentioned above AMD&#x27;s Strix Halo has quad channel memory and is consumer level. But yes, other than that everything is segmented away.
                2. torginus · · focus · HN ↗
                  Afaik steam deck is quad channel, despite it being a pretty low end chip (with a decent GPU though)
                  1. Rohansi · · focus · HN ↗
                    Kind of but not really. DDR5 splits your typical 64-bit channel into two 32-bit subchannels meaning the bus width is not increased. These subchannels are not always advertised because it&#x27;s a just a feature of DDR5. Actually adding more channels increases bus width, which is what meaningfully improves memory bandwidth.
            3. pdpi · · focus · HN ↗
              There&#x27;s a qualitative difference between &quot;they&#x27;re doing a different thing&quot; and &quot;they&#x27;re doing the same thing, tuned differently&quot;. GP is saying that this is a case of &quot;they just tuned it differently&quot;.

              This distinction doesn&#x27;t change what the performance numbers look like today, but it does inform what changes would be necessary for those numbers to look different tomorrow. E.g. Apple Silicon isn&#x27;t fundamentally orders of magnitude more efficient than x86, they just used smaller features. Newer Intel and AMD chips made on equivalent processes _also_ get similar efficiency gains.

              1. sroussey · · focus · HN ↗
                There are AMD and Intel devices on similar process (not talking about A20Pro or M6 which are set to ship later this week), and they do not get the same gains.

                And honestly, they have historically had different markets.

                When the design is for only one customer, you don&#x27;t need to generalize things, and those things you generalize to give different customers different options has costs.

                AMD will soon be a larger customer for TSMC than Apple (NVIDIA is already there) so Apple&#x27;s pre-booking new processes is likely to be gone in the near future.

            4. JohnBooty · · focus · HN ↗
              &quot;You&#x27;re only better than me at sports because you practice more and try harder!&quot;
            5. Dylan16807 · · focus · HN ↗
              The point is that &quot;unified architecture&quot; is a buzzword unrelated to what actually makes things fast.
              1. Rohansi · · focus · HN ↗
                It does make some things faster because you can share memory between CPU, GPU, etc. without copying. But not everything.
          2. Kon5ole · · focus · HN ↗
            The memory bandwidth is a small thing compared to the massive win you get by not having to move data between two memory pools at all.
            1. Rohansi · · focus · HN ↗
              Depends on your workload. And AMD has supported unified memory long before Apple Silicon existed anyway.
            2. fooker · · focus · HN ↗
              Often you don&#x27;t move memory around as a programmer, but that&#x27;s exactly what happens in the background.

              It&#x27;s the address space that&#x27;s unified, not always the physical hardware.

              The data movement (when needed) is handled transparently in the background by page faults and other tricks.

              1. Rohansi · · focus · HN ↗
                Moving memory around is a bottleneck. Non-unified memory usually means copying over the PCIe bus which is way slower than RAM (and way way slower than VRAM). Actually unified memory means you don&#x27;t need to copy anything at all though which is the absolute best case for performance.
                1. [deleted] · · focus · HN ↗

                  [deleted]

          3. throwaway85825 · · focus · HN ↗
            Except strix halo.
          4. PunchyHamster · · focus · HN ↗
            Threadripper have that extra bandwidth and M5 still is faster
            1. Dylan16807 · · focus · HN ↗
              Can you link a specific benchmark?

              Keep in mind that non-Pro threadripper is still only 256 bits wide and Pro is 512. And the memory is 30% slower than with an M5. So an M5 Ultra has 3x the memory bandwidth of the best threadripper.

              1. Melatonic · · focus · HN ↗
                Don&#x27;t some AMD full server CPUs have 12 memory channels ?
                1. Dylan16807 · · focus · HN ↗
                  Yes, some do. So similar bandwidth to a Max but with lots of slower cores. Half as much as an Ultra.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.