‹ BackHN Continuity

Thread

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

562 points · 458 comments · giuliomagnifico

  1. fooker · · focus · HN ↗
    What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?
    1. nutjob2 · · focus · HN ↗
      Nothing except CPU manufacturer choices. Mac laptops use it and they're consumer products.

      People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.

      1. chessgecko · · focus · HN ↗
        pretty sure its lpddr not hbm.
      2. addaon · · focus · HN ↗
        > Mac laptops use it and they're consumer products

        No, Mac laptops use LPDDR, currently LPDDR5X.

      3. nomorewords · · focus · HN ↗
        The normal non-tech-savvy person already does this. They simply don't know that their ram is upgradeable or something else breaks first, before having to touch ram.
        1. dylan604 · · focus · HN ↗
          what is this too much RAM thing you mention? I thought the only valid RAM situation you could find yourself is not enough RAM. Too much? That's just fantasy
      4. KeplerBoy · · focus · HN ↗
        MacBooks use regular soldered lpddr5(x) RAM. Same RAM as every other laptop manufacturer, they just use more lanes to achieve a higher bandwidth.
        1. bunderbunder · · focus · HN ↗
          Perhaps more noteworthy for general home and business computing, doesn’t it also allow for lower latency?
          1. Rohansi · · focus · HN ↗
            More channels or soldered memory? Channels are basically RAID 0 so it depends what you're measuring. Soldering memory down was the only way to use LPDDR5X so if you wanted the best memory you had to solder it down. LPCAMM2 exists though so newer devices can use that instead of soldering them down, but not all devices would be able to fit the required LPCAMM2 slots.
            1. throwaway85825 · · focus · HN ↗
              SOCAMM2 allows for removable ram in nearly 0 added space.
              1. Rohansi · · focus · HN ↗
                That's tiny! But it still depends a lot on form factor. LPDDR5X is used in phones and making memory removable would have its compromises. You may even have compromises in laptops. Look at how tiny a MacBook Air's mainboard is and you'll see the RAM modules on the same package as the SoC. SOCAMM2 is too large for that but a variant with only two modules could possibly work.
                1. Melatonic · · focus · HN ↗
                  I bet it could fit. It just do a combo of soldered ram and non soldered like many laptops used to.
      5. fooker · · focus · HN ↗
        Apple's "unified memory" marketing is so strong that even tech literate people seem to have this misconception!

        They have managed to pull this sort of thing off many many times. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Reality_distortion_field" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Reality_distortion_field

        1. Rohansi · · focus · HN ↗
          Yup, the only reason Macs have higher memory bandwidth is because they use more memory channels, which gives them a wider bus. Both Intel and AMD only allow more than dual channel memory on server class processors these days.
          1. kstrauser · · focus · HN ↗
            The “Apple only does X better because they do Y” thing has been a meme for ages. I remember dismissals like “PowerPC is only faster at math because it has more integer units” or something along those lines, and thinking, uh, isn’t that a good thing?
            1. pixl97 · · focus · HN ↗
              Depends on the expense trade off.

              If I get 10% more performance for 50% more cost it really depends on one&#x27;s needs, for example.

            2. Rohansi · · focus · HN ↗
              I am just saying it&#x27;s not magic and x86 is capable of doing the same. Quad channel memory used to be more common in consumer hardware but now it looks like you don&#x27;t even have the option anymore for desktops. AMD&#x27;s Strix Halo was the first sign to reversing that (it has quad channel memory!) and hopefully we see more of that in the future.
              1. colejohnson66 · · focus · HN ↗
                AMD does market segmentation and limits consumer chips to dual-channel. You need to cough up the dough and get Threadripper for quad-channel. Or even more for Threadripper Pro to get octa-channel.
                1. Rohansi · · focus · HN ↗
                  As I mentioned above AMD&#x27;s Strix Halo has quad channel memory and is consumer level. But yes, other than that everything is segmented away.
                2. torginus · · focus · HN ↗
                  Afaik steam deck is quad channel, despite it being a pretty low end chip (with a decent GPU though)
                  1. Rohansi · · focus · HN ↗
                    Kind of but not really. DDR5 splits your typical 64-bit channel into two 32-bit subchannels meaning the bus width is not increased. These subchannels are not always advertised because it&#x27;s a just a feature of DDR5. Actually adding more channels increases bus width, which is what meaningfully improves memory bandwidth.
            3. pdpi · · focus · HN ↗
              There&#x27;s a qualitative difference between &quot;they&#x27;re doing a different thing&quot; and &quot;they&#x27;re doing the same thing, tuned differently&quot;. GP is saying that this is a case of &quot;they just tuned it differently&quot;.

              This distinction doesn&#x27;t change what the performance numbers look like today, but it does inform what changes would be necessary for those numbers to look different tomorrow. E.g. Apple Silicon isn&#x27;t fundamentally orders of magnitude more efficient than x86, they just used smaller features. Newer Intel and AMD chips made on equivalent processes _also_ get similar efficiency gains.

              1. sroussey · · focus · HN ↗
                There are AMD and Intel devices on similar process (not talking about A20Pro or M6 which are set to ship later this week), and they do not get the same gains.

                And honestly, they have historically had different markets.

                When the design is for only one customer, you don&#x27;t need to generalize things, and those things you generalize to give different customers different options has costs.

                AMD will soon be a larger customer for TSMC than Apple (NVIDIA is already there) so Apple&#x27;s pre-booking new processes is likely to be gone in the near future.

            4. JohnBooty · · focus · HN ↗
              &quot;You&#x27;re only better than me at sports because you practice more and try harder!&quot;
            5. Dylan16807 · · focus · HN ↗
              The point is that &quot;unified architecture&quot; is a buzzword unrelated to what actually makes things fast.
              1. Rohansi · · focus · HN ↗
                It does make some things faster because you can share memory between CPU, GPU, etc. without copying. But not everything.
          2. Kon5ole · · focus · HN ↗
            The memory bandwidth is a small thing compared to the massive win you get by not having to move data between two memory pools at all.
            1. Rohansi · · focus · HN ↗
              Depends on your workload. And AMD has supported unified memory long before Apple Silicon existed anyway.
            2. fooker · · focus · HN ↗
              Often you don&#x27;t move memory around as a programmer, but that&#x27;s exactly what happens in the background.

              It&#x27;s the address space that&#x27;s unified, not always the physical hardware.

              The data movement (when needed) is handled transparently in the background by page faults and other tricks.

              1. Rohansi · · focus · HN ↗
                Moving memory around is a bottleneck. Non-unified memory usually means copying over the PCIe bus which is way slower than RAM (and way way slower than VRAM). Actually unified memory means you don&#x27;t need to copy anything at all though which is the absolute best case for performance.
                1. [deleted] · · focus · HN ↗

                  [deleted]

          3. throwaway85825 · · focus · HN ↗
            Except strix halo.
          4. PunchyHamster · · focus · HN ↗
            Threadripper have that extra bandwidth and M5 still is faster
            1. Dylan16807 · · focus · HN ↗
              Can you link a specific benchmark?

              Keep in mind that non-Pro threadripper is still only 256 bits wide and Pro is 512. And the memory is 30% slower than with an M5. So an M5 Ultra has 3x the memory bandwidth of the best threadripper.

              1. Melatonic · · focus · HN ↗
                Don&#x27;t some AMD full server CPUs have 12 memory channels ?
                1. Dylan16807 · · focus · HN ↗
                  Yes, some do. So similar bandwidth to a Max but with lots of slower cores. Half as much as an Ultra.
        2. Kon5ole · · focus · HN ↗
          Having unified memory is a real advantage though, it&#x27;s not a reality distortion.
          1. davrosthedalek · · focus · HN ↗
            The price is that you essentially glue CPU and GPU together, which limits total compute, from a size and thermal perspective.

            This is really not a limit because of unified memory -- in principle, PCIe GPUs could read&#x2F;write main memory without the CPU. But it&#x27;s a limit for &#x2F;fast&#x2F; unified memory, because fast means close.

            So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages: a) probably faster transfer CPU&lt;-&gt;GPU (but that&#x27;s an implementation choice for the non-unified case b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.

          2. nvme0n1p1 · · focus · HN ↗
            Agreed. That&#x27;s why it&#x27;s a good thing all computers made in the past 15 years have unified memory, not just macs.

            <a href="https:&#x2F;&#x2F;x.com&#x2F;Lina_Hoshino&#x2F;status&#x2F;1820947147312820497" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;Lina_Hoshino&#x2F;status&#x2F;1820947147312820497

            1. stymaar · · focus · HN ↗
              Is that Marcan&#x27;s vtuber persona (Asahi Lina) that changed name since Marcan doesn&#x27;t work on Asahi Linux anymore?
            2. Kon5ole · · focus · HN ↗
              In theory perhaps but the benefit is not as relevant with a weak iGPU. In practice all PC&#x27;s with performance ambitions had a dGPU until Strix Halo and Panther Lake.
          3. fooker · · focus · HN ↗
            It is a real advantage.

            The reality distortion is that people seem to believe it&#x27;s HBM, or somehow it gives you extraordinary amounts of vram. Neither are really true.

          4. jorvi · · focus · HN ↗
            It isn&#x27;t really, as long as you don&#x27;t care about power consumption, physical constraints and money. Basically desktops &lt;2025 (and hopefully &gt;2027).

            DDR is optimized for latency and stability at the cost of bandwidth whilst GDDR is optimized for bandwidth at the cost of latency and stability. GDDR is pushed so hard these days that a small percentage of errors is expected and corrected because this is still faster than running it slower but more accurate.

            GDDR7 often has 10-20x the total bandwidth but 3x the latency of DDR5. Graphical workloads want as much bandwidth as possible but care relatively little for latency. Conversely, applications love low latency but don&#x27;t really see any performance benefit from higher bandwidth.

            So basicallyt you have workloads that are diametrically opposed and running unified memory forces you to compromise.

            1. Melatonic · · focus · HN ↗
              Is GDDR7 that&#x27;s ECC then just the super well binned stuff (plus the ECC parts added) ? I&#x27;ve been wondering why we don&#x27;t see more cards using it as they perform pretty damn well. Take the Nvidia 6000 Pro Blackwell for example. The compute is insanely fast assuming you can fit what you need in 96GB of ECC GDDR7
        3. nutjob2 · · focus · HN ↗
          No I just failed to check before I posted, I thought it was HBM. I don&#x27;t have any interest in Apple hardware so wasn&#x27;t properly informed, and have been duly crucified.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.