‹ BackHN Continuity

Thread

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

562 points · 458 comments · giuliomagnifico

  1. fooker · · focus · HN ↗
    What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?
    1. refulgentis · · focus · HN ↗
      In one sense, nothing, in another, everything. It is DRAM, but the bandwidth requirements mean it’s paired to a processor, i.e. no DIMMs. Not 100% sure but things like MacBooks and the Framework tower, where you have fixed RAM for the device lifetime, have ~0 tradeoff.
    2. nutjob2 · · focus · HN ↗
      Nothing except CPU manufacturer choices. Mac laptops use it and they're consumer products.

      People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.

      1. chessgecko · · focus · HN ↗
        pretty sure its lpddr not hbm.
      2. addaon · · focus · HN ↗
        > Mac laptops use it and they're consumer products

        No, Mac laptops use LPDDR, currently LPDDR5X.

      3. nomorewords · · focus · HN ↗
        The normal non-tech-savvy person already does this. They simply don't know that their ram is upgradeable or something else breaks first, before having to touch ram.
        1. dylan604 · · focus · HN ↗
          what is this too much RAM thing you mention? I thought the only valid RAM situation you could find yourself is not enough RAM. Too much? That's just fantasy
      4. KeplerBoy · · focus · HN ↗
        MacBooks use regular soldered lpddr5(x) RAM. Same RAM as every other laptop manufacturer, they just use more lanes to achieve a higher bandwidth.
        1. bunderbunder · · focus · HN ↗
          Perhaps more noteworthy for general home and business computing, doesn’t it also allow for lower latency?
          1. Rohansi · · focus · HN ↗
            More channels or soldered memory? Channels are basically RAID 0 so it depends what you're measuring. Soldering memory down was the only way to use LPDDR5X so if you wanted the best memory you had to solder it down. LPCAMM2 exists though so newer devices can use that instead of soldering them down, but not all devices would be able to fit the required LPCAMM2 slots.
            1. throwaway85825 · · focus · HN ↗
              SOCAMM2 allows for removable ram in nearly 0 added space.
              1. Rohansi · · focus · HN ↗
                That's tiny! But it still depends a lot on form factor. LPDDR5X is used in phones and making memory removable would have its compromises. You may even have compromises in laptops. Look at how tiny a MacBook Air's mainboard is and you'll see the RAM modules on the same package as the SoC. SOCAMM2 is too large for that but a variant with only two modules could possibly work.
                1. Melatonic · · focus · HN ↗
                  I bet it could fit. It just do a combo of soldered ram and non soldered like many laptops used to.
      5. fooker · · focus · HN ↗
        Apple's "unified memory" marketing is so strong that even tech literate people seem to have this misconception!

        They have managed to pull this sort of thing off many many times. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Reality_distortion_field" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Reality_distortion_field

        1. Rohansi · · focus · HN ↗
          Yup, the only reason Macs have higher memory bandwidth is because they use more memory channels, which gives them a wider bus. Both Intel and AMD only allow more than dual channel memory on server class processors these days.
          1. kstrauser · · focus · HN ↗
            The “Apple only does X better because they do Y” thing has been a meme for ages. I remember dismissals like “PowerPC is only faster at math because it has more integer units” or something along those lines, and thinking, uh, isn’t that a good thing?
            1. pixl97 · · focus · HN ↗
              Depends on the expense trade off.

              If I get 10% more performance for 50% more cost it really depends on one&#x27;s needs, for example.

            2. Rohansi · · focus · HN ↗
              I am just saying it&#x27;s not magic and x86 is capable of doing the same. Quad channel memory used to be more common in consumer hardware but now it looks like you don&#x27;t even have the option anymore for desktops. AMD&#x27;s Strix Halo was the first sign to reversing that (it has quad channel memory!) and hopefully we see more of that in the future.
              1. colejohnson66 · · focus · HN ↗
                AMD does market segmentation and limits consumer chips to dual-channel. You need to cough up the dough and get Threadripper for quad-channel. Or even more for Threadripper Pro to get octa-channel.
                1. Rohansi · · focus · HN ↗
                  As I mentioned above AMD&#x27;s Strix Halo has quad channel memory and is consumer level. But yes, other than that everything is segmented away.
                2. torginus · · focus · HN ↗
                  Afaik steam deck is quad channel, despite it being a pretty low end chip (with a decent GPU though)
                  1. Rohansi · · focus · HN ↗
                    Kind of but not really. DDR5 splits your typical 64-bit channel into two 32-bit subchannels meaning the bus width is not increased. These subchannels are not always advertised because it&#x27;s a just a feature of DDR5. Actually adding more channels increases bus width, which is what meaningfully improves memory bandwidth.
            3. pdpi · · focus · HN ↗
              There&#x27;s a qualitative difference between &quot;they&#x27;re doing a different thing&quot; and &quot;they&#x27;re doing the same thing, tuned differently&quot;. GP is saying that this is a case of &quot;they just tuned it differently&quot;.

              This distinction doesn&#x27;t change what the performance numbers look like today, but it does inform what changes would be necessary for those numbers to look different tomorrow. E.g. Apple Silicon isn&#x27;t fundamentally orders of magnitude more efficient than x86, they just used smaller features. Newer Intel and AMD chips made on equivalent processes _also_ get similar efficiency gains.

              1. sroussey · · focus · HN ↗
                There are AMD and Intel devices on similar process (not talking about A20Pro or M6 which are set to ship later this week), and they do not get the same gains.

                And honestly, they have historically had different markets.

                When the design is for only one customer, you don&#x27;t need to generalize things, and those things you generalize to give different customers different options has costs.

                AMD will soon be a larger customer for TSMC than Apple (NVIDIA is already there) so Apple&#x27;s pre-booking new processes is likely to be gone in the near future.

            4. JohnBooty · · focus · HN ↗
              &quot;You&#x27;re only better than me at sports because you practice more and try harder!&quot;
            5. Dylan16807 · · focus · HN ↗
              The point is that &quot;unified architecture&quot; is a buzzword unrelated to what actually makes things fast.
              1. Rohansi · · focus · HN ↗
                It does make some things faster because you can share memory between CPU, GPU, etc. without copying. But not everything.
          2. Kon5ole · · focus · HN ↗
            The memory bandwidth is a small thing compared to the massive win you get by not having to move data between two memory pools at all.
            1. Rohansi · · focus · HN ↗
              Depends on your workload. And AMD has supported unified memory long before Apple Silicon existed anyway.
            2. fooker · · focus · HN ↗
              Often you don&#x27;t move memory around as a programmer, but that&#x27;s exactly what happens in the background.

              It&#x27;s the address space that&#x27;s unified, not always the physical hardware.

              The data movement (when needed) is handled transparently in the background by page faults and other tricks.

              1. Rohansi · · focus · HN ↗
                Moving memory around is a bottleneck. Non-unified memory usually means copying over the PCIe bus which is way slower than RAM (and way way slower than VRAM). Actually unified memory means you don&#x27;t need to copy anything at all though which is the absolute best case for performance.
                1. [deleted] · · focus · HN ↗

                  [deleted]

          3. throwaway85825 · · focus · HN ↗
            Except strix halo.
          4. PunchyHamster · · focus · HN ↗
            Threadripper have that extra bandwidth and M5 still is faster
            1. Dylan16807 · · focus · HN ↗
              Can you link a specific benchmark?

              Keep in mind that non-Pro threadripper is still only 256 bits wide and Pro is 512. And the memory is 30% slower than with an M5. So an M5 Ultra has 3x the memory bandwidth of the best threadripper.

              1. Melatonic · · focus · HN ↗
                Don&#x27;t some AMD full server CPUs have 12 memory channels ?
                1. Dylan16807 · · focus · HN ↗
                  Yes, some do. So similar bandwidth to a Max but with lots of slower cores. Half as much as an Ultra.
        2. Kon5ole · · focus · HN ↗
          Having unified memory is a real advantage though, it&#x27;s not a reality distortion.
          1. davrosthedalek · · focus · HN ↗
            The price is that you essentially glue CPU and GPU together, which limits total compute, from a size and thermal perspective.

            This is really not a limit because of unified memory -- in principle, PCIe GPUs could read&#x2F;write main memory without the CPU. But it&#x27;s a limit for &#x2F;fast&#x2F; unified memory, because fast means close.

            So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages: a) probably faster transfer CPU&lt;-&gt;GPU (but that&#x27;s an implementation choice for the non-unified case b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.

          2. nvme0n1p1 · · focus · HN ↗
            Agreed. That&#x27;s why it&#x27;s a good thing all computers made in the past 15 years have unified memory, not just macs.

            <a href="https:&#x2F;&#x2F;x.com&#x2F;Lina_Hoshino&#x2F;status&#x2F;1820947147312820497" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;Lina_Hoshino&#x2F;status&#x2F;1820947147312820497

            1. stymaar · · focus · HN ↗
              Is that Marcan&#x27;s vtuber persona (Asahi Lina) that changed name since Marcan doesn&#x27;t work on Asahi Linux anymore?
            2. Kon5ole · · focus · HN ↗
              In theory perhaps but the benefit is not as relevant with a weak iGPU. In practice all PC&#x27;s with performance ambitions had a dGPU until Strix Halo and Panther Lake.
          3. fooker · · focus · HN ↗
            It is a real advantage.

            The reality distortion is that people seem to believe it&#x27;s HBM, or somehow it gives you extraordinary amounts of vram. Neither are really true.

          4. jorvi · · focus · HN ↗
            It isn&#x27;t really, as long as you don&#x27;t care about power consumption, physical constraints and money. Basically desktops &lt;2025 (and hopefully &gt;2027).

            DDR is optimized for latency and stability at the cost of bandwidth whilst GDDR is optimized for bandwidth at the cost of latency and stability. GDDR is pushed so hard these days that a small percentage of errors is expected and corrected because this is still faster than running it slower but more accurate.

            GDDR7 often has 10-20x the total bandwidth but 3x the latency of DDR5. Graphical workloads want as much bandwidth as possible but care relatively little for latency. Conversely, applications love low latency but don&#x27;t really see any performance benefit from higher bandwidth.

            So basicallyt you have workloads that are diametrically opposed and running unified memory forces you to compromise.

            1. Melatonic · · focus · HN ↗
              Is GDDR7 that&#x27;s ECC then just the super well binned stuff (plus the ECC parts added) ? I&#x27;ve been wondering why we don&#x27;t see more cards using it as they perform pretty damn well. Take the Nvidia 6000 Pro Blackwell for example. The compute is insanely fast assuming you can fit what you need in 96GB of ECC GDDR7
        3. nutjob2 · · focus · HN ↗
          No I just failed to check before I posted, I thought it was HBM. I don&#x27;t have any interest in Apple hardware so wasn&#x27;t properly informed, and have been duly crucified.
    3. bob1029 · · focus · HN ↗
      It&#x27;s not that there&#x27;s a blocker. It&#x27;s that it takes roughly 3x the manufacturing capacity to produce an HBM package at the same storage capacity as DRAM. We are sacrificing total bytes for bandwidth.
      1. tliltocatl · · focus · HN ↗
        How so? FEOL is pretty much the same, BEOL is almost the same save TSVs, the packaging tech is different and more advanced, but not exactly 1:1 comparable. Do TSVs really occupy 3x the area of DDR IO&#x27;s?
        1. monster_truck · · focus · HN ↗
          3x is a reasonable figure. They are not literally that large, though.
        2. Const-me · · focus · HN ↗
          See remark on the slide 11: <a href="https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architectures-for-ai-at-hot-chips-2026&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architec... That presentation is by Micron.
          1. skavi · · focus · HN ↗
            direct link: <a href="https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architectures-for-ai-slide-11&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architec...
            1. tliltocatl · · focus · HN ↗
              Yuck. Time to build an xSPI&#x2F;HyperRAM workstation (if only these had multi-bank chips).
              1. buildbot · · focus · HN ↗
                Sadly the $ per byte of xSPI and HyperRAM quite high
                1. tliltocatl · · focus · HN ↗
                  Yes, but enough to run a text editor, or even a mechanical CAD. Not an LLM, but that&#x27;s the point!
          2. Karliss · · focus · HN ↗
            It doesn&#x27;t really say that it needs to use 3x more area, but that 3x more gets consumed due to &quot;advanced packaging and manfuacturing complexity&quot;. Which doesn&#x27;t properly explain why it consumes 3x more and could simply mean they have a bad yield and 2&#x2F;3 produced is garbage.
            1. tliltocatl · · focus · HN ↗
              Yea, that&#x27;s the question. Yield situation can improve. Area overhead would not improve short of a completely new and incompatible tech.
            2. bob1029 · · focus · HN ↗
              &gt; 2&#x2F;3 produced is garbage.

              This might not be far off the mark. You are irreversibly linking the fates of these devices after a certain stage of manufacturing. If something goes wrong at final packaging time, you lose all dies instead of one.

            3. crote · · focus · HN ↗
              You make HBM by stacking a whole bunch of dies on top of each other. The signals from the upper dies need to pass through vias in the lower dies to get out - taking up valuable die space in a way which simply isn&#x27;t needed with regular DDR. Similarly, HBM has a far wider bus, so each individual die has, say, 16 banks of depth 32, rather than 4 banks of depth 128. That&#x27;s more control area needed per byte of memory.

              Those two combined already result in a huge reduction in bytes per mm2, so with the same wafer processing capacity you&#x27;re producing far less byte of memory. Add to that a complicated chain of HBM-specific packaging steps, and you&#x27;re now also losing a decent bunch of perfectly-fine dies because rather than putting it into DDR you tried making a HBM sandwich and screwed up.

              Even if the memory cells are the same and have an absolutely identical yield, HBM will always end up having a significantly lower output. That&#x27;s just the cost of stacking, but some people are willing to pay the per-gigabyte price penalty in return for the higher bandwidth.

            4. imtringued · · focus · HN ↗
              Classic DRAM stacks up to four wafers on top of each other and then is packaged with BGAs. The manufacturer can check the DRAM chips independently.

              Soldering the DRAM onto a PCB is such a reliable process that there is almost zero risk of defects and even if a defect occurs the damage is limited. If the DRAM is soldered onto a DIMM the risk of a defect on the non memory hardware is non-existent. If the DRAM is soldered straight onto an SBC or GPU, then the DRAM can be removed to save the precious SoC or GPU chips.

              Meanwhile HBM is the ultimate nightmare scenario. You stack up to 16 DRAM wafers on top of each other. One defect and the whole stack is worthless and that was actually the easy part.

              In stage two things get even worse. You now have your accelerator chip and you must place the HBM on that chip. E.g. Blackwell GB300 has eight HBM stacks and the accelerator chip has a bigger area than the HBM. You must get the packaging right eight times in a row or you have wasted not only the DRAM silicon, but also the accelerator silicon because HBM cannot be removed and defects are permanent.

              The issue here isn&#x27;t just the yield of the HBM (which is obviously lower if you have taller stacks) but rather the yield of the combined HBM-based product, which is why doesn&#x27;t make sense to say it needs more area but it is completely correct to say that HBM leads to more silicon being consumed. Hence it doesn&#x27;t make sense to talk about yield of the HBM itself, because it is always part of an integrated product.

      2. rkagerer · · focus · HN ↗
        I don&#x27;t fully understand the source of the &quot;total bytes&quot; constraint, but a major factor may be because HBM4 &#x2F; HBM4E can only make use of the footprint directly above the processor&#x2F;logic die (or in direct vicinity of its interconnect), while traditional DRAM can be placed further away where there&#x27;s lots of real estate on the motherboard.

        I gather a practical max ceiling today is a stack of 16 chips in height yielding 64GB?

        These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5. That&#x27;s what gives them their order-of-magnitude bandwidth speedup. But even though they technically pack in more capacity per square millimeter of motherboard, I gather they take up more space than older technologies once you account for the vias and interconnects to route all those signals.

        1. threecheese · · focus · HN ↗
          Thanks for that, just went down an interesting rabbit hole. Many of us were hoping this re-tooling would eventually trickle some fast RAM down to DRAM-exhausted PCs, but given it would require a rearchitecture of the motherboard it&#x27;s unlikely.
          1. sroussey · · focus · HN ↗
            HBM also trades bandwidth for latency, and your regular computing is much more sensitive to latency than bandwidth.
          2. craigjb · · focus · HN ↗
            HBM4 has over 2048 signals to the processor’s PHY with tight signal integrity requirements that require the HBM stack to be &lt; 0.5 mm from the processor die. That’s why HBM integration is done with interposers (soldered on the package). So, it’d be the CPU package that integrates it. Motherboard is too far away.
            1. Melatonic · · focus · HN ↗
              Kind of seems like we should be making chips with both. Big HBM stack on top as a sort of huge L5 cache like thing. And then a bunch of DRAM type sockets (like LPCAMM) around the exterior.
              1. craigjb · · focus · HN ↗
                For chips with integrated CPU+GPU+NPU, it could be worth it tech-wise. The GPU and NPU can eat HBM bandwidth. For general purpose CPU code, the HBM would likely not be worth it. It&#x27;s high bandwidth, but you trade latency, and general CPU code is branchy. Economics-wise, the HBM stacks alone will cost more than a consumer CPU (or APU).

                [edit] The packaging needed to support HBM is also much more expensive too. If demand for current HBM applications tanks and the manufacturing lines need filled, then maybe. Currently, the price point would make it very very niche.

    4. chessgecko · · focus · HN ↗
      I think people might prefer the lower idle power consumption from lpddr over the better bandwidth in hbm in battery powered stuff. That said right now the price is definitely preventing us from finding out.
      1. fooker · · focus · HN ↗
        Idle yes, but HBM energy consumption &#x2F; memory operations seems to be a bit better than DRAM.
        1. vlovich123 · · focus · HN ↗
          Only if you’re running at 100%. Consumers generally do not.
          1. monster_truck · · focus · HN ↗
            That hasn&#x27;t been true since early HBM2 days, before the controllers standardized on power&#x2F;voltage management and did things like leave them in P0 to ship on time
            1. vlovich123 · · focus · HN ↗
              LPDDR&#x2F;DDR&#x2F;GDDR generally still win over HBM when there’s no data being transferred. HBM is primarily better in watts&#x2F;byte transferred. Consumer electronics spend most of their time idle.
              1. Zagitta · · focus · HN ↗
                Racing to idle is a very common power optimization technique
                1. vlovich123 · · focus · HN ↗
                  Right, but HBM idle is significantly worse than LPDDR idle or even DDR idle for that matter. That matters a lot precisely because the device is idle most of the time. Your idle power draw dominates.
                  1. monster_truck · · focus · HN ↗
                    Comparing HBM to LPDDR is pretty silly. That&#x27;s like a bath tub vs a urinal
                    1. imtringued · · focus · HN ↗
                      If LPDDR6 PIM ever becomes a thing then the advantage of HBM will shrink. I&#x27;m not saying LPDDR6 PIM will make HBM obsolete in the datacenter or where it is currently shining, I&#x27;m saying that large volumes have their own charm and it is more likely for LPDDR6 PIM to be in your laptop or smartphone or SBC than HBM.

                      LPDDR6 PIM would primarily help the low end and mid range accelerator market. E.g. embedded models running on SBCs can be up to 1 GiB in size with acceptable performance, small models at 8 GiB become really easy to run at reasonable speeds on a smartphone and PIM enabled laptops or mini PCs make it possible to run 32 GiB models locally without compromise.

                      Of course this also assumes that the associated accelerators (NPUs) will catch up too, but the general point is that you won&#x27;t need a 5090 or a 4090 anymore.

                    2. vlovich123 · · focus · HN ↗
                      This is literally a thread where people are saying “but why laptops no HBM”. It’s a perfectly cromulent response that meets the question where it is. It might be mildly a more defensible question for normal plugged in desktops but that would require a massive ecosystem change since people who build those traditionally really like their DIMMs and less buying a CPU that has non-swappable RAM built in, not to mention there’s no sustainable market for the higher cost and low volume product.
                      1. monster_truck · · focus · HN ↗
                        No, it really isn&#x27;t. You&#x27;re barking up the wrong trees for incorrect reasons
    5. phkahler · · focus · HN ↗
      &gt;&gt; What&#x27;s the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?

      HBM is meant to be integrated into the same package as the CPU, so no more DIMM sockets. It also has higher latency apparently.

      1. MarleTangible · · focus · HN ↗
        One of the comments mentioned that they have a bus size of 2048, which may be why the latency is higher.

        &gt; These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5.

    6. reliabilityguy · · focus · HN ↗
      HBM is a stack of DRAMs, so there is no “instead”.
      1. buckle8017 · · focus · HN ↗
        The vias to enable stacking is a significant amount of the die area.
        1. saltcured · · focus · HN ↗
          If they&#x27;re talking about production capacity, that is some product of die area and process steps, right? It doesn&#x27;t have to be 3x die area, just 3x lower factory throughput for the same number of functioning memory bits.
          1. buckle8017 · · focus · HN ↗
            HBM is less dense at a water scale than DDR because of all the vias, but each memory but is as dense or denser.

            You make HBM instead of DDR and the number of bits you&#x27;re making goes down.

            It&#x27;s really that simple.

    7. HarHarVeryFunny · · focus · HN ↗
      Why would you want&#x2F;need to?

      The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU.

      So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second!

      These are not normal requirements, other than for a GPU.

      1. fooker · · focus · HN ↗
        SIMD (and especially the modern matrix extensions) can use as much bandwidth you can throw at it.
        1. b112 · · focus · HN ↗
          Without further clarification, that statement seems impossible. &quot;As much&quot; being unbounded and all. You should expand what you mean.
          1. fooker · · focus · HN ↗
            This will blow your mind, but it actually is pretty close to being unbounded. :)

            Consider the &#x27;MMA N matrices&#x27; primitive modern CPUs are starting to support. For the current generation of CPUs, N is a constant like 16 or 32, but there&#x27;s nothing preventing it from being 1024 or larger if we have more memory bandwidth.

            All this with a single instruction.

            1. articulatepang · · focus · HN ↗
              Surely something prevents it being 1 quadrillion bits per instruction? Since that’s well within “unbounded”.
              1. pixl97 · · focus · HN ↗
                Most likely the chip running at the core temperature of the sun.

                We&#x27;ll have to figure out how to read and right to the surface of a black hole to get speeds that high.

            2. imtringued · · focus · HN ↗
              Nah his point is a bit simpler. Vector units can process a limited amount of data per time unit. Memory can load a limited amount of data per time unit.

              If you have infinite memory bandwidth you just move the bottleneck back to compute so both have to grow simultaneously in lockstep.

              What you should have said is that CPUs have so much compute headroom for matrix vector multiplication that simply adding more memory bandwidth would make them faster so every improvement in memory bandwidth is welcome.

              1. fooker · · focus · HN ↗
                Agreed.

                The &quot;move the bottleneck back to compute&quot; bit is changing rapidly though. The first time a major hardware company ships a PIM chip, you can push for a few orders of magnitude more data through without being compute bound.

      2. XorNot · · focus · HN ↗
        Right but if HBM memory is all that people want to produce, then building a CPU which can use it use it would be useful on it&#x27;s own merits.

        But in reality we also already have unified memory architecture systems, integrated graphics etc.

        1. p1esk · · focus · HN ↗
          People want to produce hbm because it’s more expensive and more profitable than regular memory.
          1. XorNot · · focus · HN ↗
            There is a world where scale and experience means it&#x27;s about the same though, is the thing.

            And memory is already expensive. It&#x27;s downright hard to even get it though - you frequently would prefer not what&#x27;s cheapest, but whatever is in largest scale production.

            1. Dylan16807 · · focus · HN ↗
              Scale and experience almost entirely share between HBM and normal memory. And they&#x27;re both in large-enough scale production to not have a big difference on availability; if you&#x27;re willing to pay HBM prices you should find even more sellers of DDR.

              The only way I see HBM becoming competitive for consumer CPUs is if they solve the yield issues. Or if AI crashes so hard that people are putting those GPUs on fire sale and salvaging mass quantities of HBM off of them.

              1. zeristor · · focus · HN ↗
                I was thinking this, but I doubt that the HBM is so easy to repurpose.

                I’m assuming that it’s mounted in the same unit as the GPU, not in a handy-dandy socket.

                It could be cracked open and extracted no doubt but if it’s glued in that’s going to be nigh on impossible to extract.

                I had been pinning my hopes on HBM coming onto the second hand markets after there three years or so of use, perhaps I’m wrong.

                1. pixl97 · · focus · HN ↗
                  It seems unlikely as the entire machines with HBM on them are apt to be sold whole on the market and snapped up quickly. With how much demand is in the current market machines 3 years old may not be sold if they can&#x27;t get newer faster machines fast enough.
      3. Dylan16807 · · focus · HN ↗
        &gt; Compare that to the 128 bit wide bus of a modern CPU.

        Or 256-512 bits on medium to high end consumer CPUs if you&#x27;re apple.

        At least DDR6 is probably widening things 50%.

      4. bobmcnamara · · focus · HN ↗
        Intel and IBM have already done this almost this with their wide eDRAM caches.

        The advantage in going wide is transferring cache lines rapidly, not the CPU bus interface.

    8. torginus · · focus · HN ↗
      Mainly bus width. Afaik HBM is like 1024 bits vs DDRs 64 so you need lots of transfers in parallel to saturate the bus, and CPUs kinda want 64 bytes of data as that&#x27;s the size of a cache line ASAP. So you need a ton of in flight transfers which isn&#x27;t a thing CPUs provide, maybe multicore workloads.

      Buy the way you win with CPUs is with latency, and not bandwidth, which is why Apple M series actually uses DDR with lower latency because of the stacking.

    9. tjwebbnorfolk · · focus · HN ↗
      Even if the cost were the same of the RAM itself, you&#x27;d need much bigger and more expensive CPU to deal with it.

      Running 17 chrome tabs doesn&#x27;t benefit at all from that HBM and all the additional hardware+software complexities that come with it. You want a specialized coprocessor to handle specialized workloads. The GPU exists separately from the CPU for a reason.

    10. KurSix · · focus · HN ↗
      HBM isn&#x27;t dramatically better in every dimension. You get huge bandwidth and good energy efficiency per bit transferred, but not necessarily a meaningful latency improvement and capacity expansion becomes tied to the package
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.