‹ BackHN Continuity

Thread

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

562 points · 458 comments · giuliomagnifico

  1. fooker · · focus · HN ↗
    What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?
    1. bob1029 · · focus · HN ↗
      It's not that there's a blocker. It's that it takes roughly 3x the manufacturing capacity to produce an HBM package at the same storage capacity as DRAM. We are sacrificing total bytes for bandwidth.
      1. tliltocatl · · focus · HN ↗
        How so? FEOL is pretty much the same, BEOL is almost the same save TSVs, the packaging tech is different and more advanced, but not exactly 1:1 comparable. Do TSVs really occupy 3x the area of DDR IO's?
        1. monster_truck · · focus · HN ↗
          3x is a reasonable figure. They are not literally that large, though.
        2. Const-me · · focus · HN ↗
          See remark on the slide 11: <a href="https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architectures-for-ai-at-hot-chips-2026&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architec... That presentation is by Micron.
          1. skavi · · focus · HN ↗
            direct link: <a href="https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architectures-for-ai-slide-11&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.servethehome.com&#x2F;micron-evolving-memory-architec...
            1. tliltocatl · · focus · HN ↗
              Yuck. Time to build an xSPI&#x2F;HyperRAM workstation (if only these had multi-bank chips).
              1. buildbot · · focus · HN ↗
                Sadly the $ per byte of xSPI and HyperRAM quite high
                1. tliltocatl · · focus · HN ↗
                  Yes, but enough to run a text editor, or even a mechanical CAD. Not an LLM, but that&#x27;s the point!
          2. Karliss · · focus · HN ↗
            It doesn&#x27;t really say that it needs to use 3x more area, but that 3x more gets consumed due to &quot;advanced packaging and manfuacturing complexity&quot;. Which doesn&#x27;t properly explain why it consumes 3x more and could simply mean they have a bad yield and 2&#x2F;3 produced is garbage.
            1. tliltocatl · · focus · HN ↗
              Yea, that&#x27;s the question. Yield situation can improve. Area overhead would not improve short of a completely new and incompatible tech.
            2. bob1029 · · focus · HN ↗
              &gt; 2&#x2F;3 produced is garbage.

              This might not be far off the mark. You are irreversibly linking the fates of these devices after a certain stage of manufacturing. If something goes wrong at final packaging time, you lose all dies instead of one.

            3. crote · · focus · HN ↗
              You make HBM by stacking a whole bunch of dies on top of each other. The signals from the upper dies need to pass through vias in the lower dies to get out - taking up valuable die space in a way which simply isn&#x27;t needed with regular DDR. Similarly, HBM has a far wider bus, so each individual die has, say, 16 banks of depth 32, rather than 4 banks of depth 128. That&#x27;s more control area needed per byte of memory.

              Those two combined already result in a huge reduction in bytes per mm2, so with the same wafer processing capacity you&#x27;re producing far less byte of memory. Add to that a complicated chain of HBM-specific packaging steps, and you&#x27;re now also losing a decent bunch of perfectly-fine dies because rather than putting it into DDR you tried making a HBM sandwich and screwed up.

              Even if the memory cells are the same and have an absolutely identical yield, HBM will always end up having a significantly lower output. That&#x27;s just the cost of stacking, but some people are willing to pay the per-gigabyte price penalty in return for the higher bandwidth.

            4. imtringued · · focus · HN ↗
              Classic DRAM stacks up to four wafers on top of each other and then is packaged with BGAs. The manufacturer can check the DRAM chips independently.

              Soldering the DRAM onto a PCB is such a reliable process that there is almost zero risk of defects and even if a defect occurs the damage is limited. If the DRAM is soldered onto a DIMM the risk of a defect on the non memory hardware is non-existent. If the DRAM is soldered straight onto an SBC or GPU, then the DRAM can be removed to save the precious SoC or GPU chips.

              Meanwhile HBM is the ultimate nightmare scenario. You stack up to 16 DRAM wafers on top of each other. One defect and the whole stack is worthless and that was actually the easy part.

              In stage two things get even worse. You now have your accelerator chip and you must place the HBM on that chip. E.g. Blackwell GB300 has eight HBM stacks and the accelerator chip has a bigger area than the HBM. You must get the packaging right eight times in a row or you have wasted not only the DRAM silicon, but also the accelerator silicon because HBM cannot be removed and defects are permanent.

              The issue here isn&#x27;t just the yield of the HBM (which is obviously lower if you have taller stacks) but rather the yield of the combined HBM-based product, which is why doesn&#x27;t make sense to say it needs more area but it is completely correct to say that HBM leads to more silicon being consumed. Hence it doesn&#x27;t make sense to talk about yield of the HBM itself, because it is always part of an integrated product.

      2. rkagerer · · focus · HN ↗
        I don&#x27;t fully understand the source of the &quot;total bytes&quot; constraint, but a major factor may be because HBM4 &#x2F; HBM4E can only make use of the footprint directly above the processor&#x2F;logic die (or in direct vicinity of its interconnect), while traditional DRAM can be placed further away where there&#x27;s lots of real estate on the motherboard.

        I gather a practical max ceiling today is a stack of 16 chips in height yielding 64GB?

        These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5. That&#x27;s what gives them their order-of-magnitude bandwidth speedup. But even though they technically pack in more capacity per square millimeter of motherboard, I gather they take up more space than older technologies once you account for the vias and interconnects to route all those signals.

        1. threecheese · · focus · HN ↗
          Thanks for that, just went down an interesting rabbit hole. Many of us were hoping this re-tooling would eventually trickle some fast RAM down to DRAM-exhausted PCs, but given it would require a rearchitecture of the motherboard it&#x27;s unlikely.
          1. sroussey · · focus · HN ↗
            HBM also trades bandwidth for latency, and your regular computing is much more sensitive to latency than bandwidth.
          2. craigjb · · focus · HN ↗
            HBM4 has over 2048 signals to the processor’s PHY with tight signal integrity requirements that require the HBM stack to be &lt; 0.5 mm from the processor die. That’s why HBM integration is done with interposers (soldered on the package). So, it’d be the CPU package that integrates it. Motherboard is too far away.
            1. Melatonic · · focus · HN ↗
              Kind of seems like we should be making chips with both. Big HBM stack on top as a sort of huge L5 cache like thing. And then a bunch of DRAM type sockets (like LPCAMM) around the exterior.
              1. craigjb · · focus · HN ↗
                For chips with integrated CPU+GPU+NPU, it could be worth it tech-wise. The GPU and NPU can eat HBM bandwidth. For general purpose CPU code, the HBM would likely not be worth it. It&#x27;s high bandwidth, but you trade latency, and general CPU code is branchy. Economics-wise, the HBM stacks alone will cost more than a consumer CPU (or APU).

                [edit] The packaging needed to support HBM is also much more expensive too. If demand for current HBM applications tanks and the manufacturing lines need filled, then maybe. Currently, the price point would make it very very niche.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.