‹ BackHN Continuity

Thread

The Painful Truth: The RAM Crisis Is Only Just the Beginning

62 points · 69 comments · perelin

  1. YuechenLi · · focus · HN ↗
    DRAM pricing has historically been cyclical in nature, and the only reason for the current high DRAM pricing is that current AI inference is unnecessarily RAM intensive/inefficient.

    Interesting effect is that since DRAM production tooling has been switched from DDR/GDDR to HBM, we may finally see the proliferation of HBM in consumer GPUs/accelerators after the bust cycle starts.

    1. Catloafdev · · focus · HN ↗
      >the only reason ... is that current AI inference is unnecessarily RAM intensive

      There's no indication that this will change, and every indication this will continue to grow. This is not some temporary thing. Demand is already exponentially higher than what is possible to produce, and there's no current reason to believe it has any ceiling.

      1. YuechenLi · · focus · HN ↗
        I've actually done some work and was able to get the full 12GB Z-image Turbo to run on my 8Gb 3070, and throughout the generation only 600MB of VRAM was allocated (you can get a speed-up by double buffering, but the maximum VRAM was still only ~1GB through inference). It's very experimental and pretty much requires you to write all the compute kernels directly specifically against that particular weight and statically allocate the memory at compile time instead of using current ML frameworks like Pytorch/JAX, but I think in time somebody else would figure it out.
        1. ASalazarMX · · focus · HN ↗
          The hardware savings at the scale Anthropic or Google use would me immense, it makes me wonder why no big player has done more optimization already. When DeepSeek showed how inneficient were the models of its time, I'd expected each to create a permanent optimization team with all the talent they have hired.

          I guess hardware is not that expensive to them in the grand scheme of things, at least not at this stage. OTOH, their propietary models might be thoughly optimized and we can't know, because they're still bound by supply contracts to buy the same amount of hardware nevertheless.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.