Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
Unofficial Hacker News client; not affiliated with Y Combinator.
fooker · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU.
So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second!
These are not normal requirements, other than for a GPU.
fooker · · focus · HN ↗
b112 · · focus · HN ↗
fooker · · focus · HN ↗
Consider the 'MMA N matrices' primitive modern CPUs are starting to support. For the current generation of CPUs, N is a constant like 16 or 32, but there's nothing preventing it from being 1024 or larger if we have more memory bandwidth.
All this with a single instruction.
articulatepang · · focus · HN ↗
pixl97 · · focus · HN ↗
We'll have to figure out how to read and right to the surface of a black hole to get speeds that high.
imtringued · · focus · HN ↗
If you have infinite memory bandwidth you just move the bottleneck back to compute so both have to grow simultaneously in lockstep.
What you should have said is that CPUs have so much compute headroom for matrix vector multiplication that simply adding more memory bandwidth would make them faster so every improvement in memory bandwidth is welcome.
fooker · · focus · HN ↗
The "move the bottleneck back to compute" bit is changing rapidly though. The first time a major hardware company ships a PIM chip, you can push for a few orders of magnitude more data through without being compute bound.