Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
Unofficial Hacker News client; not affiliated with Y Combinator.
fooker · · focus · HN ↗
nutjob2 · · focus · HN ↗
People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.
fooker · · focus · HN ↗
They have managed to pull this sort of thing off many many times. <a href="https://en.wikipedia.org/wiki/Reality_distortion_field" rel="nofollow">https://en.wikipedia.org/wiki/Reality_distortion_field
Kon5ole · · focus · HN ↗
davrosthedalek · · focus · HN ↗
This is really not a limit because of unified memory -- in principle, PCIe GPUs could read/write main memory without the CPU. But it's a limit for /fast/ unified memory, because fast means close.
So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages: a) probably faster transfer CPU<->GPU (but that's an implementation choice for the non-unified case b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.
nvme0n1p1 · · focus · HN ↗
<a href="https://x.com/Lina_Hoshino/status/1820947147312820497" rel="nofollow">https://x.com/Lina_Hoshino/status/1820947147312820497
stymaar · · focus · HN ↗
Kon5ole · · focus · HN ↗
fooker · · focus · HN ↗
The reality distortion is that people seem to believe it's HBM, or somehow it gives you extraordinary amounts of vram. Neither are really true.
jorvi · · focus · HN ↗
DDR is optimized for latency and stability at the cost of bandwidth whilst GDDR is optimized for bandwidth at the cost of latency and stability. GDDR is pushed so hard these days that a small percentage of errors is expected and corrected because this is still faster than running it slower but more accurate.
GDDR7 often has 10-20x the total bandwidth but 3x the latency of DDR5. Graphical workloads want as much bandwidth as possible but care relatively little for latency. Conversely, applications love low latency but don't really see any performance benefit from higher bandwidth.
So basicallyt you have workloads that are diametrically opposed and running unified memory forces you to compromise.
Melatonic · · focus · HN ↗