Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
Unofficial Hacker News client; not affiliated with Y Combinator.
HarHarVeryFunny · · focus · HN ↗
The lack of ASML EUV machines certainly hurts, and pushing DUV so hard results in abysmal yields of good chips, but you can compensate by running more wafers or making smaller chips, and the net result is that Huawei's Ascend production volume is limited by CXMT's HBM capacity not processor dies.
The problem is that HBM manufacture requires many steps (die thinning, via drilling, plating, alignment) where the equipment used by everyone else (Samsung, SK Hynix, Micron) is also blocked by sanctions, so the Chinese are having to develop all of this themselves too, which they have, but yields are currently low, even when using shorter HBM stacks.
andy_ppp · · focus · HN ↗
martinald · · focus · HN ↗
Prefill (input tokens) is heavily compute bound. And the ratio of input to output continues to rise, as typically in agentic sessions you have a few tokens output for a tool call and (many) thousands of input from the tool result.
Then you have cached input tokens, which is a totally different issue, system RAM or NVMe bound.
Obviously output tokens is VRAM memory bandwidth bound, but this is less and less of the bottleneck these days for overall agentic speed.
com2kid · · focus · HN ↗
andy_ppp · · focus · HN ↗