This comes at a perfect time. The public doesn't need computers, they can just use their phones.
If the consumer had access to this RAM, they might all just run local or semi-local AI. It's important to outbid them so you can rent AI to them, and extract money from them in a million other ways while they use it.
Make RAM half the price it was before and you'd still have very few people trying to do local AI on it because it's just so damn slow for the task. Having 12 channels of DDR5 in an Epyc system with NVMe storage is still one of the slowest ways I can run LLMs locally.
Funny, I'm running local LLMs in a modest iMac M4, they are slow indeed but not utterly slow. A dedicated system should be way faster I'd guess.
The M4 (soldered wide LPDDR5X + unified architecture + GPU with neural cores) is also closer in design to a dedicated GPU than an Epyc server with DDR5 or the components in the article.
But yes, a dedicated consumer GPU can still typically outpace the M4 quite well (if things fit in VRAM) and both will have their socks blown clean off by a hyperscaler GPU cluster. For some hard numbers I can run a ~24 GB model on my 5090 (about as fast as you can get on a single consumer class device) about 3x-4x faster than on my M4 mini and I'd still consider that pretty slow to running models 10x the size in the cloud.
No one is running cpu-only inference anymore. Deepseek flash / Qwen flash next et al only need ~12-24 gb of VRAM for the active experts, so anyone with a decent gpu + lots of RAM + server cpu can run them at okay speeds.
Or one can also run Strix Halo / DGX spark / Apple. They handle MoEs well and use normal (not HBM) ram.
some benchmarks from locallama:
Deepseek V4:
"~25t/s decode at full quant, no speculative decoding, 2x3090 (only one being used for this model) + 9684x w/ 12 channel ddr5 4800, latest llama.cpp"
"Getting about 300 tk/sec pp and 23 tk/sec decode on Mac M3 Ultra 256 GB. Using original full precision model weights."
pessimizer · · focus · HN ↗
If the consumer had access to this RAM, they might all just run local or semi-local AI. It's important to outbid them so you can rent AI to them, and extract money from them in a million other ways while they use it.
zamadatix · · focus · HN ↗
ASalazarMX · · focus · HN ↗
zamadatix · · focus · HN ↗
But yes, a dedicated consumer GPU can still typically outpace the M4 quite well (if things fit in VRAM) and both will have their socks blown clean off by a hyperscaler GPU cluster. For some hard numbers I can run a ~24 GB model on my 5090 (about as fast as you can get on a single consumer class device) about 3x-4x faster than on my M4 mini and I'd still consider that pretty slow to running models 10x the size in the cloud.
upboundspiral · · focus · HN ↗
No one is running cpu-only inference anymore. Deepseek flash / Qwen flash next et al only need ~12-24 gb of VRAM for the active experts, so anyone with a decent gpu + lots of RAM + server cpu can run them at okay speeds.
Or one can also run Strix Halo / DGX spark / Apple. They handle MoEs well and use normal (not HBM) ram.
some benchmarks from locallama:
Deepseek V4: "~25t/s decode at full quant, no speculative decoding, 2x3090 (only one being used for this model) + 9684x w/ 12 channel ddr5 4800, latest llama.cpp"
"Getting about 300 tk/sec pp and 23 tk/sec decode on Mac M3 Ultra 256 GB. Using original full precision model weights."
Qwen-Flash-Next: <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wom3fe/qwen_38_flash_next_q4_k_m_130k_context_q8_cache/" rel="nofollow">https://www.reddit.com/r/LocalLLaMA/comments/1wom3fe/qwen_38...
<a href="https://www.reddit.com/r/LocalLLaMA/comments/1vcaztx/what_speeds_are_everyone_getting_with_deepseek_v4/" rel="nofollow">https://www.reddit.com/r/LocalLLaMA/comments/1vcaztx/what_sp...
striking · · focus · HN ↗