CPU DRAM can't really be used for inference efficiently -- inference mostly wants memory bandwidth, not memory capacity, and GPU DRAM has >10x more bandwidth. The fabs can switch between them but you can't switch after the fact.
This is just wrong unless you're using some confusing definition. Notice that companies trying to do lots of inference aren't looking for CPUs, they're looking for GPUs/TPUs.
Ah well, the opposite actually. Your definition, while traditional, is wrong. It was right some time ago. But it's wrong now. At least according to Intel and Nvidia in their disclosures. Intel was literally talking about their sales to "companies that do lots of inference" when they said it... so...
sroussey · · focus · HN ↗
why_only_15 · · focus · HN ↗
halJordan · · focus · HN ↗
why_only_15 · · focus · HN ↗
halJordan · · focus · HN ↗