I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
I wonder what "compute" would mean if CPUs were more efficient at matrix multiplication and vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.
On the compute side there's the issue of scalability. CPUs are designed to perform a handful of operations at one time. They typically have a small number of dedicated integer, float, and other ALU configurations. Having dedicated matrix multiplication instructions would still lock that to how many matrix-capable ALUs there are in the CPU.
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
<a href="https://unsloth.ai">https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
<a href="https://medium.com/@kailaspsudheer/the-transformers-arithmetic-527111099527" rel="nofollow">https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
The ideas survive in AVX-512 a.k.a. AVX10, not in AVX, which was a project parallel to Larrabee and resulting in an inferior ISA, which was adopted in the mainline Intel CPUs due to internal politics, not due to technical superiority.
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
tolugenius · · focus · HN ↗
actionfromafar · · focus · HN ↗
pjmlp · · focus · HN ↗
There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
aeve890 · · focus · HN ↗
Like in the existing SIMD technology?
pjmlp · · focus · HN ↗
rhdunn · · focus · HN ↗
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
<a href="https://unsloth.ai">https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
<a href="https://medium.com/@kailaspsudheer/the-transformers-arithmetic-527111099527" rel="nofollow">https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
pjmlp · · focus · HN ↗
<a href="https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_larrabee.pdf" rel="nofollow">https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_la...
adrian_b · · focus · HN ↗
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
pjmlp · · focus · HN ↗
Dwedit · · focus · HN ↗
alfiedotwtf · · focus · HN ↗
Dwedit · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]