I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
The ideas survive in AVX-512 a.k.a. AVX10, not in AVX, which was a project parallel to Larrabee and resulting in an inferior ISA, which was adopted in the mainline Intel CPUs due to internal politics, not due to technical superiority.
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
tolugenius · · focus · HN ↗
pjmlp · · focus · HN ↗
<a href="https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_larrabee.pdf" rel="nofollow">https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_la...
adrian_b · · focus · HN ↗
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
pjmlp · · focus · HN ↗