"It used Voronoi Cells and reduced most math to 8- and 16-bit integer calculations with one or two single-precision floating-point calculations."
I wonder if it can be used to speed up LLM inference.
This seems worth the read but I didn't dig too deep <a href="https://medium.com/@sergiopr89/float-point-quantization-the-maths-behind-it-explained-for-everyone-fc674d313d67" rel="nofollow">https://medium.com/@sergiopr89/float-point-quantization-the-...
JPLeRouzic · · focus · HN ↗
"It used Voronoi Cells and reduced most math to 8- and 16-bit integer calculations with one or two single-precision floating-point calculations."
I wonder if it can be used to speed up LLM inference.
Neywiny · · focus · HN ↗
JPLeRouzic · · focus · HN ↗
Neywiny · · focus · HN ↗