PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.
If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.
in the case of large language models, the wire is the communication of your parameters between your layers of memory that is often the bottleneck. To do a forward pass, you need to use all parameters once, and so the communication between the compute and the storage is the bottleneck, and that bottleneck is also a bunch of wires.
om8 · · focus · HN ↗
janalsncm · · focus · HN ↗
If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.
mitxela · · focus · HN ↗
Kerbonut · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
317070 · · focus · HN ↗
mitxela · · focus · HN ↗