‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. om8 · · focus · HN ↗
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    1. janalsncm · · focus · HN ↗
      PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

      If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

      1. om8 · · focus · HN ↗
        > If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

        That’s why you need to use efficient gemm kernels like FLUTE for inference. They are ~as good as what you can do with ternary quantization.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.