‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. om8 · · focus · HN ↗
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    1. janalsncm · · focus · HN ↗
      PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

      If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

      1. mitxela · · focus · HN ↗
        which is important though since sending it across the wire over and over and over is actually the main bottleneck.
        1. Kerbonut · · focus · HN ↗
          Wire typically means internet connection, and it’s hardly the bottleneck
          1. [deleted] · · focus · HN ↗

            [deleted]

          2. 317070 · · focus · HN ↗
            in the case of large language models, the wire is the communication of your parameters between your layers of memory that is often the bottleneck. To do a forward pass, you need to use all parameters once, and so the communication between the compute and the storage is the bottleneck, and that bottleneck is also a bunch of wires.
            1. mitxela · · focus · HN ↗
              The other bottleneck is the amount of fast storage, which compression also improves.
      2. om8 · · focus · HN ↗
        > If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

        That’s why you need to use efficient gemm kernels like FLUTE for inference. They are ~as good as what you can do with ternary quantization.

      3. WithinReason · · focus · HN ↗
        And storing it in memory. Memory is expensive.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.