‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. om8 · · focus · HN ↗
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    1. janalsncm · · focus · HN ↗
      PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

      If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

      1. mitxela · · focus · HN ↗
        which is important though since sending it across the wire over and over and over is actually the main bottleneck.
        1. Kerbonut · · focus · HN ↗
          Wire typically means internet connection, and it’s hardly the bottleneck
          1. [deleted] · · focus · HN ↗

            [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.