‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. om8 · · focus · HN ↗
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    1. om8 · · focus · HN ↗
      If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.