‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. yalok · · focus · HN ↗
    sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.

    And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...

    0. <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2402.17764" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.