‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. infogulch · · focus · HN ↗
    So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

    1. NostraDavid · · focus · HN ↗
      >baked into hardware as custom silicon

      Taalas (Acquired by AMD, back in August) created Jimmy[0], a little chat app that runs on a POC chip with ~14k tps. Yes, 14,000 tokens per second. Sure, it's just a 8B model or so (Llama 3.1 8B), but I can imagine that having a 1.58-bit model might be helpful for their next chip.

      Heck, what would happen if you used a dLLM (d for diffusion)?

      [0]: <a href="https:&#x2F;&#x2F;chatjimmy.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;chatjimmy.ai&#x2F;

      1. nwah1 · · focus · HN ↗
        They don&#x27;t use ternary quantization. But, they could, if they wanted.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.