‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. infogulch · · focus · HN ↗
    So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

    1. kadushka · · focus · HN ↗
      By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
      1. mixermachine · · focus · HN ↗
        I also no longer trust benchmarks on this one. When the context gets a bit longer and the problem harder low quant models often produce worse output for me. Sometimes they even loop.

        Interestingly different formats also often behave differently. GGUF unsloth is so far the best for me.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.