‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. infogulch · · focus · HN ↗
    So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

    1. kadushka · · focus · HN ↗
      By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
      1. FuckButtons · · focus · HN ↗
        Quantization is a category error, the thing you care about is not in weight space, so there’s unbounded error introduced by doing it. The thing you actually want to preserve is the knowledge manifold, but that is in a different vector space. Until we have some better understanding of how to interact with that space directly, rather than inferring it through distillation of reasoning traces, I would not anticipate truly low bit models to be useful.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.