‹ BackHN Continuity

Thread

Breaking the 1.58-bit Barrier for Ternary LLMs

245 points · 41 comments · matt_d

  1. infogulch · · focus · HN ↗
    So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

    1. kadushka · · focus · HN ↗
      By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
      1. montroser · · focus · HN ↗
        Well, you could train directly at this bitrate.
        1. brookst · · focus · HN ↗
          Not an expert, but doesn’t that produce lower quality results, the same way a 1MP image isn’t lower quality than a 20mp image downscaled to 1MP? (Everything else equal)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.