‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. adrian17 · · focus · HN ↗
    > Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

    If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

    <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128

    1. 0x457 · · focus · HN ↗
      1.76 bpw number is kinda misleading if you compare it directly to IQ2&#x2F;Q2. The encoding is ternary, but the quantization procedure is way more sophisticated than &quot;round Qwen weights to {-1,0,+1}.&quot;

      They rotate the weights into a quantization-friendly basis first, then ternarize with per-group scales and error compensation.

      1. smallerize · · focus · HN ↗
        I think the 1.76 includes that. Plain ternary packed into bytes would be around 1.58.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.