‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. adrian17 · · focus · HN ↗
    > Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

    If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

    <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128

    1. edflsafoiewq · · focus · HN ↗
      I think the general idea is naive quantization falls apart below 4bpw but you can go lower with more sophisticated QAT-adjacent methods. Bonsai&#x27;s quantization method is proprietary though.
      1. om8 · · focus · HN ↗
        Could&#x27;ve been better if GGUF implemented QTIP format. GGUF representation is a major limitation for llama.cpp quantization performance
        1. edflsafoiewq · · focus · HN ↗
          They use their own llama fork anyway, so that shouldn&#x27;t matter.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.