‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. adrian17 · · focus · HN ↗
    > Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

    If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

    <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128

    1. edflsafoiewq · · focus · HN ↗
      I think the general idea is naive quantization falls apart below 4bpw but you can go lower with more sophisticated QAT-adjacent methods. Bonsai&#x27;s quantization method is proprietary though.
      1. yowlingcat · · focus · HN ↗
        That&#x27;s correct. I think there can certainly be issues even with 4bpw with naive quantization (IE you&#x27;ll notice far better results from a QAT 4bpw vs a naive 4bpw).

        One such method that I&#x27;ve been meaning to look into further is Tencent&#x27;s AngelSlim QAT&#x2F;PTQ approach. They did a Hy4 preview release thats an STQ_1_0 at 2.38 bpw:

        <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;AngelSlim&#x2F;Hy4-preview-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;AngelSlim&#x2F;Hy4-preview-GGUF <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2602.21233" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2602.21233

        Of course, it&#x27;s still 213g of VRAM I&#x27;d need so it&#x27;s somewhat out of the range of what I can run locally. In contrast, this new Bonsai is nice because the original was already exciting for making use of low VRAM devices. Could breath new life into some of the older GPUs that were previously close to top of the line just quite VRAM constrained by modern standards and still quite cost effective for now.

      2. om8 · · focus · HN ↗
        Could&#x27;ve been better if GGUF implemented QTIP format. GGUF representation is a major limitation for llama.cpp quantization performance
        1. edflsafoiewq · · focus · HN ↗
          They use their own llama fork anyway, so that shouldn&#x27;t matter.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.