‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. adrian17 · · focus · HN ↗
    > Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

    If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

    <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49611128

    1. nulld3v · · focus · HN ↗
      There&#x27;s a table on the HF page that compares it against UD-Q4_K_XL and IQ2_XXS (you need to expand the dropdown): <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#full-per-benchmark-results" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#fu...

      The table claims it performs on par with UD-Q4_K_XL except on OCR.

      1. Balinares · · focus · HN ↗
        I wonder how well it performs in practice, because I can&#x27;t help seriously doubting those benchmarks. That would put this 6GB model in Opus 4.6+ ballpark. Granted, that&#x27;s mostly to Qwen 3.8&#x27;s credit, but it&#x27;s hard to believe that Qwen&#x27;s already unbelievable capability density can still be compressed this much more.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.