Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Unofficial Hacker News client; not affiliated with Y Combinator.
adrian17 · · focus · HN ↗
If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?
<a href="https://news.ycombinator.com/item?id=49611128">https://news.ycombinator.com/item?id=49611128
0x457 · · focus · HN ↗
They rotate the weights into a quantization-friendly basis first, then ternarize with per-group scales and error compensation.
smallerize · · focus · HN ↗