Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Unofficial Hacker News client; not affiliated with Y Combinator.
adrian17 · · focus · HN ↗
If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?
<a href="https://news.ycombinator.com/item?id=49611128">https://news.ycombinator.com/item?id=49611128
nulld3v · · focus · HN ↗
The table claims it performs on par with UD-Q4_K_XL except on OCR.
Balinares · · focus · HN ↗