Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Unofficial Hacker News client; not affiliated with Y Combinator.
adrian17 · · focus · HN ↗
If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?
<a href="https://news.ycombinator.com/item?id=49611128">https://news.ycombinator.com/item?id=49611128
edflsafoiewq · · focus · HN ↗
yowlingcat · · focus · HN ↗
One such method that I've been meaning to look into further is Tencent's AngelSlim QAT/PTQ approach. They did a Hy4 preview release thats an STQ_1_0 at 2.38 bpw:
<a href="https://huggingface.co/AngelSlim/Hy4-preview-GGUF" rel="nofollow">https://huggingface.co/AngelSlim/Hy4-preview-GGUF <a href="https://arxiv.org/abs/2602.21233" rel="nofollow">https://arxiv.org/abs/2602.21233
Of course, it's still 213g of VRAM I'd need so it's somewhat out of the range of what I can run locally. In contrast, this new Bonsai is nice because the original was already exciting for making use of low VRAM devices. Could breath new life into some of the older GPUs that were previously close to top of the line just quite VRAM constrained by modern standards and still quite cost effective for now.
om8 · · focus · HN ↗
edflsafoiewq · · focus · HN ↗