‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. avaer · · focus · HN ↗
    LLM quants seem to eerily converge to modern/not so modern graphics techniques. You wouldn't think it would apply but it's obvious in hindsight. In fact mining graphics ideas is probably a good inspiration for efficient LLM architecture.

    For example, the Hadamard activation transform used here feels a lot like multiplying Fourier basis ala DFT; strong parallels to how image codecs work to make the residuals more compressible (especially discrete block codecs like are used in GPU compressed textures).

    I thought I was being clever suggesting that you could even abuse texture decode units to efficiently sample compressed LLMs with hardware; turns out Apple foundation models are already doing this [1].

    [1] <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2507.13575" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2507.13575

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.