‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. Chance-Device · · focus · HN ↗
    Let’s see, so if you get the same 1/9th the size compression ratio with GLM-5.3-Flash, then you’d end up with a ~72GB model that’s about as good as GPT-5.6 Sol (high), according to artificialanalysis.ai

    Which is within reach of some higher end consumer hardware, especially with layer offloading.

    You have to wonder what kind of trouble the “labs” are in when this is becoming possible. Lots of money, where’s the moat?

    1. ctolsen · · focus · HN ↗
      If we go with AA's benchmarks Qwen 3.8 27B is already slightly below Luna level which is in itself impressive, but with this compression it should be just slightly more below Luna level and could run on my old GTX 1070 that I'm now tempted to fire up. That's kinda nuts even allowing for small-model problems that I'm sure I'd see quite clearly.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.