‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. Chance-Device · · focus · HN ↗
    Let’s see, so if you get the same 1/9th the size compression ratio with GLM-5.3-Flash, then you’d end up with a ~72GB model that’s about as good as GPT-5.6 Sol (high), according to artificialanalysis.ai

    Which is within reach of some higher end consumer hardware, especially with layer offloading.

    You have to wonder what kind of trouble the “labs” are in when this is becoming possible. Lots of money, where’s the moat?

    1. kllrnohj · · focus · HN ↗
      The labs still have performance as a differentiator for coding usages and similar, and for other things there's still all the same reasons people switched cloud hosted stuff in the first place. AWS & friends didn't get popular because the hardware was out of reach, after all.
      1. Chance-Device · · focus · HN ↗
        That’s an argument for a cloud LLM service, sure, but the hyperscalers can do that by themselves with the weights.

        What’s the moat for trillion dollar AI companies? Access-anywhere convenience for models as good as everyone else’s?

        1. kllrnohj · · focus · HN ↗
          the trillion dollar AI companies will be the chip designers for the hyperscalers, and probably not worth a trillion dollars as a result
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.