‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. 2001zhaozhao · · focus · HN ↗
    I think if they made this for Qwen3.8-Next it could fit in a single 5090?
    1. kennywinker · · focus · HN ↗
      180b * 1.76 bits per weight = 39.6 gigabytes.

      Best you could realistically run in 32gb is like 28gb, or a 127B param model

      1. jokethrowaway · · focus · HN ↗
        Qwen3.8-Next, thanks to its new architecture, is quite fast even if part of it is streaming from disk
        1. kennywinker · · focus · HN ↗
          Totally. Any MoE model can have experts swapped in and out from disk or system ram. I only framed it this way because the question was about the model fitting in vram.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.