‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. abraxas · · focus · HN ↗
    I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
    1. kamranjon · · focus · HN ↗
      "Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."
      1. pizza234 · · focus · HN ↗
        Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!
        1. sisve · · focus · HN ↗
          They mention 5090 with regards to speed, Q6 will not have that speed?

          And speed matters a lot for many use cases

          1. wincy · · focus · HN ↗
            With Ninfer and Qwen 3.8 27b it uses a groupwise int mixed quant, and it gets 160 tokens/sec. The mixed quant is between 4 and 6 bits.
            1. Foobar8568 · · focus · HN ↗
              Ninfer is compatible with a nvfp4 model for the 27b. Also nowadays I prefer to use the byteshape one, I get less loops, and I am not sure if I really see a difference in speed or quality. Pure vibe agentic coding on a C++ codebase or ocaml one, ocaml one has codex as reviewer as I am more interested by that project, the other is more for fun.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.