‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. huseyinkeles · · focus · HN ↗
    Testing on a MBP m4 pro 24gb

    ~100t/s prefill, ~15t/s, dropping to ~10t/s later with 64k context.

    The issue is I have yet to find a useful agentic local llm that I can run on this machine.

    Just given a relatively simple task on a swift app, took 25 minutes, brainstorming like crazy but can not decide on what to do. Eventually I killed it. GPT 5.6 sol-medium took 3 minutes to complete the same task for reference.

    1. aetherspawn · · focus · HN ↗
      Gemma 30B with 256K context runs at 20 tok/sec on my M3 Max with 128GB RAM so I think there’s something wrong with your setup. This should run at ~30-40 toks. Maybe your inference engine is not optimised for Mac.
      1. piyh · · focus · HN ↗
        30B is useless on 24 gigs of ram as there's ~4 gigs of ram left for everything else even with unsloth quants
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.