‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. huseyinkeles · · focus · HN ↗
    Testing on a MBP m4 pro 24gb

    ~100t/s prefill, ~15t/s, dropping to ~10t/s later with 64k context.

    The issue is I have yet to find a useful agentic local llm that I can run on this machine.

    Just given a relatively simple task on a swift app, took 25 minutes, brainstorming like crazy but can not decide on what to do. Eventually I killed it. GPT 5.6 sol-medium took 3 minutes to complete the same task for reference.

    1. aetherspawn · · focus · HN ↗
      Gemma 30B with 256K context runs at 20 tok/sec on my M3 Max with 128GB RAM so I think there’s something wrong with your setup. This should run at ~30-40 toks. Maybe your inference engine is not optimised for Mac.
      1. huseyinkeles · · focus · HN ↗
        I just used their `Bonsai-demo` repo like this;

        `cd ~/Code/Bonsai-demo && BONSAI_CTX=65536 ./scripts/start_llama_server.sh`

        then used it in a very minimalistic pi with a very small system prompt.

        Didn't spend much time to try to optimize it tbh, but my issue was not the speed. it just could not make a decision on how to implement the task, kept going on an on.

        1. aetherspawn · · focus · HN ↗
          Yes, I use LM Studio with MLX support, which is specifically faster on M series Macs. I am not sure if llama is the same, but I guess what I’m saying is if you want the performance to be good on M series you have to use models packaged in the right format.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.