‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. simonw · · focus · HN ↗
    If you want to try out out the GGUFs from <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#these-files-need-our-llamacpp-build" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism&#x27;s llama.cpp fork to get them to work, from <a href="https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp&#x2F;releases&#x2F;tag&#x2F;prism-b10685-7dffb15" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp&#x2F;releases&#x2F;tag&#x2F;prism-...

    This should work:

      cd &#x2F;tmp
    
      # Get the Prism macOS runtime
      curl -fL https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp&#x2F;releases&#x2F;download&#x2F;prism-b10685-7dffb15&#x2F;llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz
      tar -xzf bonsai-runtime.tar.gz
    
      # Get the ~5.95 GB GGUF model:
      curl -fL https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf&#x2F;resolve&#x2F;main&#x2F;Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf
    
      # Run the server, I used port 8331
      .&#x2F;llama-prism-b10685-7dffb15&#x2F;llama-server \
        -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
        --port 8331 -ngl 99 -fa on -c 32768
    
    Then open http:&#x2F;&#x2F;localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this:

      uvx llm openai endpoint http:&#x2F;&#x2F;127.0.0.1:8331&#x2F;v1 \
        --model bonsai-2-27b --responses hi
    
    That&#x27;s running at ~20 token&#x2F;second for me on an M5 Pro (after a server restart I got 44 token&#x2F;second, not sure why), but I&#x27;m pretty sure something isn&#x27;t working right, on startup the server said &quot;ggml_metal_device_init: - the tensor API is not supported in this environment - disabling&quot;.
    1. [deleted] · · focus · HN ↗

      [deleted]

    2. simonw · · focus · HN ↗
      I used that to Generate an SVG of a pelican riding a bicycle:

      <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fba4cf3a88f4e7dc32994f2672150f770" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

      It took 18 minutes 20 seconds. Pretty decent for a 5.5GB model file.

      1. kadoban · · focus · HN ↗
        Honestly looks pretty good except whatever is going on with its booty. Is that an ass helmet? I cannot parse what&#x27;s going on there.
        1. Forgeties79 · · focus · HN ↗
          I think it’s supposed to be a wing
        2. bigwheels · · focus · HN ↗
          I like the lens effect behind the rear tire.
        3. tomcam · · focus · HN ↗
          Like you&#x27;ve never worn an ass helmet
          1. kadoban · · focus · HN ↗
            Only because I hadn&#x27;t previously thought of it xD Step up from the standard ass-hat for sure.
      2. [deleted] · · focus · HN ↗

        [deleted]

      3. rahimnathwani · · focus · HN ↗
        M1 Pro, same prompt, same cli options:

          32,706 tokens
          38min 19s
          14.22 t&#x2F;s
      4. raylad · · focus · HN ↗
        How does that compare with the bf16 version?

        For my &quot;Please recite Jabberwocky&quot; test the bf16 almost passes but the ternary and even fp8 versions fail badly.

        1. ctolsen · · focus · HN ↗
          Definitely a little better.

          <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;ctolsen&#x2F;b2883e7cbf5e4357fa04366019e60bfe" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;ctolsen&#x2F;b2883e7cbf5e4357fa04366019e6...

      5. shmoil · · focus · HN ↗
        Can you ask it for an SVG of a bicycle riding a pelican? Thanks.
    3. refibrillator · · focus · HN ↗
      Where did you get these instructions?

      They have a demo repo with a setup.sh script:

      <a href="https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;Bonsai-demo" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;Bonsai-demo

      The release tag and weight file you suggest doesn’t match what they wrote.

      1. simonw · · focus · HN ↗
        I figured them out, starting from the GGUF on Hugging Face.

        If you have found better instructions and they work then use those instead!

        Personally I prefer to download models directly rather than running some `.&#x2F;setup.sh` script where I need to then review what it does first.

        1. refibrillator · · focus · HN ↗
          Yeah just wanted to mention in case it explains the 2x lower throughout you are seeing on M5. To be fair their documentation is a bit inconsistent in some spots.

          Would be good to know if the release and weights from their demo repo work better. I’m trying on a 4090 and will report back.

          1. fnordpiglet · · focus · HN ↗
            They’re fairly useful as constrain domain classifiers due to the low memory requirements means you can stuff a lot of them into a less expensive GPU farm and get really decent throughout with pretty good results over all. At least that’s my experience. I wouldn’t bother using a tiny model for coding - but the world is full of abductive reasoning tasks that don’t involve coding.
    4. nikwen · · focus · HN ↗
      It would be great to have upstream llama.cpp support for this!
      1. iJohnDoe · · focus · HN ↗
        Agreed. They always sound exciting to try out but are such a pain to get working.
        1. Zetaphor · · focus · HN ↗
          I always just throw an agent at it. Is this the RSI I keep hearing about
          1. hedgehog · · focus · HN ↗
            RSI saves you from RSI
      2. xlazom00 · · focus · HN ↗
        They opened some pull requests and some of them are already merged
    5. rahimnathwani · · focus · HN ↗
      If you want to download the gguf to your regular huggingface cache directory instead of to &#x2F;tmp, you can download the model and run the server in one step:

        export HF_TOKEN=xxx # optional, speeds up the download
        
        .&#x2F;llama-prism-b10685-7dffb15&#x2F;llama serve \
          -hf prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf:PTQ1_0 \
          --port 8331 -ngl 99 -fa on -c 32768
    6. francisjp · · focus · HN ↗
      Thanks for all of your exploration in public Simon.

      Commenting because the fix I proposed was merged in roughly 49 commits after the PrismML Fork. The “tensor API is not supported” warning occurred because llama.cpp’s startup probe fails to compile a matmul2d kernel: Metal’s tensor headers require language version 4.0, but ggml-metal-device.m previously omitted MTLCompileOptions.languageVersion, disabling the API universally.

      Here’s a link to the diff if you want to try and update that fork to take advantage of the prefill gains afforded by the hardware: <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;pull&#x2F;27461&#x2F;changes" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;pull&#x2F;27461&#x2F;changes

      1. jb_briant · · focus · HN ↗
        That kind of issue is exactly why Im so happy to have LLMs, let it take one hour or trial and error instead of me spending a day digging traces
    7. wombat23 · · focus · HN ↗
      I managed to run it with RTX 3070 (8GB VRAM) following the &quot;Quickstart&quot; on HF model card with minor modifications (modify the architecture 86 for your own hardware):

        git clone https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp &amp;&amp; cd llama.cpp
        cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 &amp;&amp; cmake --build build -j
      
      Then downloaded &amp; verified `Ternary-Bonsai-2-27B-PTQ1_0.gguf` from HF and ran

        .&#x2F;llama.cpp&#x2F;build&#x2F;bin&#x2F;llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -fa on -c 32768 -b 256 -ub 64 --temp 1.0 --top-p 0.95 --top-k 20 -ctk q8_0 -ctv q8_0 -p &quot;hello world in x86 assembler&quot; -n 256
      
      the parameters were suggested by gpt-5.6-luna to reduce memory footprint, as the defaults ran OOM on my gpu. result looks good:

        [ Prompt: 165.6 t&#x2F;s | Generation: 40.6 t&#x2F;s ]
      
      would be nice if they upstreamed their changes so that it runs with the original llama.cpp
      1. aktenlage · · focus · HN ↗
        Would it speed up prompt processing if you increased the -ub (and -b) parameters.
        1. wombat23 · · focus · HN ↗
          I don&#x27;t see any significant speed up. also with defaults according to --help

            -b,    --batch-size N                   logical maximum batch size (default: 2048)
            -ub,   --ubatch-size N                  physical maximum batch size (default: 512)
          
          I also tried double the default. that also means that they can be left to default settings, apparently.
      2. ekianjo · · focus · HN ↗
        wow why is prompt processing so slow?
      3. zepearl · · focus · HN ↗
        Exact same test executed on my RTX 3060 (12 GiB VRAM, PCIe 3.0 4x slot):

          [ Prompt: 95.0 t&#x2F;s | Generation: 26.5 t&#x2F;s ]
        
        (the test&#x27;s prompt is very short but with longer ones the I get ~200 prompt processing rate, but I was hoping for a better token generation rate...)

        Am I understanding correctly that no draft model exists (will never exist or just currently does not exist yet)?

        There is no draft file in Huggingface&#x27;s repository ( <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf&#x2F;tree&#x2F;main" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf&#x2F;tr... ) and in the file &quot;scripts&#x2F;download_models.sh&quot; of the demo repository ( <a href="https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;Bonsai-demo&#x2F;blob&#x2F;main&#x2F;scripts&#x2F;download_models.sh" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;Bonsai-demo&#x2F;blob&#x2F;main&#x2F;scripts... ) I see this remark:

          if [ &quot;$_family&quot; = &quot;bonsai2&quot; ]; then
            the projector ships in the same repo; Bonsai 2 has no dspark drafter
      4. jimmySixDOF · · focus · HN ↗
        hummm wonder what the context length limit will be like this
      5. wombat23 · · focus · HN ↗
        UPDATE: I did more experiments - this time with llama-server and pi harness. i&#x27;ll leave the results here for posteriority:

          .&#x2F;llama.cpp&#x2F;build&#x2F;bin&#x2F;llama-server \
            -c &quot;$context_size&quot; \
            -ctk q4_0 \
            -ctv q4_0 \
            --no-models-autoload \
            --models-dir ~&#x2F;bonsai \
            --reasoning-preserve \
            -ngl 99
        
        the max context size i could serve is 64K on GPU only (the -ngl 99 setting). tested with pi harness and it is very fast. the &#x2F;thinking level always gets reset to off though and it is not very smart like this. haven&#x27;t figure out a way to fix that.
      6. ranger_danger · · focus · HN ↗
        It&#x27;s being worked on for upstream: <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;issues&#x2F;29058" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;issues&#x2F;29058
    8. ricardobayes · · focus · HN ↗
      Thanks for this, I never knew llama-server has a web ui until now.
      1. jakswa · · focus · HN ↗
        I love their web UI so much that I had AI slop all over it in a fit of fanboy-ism <a href="https:&#x2F;&#x2F;inkcap.click" rel="nofollow">https:&#x2F;&#x2F;inkcap.click
    9. ithkai92 · · focus · HN ↗
      Thanks for the headstart, I saw hf also has PQ2_0 and able to finetune the command and in a MBA M4 24GB averages around 10t&#x2F;s with the command.

      .&#x2F;llama-prism-b10685-7dffb15&#x2F;llama-server \ -m Ternary-Bonsai-2-27B-PQ2_0.gguf \ --port 8331 -ngl 99 -fa on -c 65536 --jinja \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.