‹ BackHN Continuity

Thread

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

589 points · 200 comments · JonSchneider

  1. simonw · · focus · HN ↗
    If you want to try out out the GGUFs from <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#these-files-need-our-llamacpp-build" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism&#x27;s llama.cpp fork to get them to work, from <a href="https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp&#x2F;releases&#x2F;tag&#x2F;prism-b10685-7dffb15" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp&#x2F;releases&#x2F;tag&#x2F;prism-...

    This should work:

      cd &#x2F;tmp
    
      # Get the Prism macOS runtime
      curl -fL https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp&#x2F;releases&#x2F;download&#x2F;prism-b10685-7dffb15&#x2F;llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz
      tar -xzf bonsai-runtime.tar.gz
    
      # Get the ~5.95 GB GGUF model:
      curl -fL https:&#x2F;&#x2F;huggingface.co&#x2F;prism-ml&#x2F;Ternary-Bonsai-2-27B-gguf&#x2F;resolve&#x2F;main&#x2F;Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf
    
      # Run the server, I used port 8331
      .&#x2F;llama-prism-b10685-7dffb15&#x2F;llama-server \
        -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
        --port 8331 -ngl 99 -fa on -c 32768
    
    Then open http:&#x2F;&#x2F;localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this:

      uvx llm openai endpoint http:&#x2F;&#x2F;127.0.0.1:8331&#x2F;v1 \
        --model bonsai-2-27b --responses hi
    
    That&#x27;s running at ~20 token&#x2F;second for me on an M5 Pro (after a server restart I got 44 token&#x2F;second, not sure why), but I&#x27;m pretty sure something isn&#x27;t working right, on startup the server said &quot;ggml_metal_device_init: - the tensor API is not supported in this environment - disabling&quot;.
    1. wombat23 · · focus · HN ↗
      I managed to run it with RTX 3070 (8GB VRAM) following the &quot;Quickstart&quot; on HF model card with minor modifications (modify the architecture 86 for your own hardware):

        git clone https:&#x2F;&#x2F;github.com&#x2F;PrismML-Eng&#x2F;llama.cpp &amp;&amp; cd llama.cpp
        cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 &amp;&amp; cmake --build build -j
      
      Then downloaded &amp; verified `Ternary-Bonsai-2-27B-PTQ1_0.gguf` from HF and ran

        .&#x2F;llama.cpp&#x2F;build&#x2F;bin&#x2F;llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -fa on -c 32768 -b 256 -ub 64 --temp 1.0 --top-p 0.95 --top-k 20 -ctk q8_0 -ctv q8_0 -p &quot;hello world in x86 assembler&quot; -n 256
      
      the parameters were suggested by gpt-5.6-luna to reduce memory footprint, as the defaults ran OOM on my gpu. result looks good:

        [ Prompt: 165.6 t&#x2F;s | Generation: 40.6 t&#x2F;s ]
      
      would be nice if they upstreamed their changes so that it runs with the original llama.cpp
      1. ekianjo · · focus · HN ↗
        wow why is prompt processing so slow?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.