Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Unofficial Hacker News client; not affiliated with Y Combinator.
simonw · · focus · HN ↗
This should work:
Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this: That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".wombat23 · · focus · HN ↗
zepearl · · focus · HN ↗
Am I understanding correctly that no draft model exists (will never exist or just currently does not exist yet)?
There is no draft file in Huggingface's repository ( <a href="https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tree/main" rel="nofollow">https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tr... ) and in the file "scripts/download_models.sh" of the demo repository ( <a href="https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts/download_models.sh" rel="nofollow">https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts... ) I see this remark: