Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Unofficial Hacker News client; not affiliated with Y Combinator.
simonw · · focus · HN ↗
This should work:
Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this: That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".simonw · · focus · HN ↗
<a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fba4cf3a88f4e7dc32994f2672150f770" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It took 18 minutes 20 seconds. Pretty decent for a 5.5GB model file.
[deleted] · · focus · HN ↗
[deleted]