Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Unofficial Hacker News client; not affiliated with Y Combinator.
simonw · · focus · HN ↗
This should work:
Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this: That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".simonw · · focus · HN ↗
<a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fba4cf3a88f4e7dc32994f2672150f770" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It took 18 minutes 20 seconds. Pretty decent for a 5.5GB model file.
kadoban · · focus · HN ↗
Forgeties79 · · focus · HN ↗
bigwheels · · focus · HN ↗
tomcam · · focus · HN ↗
kadoban · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
rahimnathwani · · focus · HN ↗
raylad · · focus · HN ↗
For my "Please recite Jabberwocky" test the bf16 almost passes but the ternary and even fp8 versions fail badly.
ctolsen · · focus · HN ↗
<a href="https://gist.github.com/ctolsen/b2883e7cbf5e4357fa04366019e60bfe" rel="nofollow">https://gist.github.com/ctolsen/b2883e7cbf5e4357fa04366019e6...
shmoil · · focus · HN ↗