Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Unofficial Hacker News client; not affiliated with Y Combinator.
simonw · · focus · HN ↗
This should work:
Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this: That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".francisjp · · focus · HN ↗
Commenting because the fix I proposed was merged in roughly 49 commits after the PrismML Fork. The “tensor API is not supported” warning occurred because llama.cpp’s startup probe fails to compile a matmul2d kernel: Metal’s tensor headers require language version 4.0, but ggml-metal-device.m previously omitted MTLCompileOptions.languageVersion, disabling the API universally.
Here’s a link to the diff if you want to try and update that fork to take advantage of the prefill gains afforded by the hardware: <a href="https://github.com/ggml-org/llama.cpp/pull/27461/changes" rel="nofollow">https://github.com/ggml-org/llama.cpp/pull/27461/changes
jb_briant · · focus · HN ↗