Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
kamranjon · · focus · HN ↗
kadoban · · focus · HN ↗
This will hopefully be better, though it'd be a _very_ surprising increase in performace at the size they say. Would love to see more about how it benchmarks.
spijdar · · focus · HN ↗
orsorna · · focus · HN ↗
kadoban · · focus · HN ↗
orsorna · · focus · HN ↗
lta · · focus · HN ↗
redox99 · · focus · HN ↗
Zambyte · · focus · HN ↗
kadoban · · focus · HN ↗
djkoolaide · · focus · HN ↗
Prism's llama.cpp fork only has the kernels for CUDA, CPU and Vulkan. No SYCL at all :(
kennywinker · · focus · HN ↗
abraxas · · focus · HN ↗
kamranjon · · focus · HN ↗
pizza234 · · focus · HN ↗
sisve · · focus · HN ↗
And speed matters a lot for many use cases
selectodude · · focus · HN ↗
wincy · · focus · HN ↗
Foobar8568 · · focus · HN ↗
pwython · · focus · HN ↗
blurbleblurble · · focus · HN ↗
azatom · · focus · HN ↗
_kulang · · focus · HN ↗
azatom · · focus · HN ↗
<a href="https://xkcd.com/3038/" rel="nofollow">https://xkcd.com/3038/
Havoc · · focus · HN ↗
Aurornis · · focus · HN ↗
Remember to clear the downloaded weights afterward.
Like the last model, it's amazing they work as well as they do. Use it for any longer task and they fall apart spectacularly and in interesting ways.
outofpaper · · focus · HN ↗
Aurornis · · focus · HN ↗
SXX · · focus · HN ↗
trvz · · focus · HN ↗
SXX · · focus · HN ↗
14u2c · · focus · HN ↗
jeroenhd · · focus · HN ↗
throwa356262 · · focus · HN ↗
I really hope they stop using PowerVR in pixel 11.
aschobel · · focus · HN ↗
z2 · · focus · HN ↗
all2 · · focus · HN ↗
codebje · · focus · HN ↗
JonSchneider · · focus · HN ↗
sroussey · · focus · HN ↗
verdverm · · focus · HN ↗
simonw · · focus · HN ↗
This should work:
Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this: That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".OutOfHere · · focus · HN ↗
simonw · · focus · HN ↗
<a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fba4cf3a88f4e7dc32994f2672150f770" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It took 18 minutes 20 seconds. Pretty decent for a 5.5GB model file.
kadoban · · focus · HN ↗
Forgeties79 · · focus · HN ↗
bigwheels · · focus · HN ↗
tomcam · · focus · HN ↗
kadoban · · focus · HN ↗
rahimnathwani · · focus · HN ↗
After 12 minutes and 12,000 output tokens, it's just finished its overall plan and is writing some of the code it will use later...
rahimnathwani · · focus · HN ↗
raylad · · focus · HN ↗
For my "Please recite Jabberwocky" test the bf16 almost passes but the ternary and even fp8 versions fail badly.
ctolsen · · focus · HN ↗
<a href="https://gist.github.com/ctolsen/b2883e7cbf5e4357fa04366019e60bfe" rel="nofollow">https://gist.github.com/ctolsen/b2883e7cbf5e4357fa04366019e6...
shmoil · · focus · HN ↗
refibrillator · · focus · HN ↗
They have a demo repo with a setup.sh script:
<a href="https://github.com/PrismML-Eng/Bonsai-demo" rel="nofollow">https://github.com/PrismML-Eng/Bonsai-demo
The release tag and weight file you suggest doesn’t match what you wrote.
simonw · · focus · HN ↗
If you have found better instructions and they work then use those instead!
Personally I prefer to download models directly rather than running some `./setup.sh` script where I need to then review what it does first.
refibrillator · · focus · HN ↗
Would be good to know if the release and weights from their demo repo work better. I’m trying on a 4090 and will report back.
fnordpiglet · · focus · HN ↗
nikwen · · focus · HN ↗
iJohnDoe · · focus · HN ↗
Zetaphor · · focus · HN ↗
hedgehog · · focus · HN ↗
xlazom00 · · focus · HN ↗
rahimnathwani · · focus · HN ↗
francisjp · · focus · HN ↗
Commenting because the fix I proposed was merged in roughly 49 commits after the PrismML Fork. The “tensor API is not supported” warning occurred because llama.cpp’s startup probe fails to compile a matmul2d kernel: Metal’s tensor headers require language version 4.0, but ggml-metal-device.m previously omitted MTLCompileOptions.languageVersion, disabling the API universally.
Here’s a link to the diff if you want to try and update that fork to take advantage of the prefill gains afforded by the hardware: <a href="https://github.com/ggml-org/llama.cpp/pull/27461/changes" rel="nofollow">https://github.com/ggml-org/llama.cpp/pull/27461/changes
jb_briant · · focus · HN ↗
wombat23 · · focus · HN ↗
aktenlage · · focus · HN ↗
wombat23 · · focus · HN ↗
ekianjo · · focus · HN ↗
zepearl · · focus · HN ↗
jimmySixDOF · · focus · HN ↗
wombat23 · · focus · HN ↗
ranger_danger · · focus · HN ↗
ricardobayes · · focus · HN ↗
jakswa · · focus · HN ↗
ithkai92 · · focus · HN ↗
./llama-prism-b10685-7dffb15/llama-server \ -m Ternary-Bonsai-2-27B-PQ2_0.gguf \ --port 8331 -ngl 99 -fa on -c 65536 --jinja \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
miffy900 · · focus · HN ↗
I keep seeing this being used when people talk about efficiency or performance gains and it's just very unintuitive language.
hamandcheese · · focus · HN ↗
simondotau · · focus · HN ↗
D-Machine · · focus · HN ↗
Rather, "nine times larger" means multiplied by nine, and "nine times smaller" means divided by nine. This is basic and not particularly awkward, certainly not more so than e.g. positive/negative correlation, or many much more awkward and more common linguistic constructions, IMO.
If you have to edit out words (i.e. context) to argue a phrase doesn't make sense... I am not sure what mental model you have for natural language, exactly, but it certainly isn't a very robust one.
simondotau · · focus · HN ↗
[dead]
dools · · focus · HN ↗
Likewise if I repeatedly subtract a number 9 times to make it “9 times smaller” I can omit the subtracted bit.
simondotau · · focus · HN ↗
dools · · focus · HN ↗
If you divide 81 by 9 you subtract 9 from 81 until you get to 0. The number of TIMES you do that is the result “81 divided by 9”.
simondotau · · focus · HN ↗
You chose 81 because it’s the one number that makes your argument appear to work. Try the same reasoning with 80 ÷ 9.
dools · · focus · HN ↗
[dead]
_carbyau_ · · focus · HN ↗
Conversely, 9x [filesize/natural number] is bigger. Every time. At least in the basic maths used by most people. There is no conversion into other units.
Therefore "9x smaller" when talking about a natural number like filesize is a nonsense statement in logic terms. If you strive for unambiguous phrasing - which is a significant part of the programming experience - this logical nonsense might well perturb you.
But english language is a flexible thing and if the phrase communicates your intent to your audience then that's fine by me.
dimes · · focus · HN ↗
jonhohle · · focus · HN ↗
Multiplying some scale by units per period makes sense and is both linguistically and mathematically sound.
D-Machine · · focus · HN ↗
Notice how you had to remove the word "smaller" to come to this reading. So "9x smaller" is, by one argument, perfectly logical, and you arguing that the "larger" is always implied in "9x" is not logical at all.
UI_at_80x24 · · focus · HN ↗
Kinrany · · focus · HN ↗
Taek · · focus · HN ↗
Idioms don't have to make literal sense or being linguistically/mathematically correct to be useful. All that matters is that other people know exactly what you mean when you say it.
And, pretty much universally, if I tell someone "the compressed file is 10x smaller than the original", they are going to know what I mean is that the byte size is 10% of the size of the original.
That makes it an idiom that is perfectly okay for everyday use.
zamadatix · · focus · HN ↗
One other way to map both types of linguistic statement consistently to math is to interpret "9x" as "there is a 9 times difference between these two things" and then "smaller"/"larger" tells you which end of that separation the subject is (rather than specifying whether the multiplication builds up or down).
kevinwang · · focus · HN ↗
zamadatix · · focus · HN ↗
Taek · · focus · HN ↗
This is the same. Taken literally, "9x smaller" is nonsense, but everyone who hears that phrase knows exactly what mathematical operation you are referring to, thus it's a totally acceptable way to express that you mean to say 11.11% as large.
asterix_pano · · focus · HN ↗
DoctorOetker · · focus · HN ↗
"9 times larger" => multiply the numerator of the fraction with 9
"9 times smaller" => multiply the denominator of the fraction with 9
Consider a bag of identical resistors with resistance R
Add 9 of them in series: the resistance is 9 times larger.
Add 9 of them in parallel: the resistance is 9 times smaller.
Its the series vs parallel dictionary wars all over again...
Check my other comment.
fwip · · focus · HN ↗
This is another pet peeve that I have with the way we phrase these things: It should be 9x "as large" and 8x "larger."
The way I'd usually phrase this, to avoid ambiguity, is "this new model is 11% the size (of the original)." Or, the other way, "the old model is 9x the size".
rpdillon · · focus · HN ↗
miffy900 · · focus · HN ↗
if I say 'this Apple M5 chip is 3x faster' than this intel chip, it implies two things:
- the run time of most operations that runs on it is now reduced (so one quantity is smaller)
- but also: MORE WORK is being completed per unit of time compared to the intel chip (so this quantity is greater)
So yes, a greater quantity is being measured in the apple chip compared to the intel when you say apple is N times faster. i guess this a quirk with the word 'faster' - it actually measures two things, time and work performed per unit of time. the word smaller just measures size.
like, if I give a customer a cup of coffee one day, then give them a SMALLER cup of coffee the next day, for the same price, but declare "It's now 2X more space efficient!" i.e. it's now half the size, I'm certain the customer is gonna be pissed.
rpdillon · · focus · HN ↗
atiedebee · · focus · HN ↗
fwip · · focus · HN ↗
comradesmith · · focus · HN ↗
foobarbecue · · focus · HN ↗
dofm · · focus · HN ↗
peey · · focus · HN ↗
If 9 is "9 times greater" than 1 then 1 must be "9 times smaller" than 9
It'd help if you read "9 times" with the operator which is what's being flipped instead of with the number
a3w · · focus · HN ↗
50 percent smaller means two-thirds the size. 0.999 times smaller means size is nearly zero. 9 times smaller means you get negative memory from loading it.
derefr · · focus · HN ↗
Kinrany · · focus · HN ↗
nicbor · · focus · HN ↗
I don't see what makes it hard to understand.
stkdump · · focus · HN ↗
so-cal-schemer · · focus · HN ↗
200% as large as 1 is 2.
200% larger than 1 is 3.
so-cal-schemer · · focus · HN ↗
stkdump · · focus · HN ↗
So 200% smaller is 32=6 smaller than 3, which is 3-6=-3
Which is why these kinds of statements usually don't make any sense at all. Usually something can't be more than 100% smaller, at which point it is completely gone.
trentor · · focus · HN ↗
[dead]
a3w · · focus · HN ↗
DoctorOetker · · focus · HN ↗
a∥b=((a*b)/(a+b)
A normal sum has the associative property:
a+(b+c)=(a+b)+c which justifies dropping the parentheses
a+(b+c)=a+b+c=(a+b)+c
It is just an exercise for the reader that the parallel sum ∥ operation is associative:
a∥(b∥c) = (a∥b)∥c
multiplication is typically defined as repeatedly adding:
a x b = b+b+...+b+b (a times b)
similarily one can define parallel multiplication xx :
a xx b = b∥b∥...∥b∥b (a parallel-times b)
The reader can verify that 9 parallel-times size
9 xx size = size∥size∥size∥size∥size∥size∥size∥size∥size = size / 9
So I just think the Bonsai designers meant it was 9 parallel-times smaller, which checks out...
2001zhaozhao · · focus · HN ↗
kennywinker · · focus · HN ↗
jokethrowaway · · focus · HN ↗
kennywinker · · focus · HN ↗
danbrooks · · focus · HN ↗
0xbadcafebee · · focus · HN ↗
nulld3v · · focus · HN ↗
kadoban · · focus · HN ↗
anana_ · · focus · HN ↗
SillyUsername · · focus · HN ↗
I wonder if there's a way to mitigate this by running it through an original Q8 draft model, attuned somehow for the PTQ1 quant, but giving it a higher threshold for the acceptance linear with the context length itself?
The longer the context, the higher the multiplier on the threshold, and more likely the draft result is used. Not ideal but it may extend the usable max context.
This model might, even without this, be amazing for short lived agents that work via generations / have changing tasks.
[deleted] · · focus · HN ↗
[deleted]
WithinReason · · focus · HN ↗
<a href="https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF" rel="nofollow">https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
The 3-bit quant is lossless based on benchmarks.
raylad · · focus · HN ↗
The bf16 only misses "snicker-snack" and this quantization becomes confused after the first stanza.
sorenjan · · focus · HN ↗
0xbadcafebee · · focus · HN ↗
sorenjan · · focus · HN ↗
WithinReason · · focus · HN ↗
[dead]
adrian17 · · focus · HN ↗
If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what makes them better?
<a href="https://news.ycombinator.com/item?id=49611128">https://news.ycombinator.com/item?id=49611128
edflsafoiewq · · focus · HN ↗
yowlingcat · · focus · HN ↗
om8 · · focus · HN ↗
edflsafoiewq · · focus · HN ↗
0x457 · · focus · HN ↗
[dead]
smallerize · · focus · HN ↗
nulld3v · · focus · HN ↗
Balinares · · focus · HN ↗
logicallee · · focus · HN ↗
[1] <a href="https://news.ycombinator.com/item?id=49732931">https://news.ycombinator.com/item?id=49732931
Havoc · · focus · HN ↗
jedbrooke · · focus · HN ↗
So far feels smarter than Bonsai 1 27B, it’s slightly larger than the Q1_0 quant. Super exciting stuff :)
flutetornado · · focus · HN ↗
Seems like we don't have a drafter model yet so it could not test with speculative decoding on. ngram speculative decoding did not help too much either - not enough accepted tokens.
Smaller size I suppose does not mean better performance in this case - we maybe limited by Spark's low memory bandwidth.
cmrdporcupine · · focus · HN ↗
flutetornado · · focus · HN ↗
mmastrac · · focus · HN ↗
flutetornado · · focus · HN ↗
ipoole_dev0 · · focus · HN ↗
[dead]
circularfoyers · · focus · HN ↗
bigyabai · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
That would bring it down to the point where it can fit in 128GB on things like the Spark or Strix Halo.
redlimetea · · focus · HN ↗
[dead]
hedora · · focus · HN ↗
Also, perf speedup?
jjcm · · focus · HN ↗
kennywinker · · focus · HN ↗
So for this one, 27B * 1.76 / 8 = 5.94 GB
For speed, far as I can tell it depends if your gpu is memory bandwidth bound or (mostly older gpus) processing bound. If it’s memory bandwidth bound, and your gpu gets 300GB/s, that’s:
300GB/s / 5.98 GB = 50.5t/s.
Realistically it’s probably a bit slower, but that is your theoretical maximum.
Dwedit · · focus · HN ↗
kennywinker · · focus · HN ↗
Dwedit · · focus · HN ↗
g023 · · focus · HN ↗
Fordec · · focus · HN ↗
ncr100 · · focus · HN ↗
blactuary · · focus · HN ↗
nullbio · · focus · HN ↗
redox99 · · focus · HN ↗
respectattentio · · focus · HN ↗
Yet, seems like there is still another year for improvements.
I like local models (but not mainly using them) for offline needs.
jocelyner · · focus · HN ↗
[dead]
avaer · · focus · HN ↗
For example, the Hadamard activation transform used here feels a lot like multiplying Fourier basis ala DFT; strong parallels to how image codecs work to make the residuals more compressible (especially discrete block codecs like are used in GPU compressed textures).
I thought I was being clever suggesting that you could even abuse texture decode units to efficiently sample compressed LLMs with hardware; turns out Apple foundation models are already doing this [1].
[1] <a href="https://arxiv.org/abs/2507.13575" rel="nofollow">https://arxiv.org/abs/2507.13575
sb057 · · focus · HN ↗
[dead]
Dwedit · · focus · HN ↗
huseyinkeles · · focus · HN ↗
~100t/s prefill, ~15t/s, dropping to ~10t/s later with 64k context.
The issue is I have yet to find a useful agentic local llm that I can run on this machine.
Just given a relatively simple task on a swift app, took 25 minutes, brainstorming like crazy but can not decide on what to do. Eventually I killed it. GPT 5.6 sol-medium took 3 minutes to complete the same task for reference.
aetherspawn · · focus · HN ↗
huseyinkeles · · focus · HN ↗
`cd ~/Code/Bonsai-demo && BONSAI_CTX=65536 ./scripts/start_llama_server.sh`
then used it in a very minimalistic pi with a very small system prompt.
Didn't spend much time to try to optimize it tbh, but my issue was not the speed. it just could not make a decision on how to implement the task, kept going on an on.
yearolinuxdsktp · · focus · HN ↗
aetherspawn · · focus · HN ↗
piyh · · focus · HN ↗
BoredomIsFun · · focus · HN ↗
sean_pedersen · · focus · HN ↗
huseyinkeles · · focus · HN ↗
zhiyan · · focus · HN ↗
throwaway_5753 · · focus · HN ↗
hilti · · focus · HN ↗
[dead]
nilsherzig · · focus · HN ↗
PTQ1_0 has no optimized MMQ-Path in their llama-cpp fork, try running PTQ2_0 (needs a bit more vram, but is about 2x faster on my 6700 XT)
jakswa · · focus · HN ↗
verytrivial · · focus · HN ↗
Aurornis · · focus · HN ↗
I agree. I don’t know how they get such good results on these benchmarks because using them gives a very different experience. They’re kind of cool for doing short free form outputs in memory constrained systems, but I don’t think they’re useful as coding agents.
UrineSqueegee · · focus · HN ↗
qingcharles · · focus · HN ↗
mpweiher · · focus · HN ↗
If true, that would be a very welcome development.
hvhvubufyvycjcx · · focus · HN ↗
hvhvubufyvycjcx · · focus · HN ↗
hvhvubufyvycjcx · · focus · HN ↗
Chance-Device · · focus · HN ↗
Which is within reach of some higher end consumer hardware, especially with layer offloading.
You have to wonder what kind of trouble the “labs” are in when this is becoming possible. Lots of money, where’s the moat?
indy · · focus · HN ↗
DoctorOetker · · focus · HN ↗
It would be a lip service agreement, and all would continue the machine learning race...
The mere suggestion of slowing down AI progress signals to adversaries that there is some novel power just identified: do you think adversaries would agree and slow down, or calculate a little harder and longer in order to also figure out what it unlocks?
indy · · focus · HN ↗
kllrnohj · · focus · HN ↗
Chance-Device · · focus · HN ↗
What’s the moat for trillion dollar AI companies? Access-anywhere convenience to models as good as everyone else’s?
kllrnohj · · focus · HN ↗
ctolsen · · focus · HN ↗
claud_ia · · focus · HN ↗
[dead]
nullbio · · focus · HN ↗
cregy · · focus · HN ↗
Assume bonsai has the same ticks
thway15269037 · · focus · HN ↗
If they did 60gb -> 6gb to Qwen-27B, could they possibly do the same to the MoE model. 36gb lossless 126B model seems like an impossible task.
v3ss0n · · focus · HN ↗
antonly · · focus · HN ↗
petrenk0n · · focus · HN ↗
itsmeduncan · · focus · HN ↗
[dead]
euroderf · · focus · HN ↗
Aurornis · · focus · HN ↗
jakswa · · focus · HN ↗
- 89 tokens/sec generation with speculative decoding - 81 tokens/sec at 20k context - 474 tokens/sec ingestion at 20k — about 42 seconds - 10.1 GiB peak VRAM with a 24k context window
ROCm 7.2.3 · PQ2_0 · Qwen Q4 MTP, draft length 2
---- versus ----
Qwen3.8-27B IQ3_S · Radeon RX 7900 XTX
- 79 tokens/sec generation on a short coding prompt - 61 tokens/sec at 60k context - 53 tokens/sec at 95k context - 558 tokens/sec ingestion at 60k — about 108 seconds - 19.9 GiB peak VRAM during coding tests with a 100k context window
Vulkan · GSQ-RCO IQ3_S · MTP, draft length 2 · vision projector loaded
imagetic · · focus · HN ↗