Love pi. I tried to run some local models and pi was the only one that actually worked decently because it didn’t have a gargantuan system prompt that would take minutes to prefill on my scrawny ass laptop.
Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.
Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.
You inspired me to try Pi out - so far it's worked flawlessly. Plugged it into OpenRouter and ~$.50 of Deepseek later I've installed llama.cpp and Llama 3.1. The local model doesn't work with Pi yet (and I know it will be bad and slow even if it does) but I'm curious to see what you can do on an 8GB consumer GPU these days...
I'd chime in with @rablackburn: mixture of experts is the way to go. I have a laptop with 6GB VRAM and I'm running KDE with a 4k display on the same machine, so there's only about 4 to 4.5GB actually available.
Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.
It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.
The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.
What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?
Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?
> Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc?
Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think I sloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if we'll done) the least important information. But with every step you gain less and need to shave of more important things.
> And also be effected by who did the quantization?
My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.
FacelessJim · · focus · HN ↗
Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.
Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.
simpaticoder · · focus · HN ↗
aktenlage · · focus · HN ↗
Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.
It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.
The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.
GCUMstlyHarmls · · focus · HN ↗
What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?
Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?
aktenlage · · focus · HN ↗
Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think I sloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if we'll done) the least important information. But with every step you gain less and need to shave of more important things.
> And also be effected by who did the quantization?
My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.