Love pi. I tried to run some local models and pi was the only one that actually worked decently because it didn’t have a gargantuan system prompt that would take minutes to prefill on my scrawny ass laptop.
Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.
Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.
"Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills."
I am also using pi exclusively after having had decent success with openhands but begrudging all of the docker infrastructure ... and all of the emojis.
My only pain point is that in my extremely common and boring workflow, which is pi inside of gnu screen inside of OSX terminal.app ... all reasoning/thinking text is blinking ... like old fashioned ANSI blink on a BBS.
I cannot figure out how to disable the blinking thought/reasoning text ...
I just fixed something very similar in my setup. Except in my case the reasoning text was shown in dark grey on a light gray background. Very ugly and hard to read.
If I remember correctly, the reasoning text was being output using the italic ANSI code, which was being formatted funny on my terminal. I fixed it by adding a font that supports italics. I recommend taking a look at the ansi codes.
I was running into similar issues where italics text was blinking. I traced it back to a bug in screen, which I patched in my own screen fork. Not sure if exactly the same bug but could be? <a href="https://github.com/Sothatsit/screen" rel="nofollow">https://github.com/Sothatsit/screen
Have you tried a different terminal app, such as Ghostty? <a href="https://ghostty.org/" rel="nofollow">https://ghostty.org/ has a Mac build, and handles italics properly. That might solve your issue without having to edit any configuration files. Plus, as a side benefit, Ghostty ignores the ANSI color codes for blinking text, so you won't ever see blinking text again.
I am wondering why people are so hyped about Ghostty? I gave it recently a try coming from kitty. And I had to configure stuff that worked out of the box with kitty, like Ctrl-Enter support and other key combos for agent harnesses. I like the development model of ghostt. And the developer really cares about creating great building blocks, e.g libghostty. But I am a bit underwhelmed.
I feel like the ghostty hype make more sense as excitement around the direction of terminal infra than a claim it is the one true terminal all others suck. Kitty's fine.
Tangentially I also kind of feel like there is some level of deification of Mitchell's products but that's a diff topic.
Apart from a different terminal app, try `/settings`, find the "TUI mode" or whatever, and set it to fullscreen. It fixed all my flickering and other issues.
Couldn't agree more. The vanilla openclaw install was this byzantine mess of MD files talking about souls and identities and such, it really put me off. Stripping back to a bare install of the underlying pi, it was delightfully minimal and easy to reason about. Excellent starting point for building an assistant agent without having to read or fight with a bunch of cruft on top.
Couple skills to integrate with an obsidian MD task tracker, small chat interface on the phone made public via tailscale, and bam, a reminder bot you can text from the grocery store.
Are you using Windows Terminal by any chance? I'm building a personal fork [0] with a patch for this exact bug (plus a few other open PRs from the upstream repo that seemed cool). Haven't tried contributing it upstream, since the patch is fully vibe-coded and I've spent almost no time trying to understand how it works, but the bug hasn't recurred since I've been using it.
You inspired me to try Pi out - so far it's worked flawlessly. Plugged it into OpenRouter and ~$.50 of Deepseek later I've installed llama.cpp and Llama 3.1. The local model doesn't work with Pi yet (and I know it will be bad and slow even if it does) but I'm curious to see what you can do on an 8GB consumer GPU these days...
> I'm curious to see what you can do on an 8GB consumer GPU these days
Running smaller 4B-7B models entirely on the GPU VRAM will get you fast inference, but you will need to scope and define the tasks well. eg, using it the model as a classifier and just feeding it from a queue.
The best performing "agent"-like model to plug into a harness that I have found so far has been Qwen3.6-35B-A3B (mixture of experts) model as I can park most of it in system RAM and CPU, while the VRAM holds the attention/shared weights.
It's definitely workable as a local AI homelab. But expect homelab levels of tuning/fiddling with it.
With the improved support for AMD GPUs I'm finally considering getting a modern 16GB card (and maybe a second one in a few years assuming prices come down)
Sure, but I'm not interested in learning about running cpp, installing CUDA, finding the right URLs for downloading llama weights. It's the best 50 cents I've spent in 20 years.
One of the dumbest sayings ever. Unless you spend all of your time doing something that makes money, the time is worth $0. You could say that you prefer to do something else during that time and would happily pay to free it up.
If you paid DeepSeek directly, that would have been 1 to 10 cents. OpenRouter has a huge overhead due to their cache logic, I'm surprised they keep business coming in the door for tasks other than system prompt - output pairs.
I thought openrouter just routes you to the same provider for the rest of the session, so that you keep hitting the same cache. Is that not the case?
Also, I’d love to use Deepseek directly (or any of the Chinese providers, at that). Seems only fair to pay the lab that built the model. Unfortunately, any requests to Chinese servers is deeply frowned upon here (Belgium, EU). For personal use: sure. As a token intelligence strategy for the company: absolutely fucking not.
I'd chime in with @rablackburn: mixture of experts is the way to go. I have a laptop with 6GB VRAM and I'm running KDE with a 4k display on the same machine, so there's only about 4 to 4.5GB actually available.
Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.
It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.
The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.
What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?
Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?
I am using the unsloth 4bit quants for both, with quantization aware training for gemma. I haven't tried other quants with these models. I also use a q4 quantized KV cache.
The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there's no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.
> Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc?
Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think I sloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if we'll done) the least important information. But with every step you gain less and need to shave of more important things.
> And also be effected by who did the quantization?
My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.
If that's 8GB is available dedicatedly for the model then a lot but if it's the sad story like my M1 Pro where even wtih 16GB unififed I've barely anything left for myself.
You should go to huggingface and maybe create an a/c with a throwaway email and enter your hardware details and that will filter the models for you.
I use oh-my-pi, not sure how it compares, but I say someone praising antigravity for being a good harness(1), and for the love of good people settle for such low standards of user experience it's almost pitiful.
Pardon me, but authorizing ls, grep, ps, find... doesn't seem like corner cases to me. And codex and agy are close to the kind of mainstream audience that they aim for.
But I agree, it's definitely a matter of workflow, because I can see omp doing untold damage in the hands of the uninitiated.
We launched Pi support in Wasmer a few days ago and reception has been great (so you can run pi in your iPhone or browser, or even embedded)
We have set up this demo, if you want to try Pi 1.0 online: <a href="https://wasmer.sh/?example=pi" rel="nofollow">https://wasmer.sh/?example=pi
(for an easter egg click on the Pi logo on the top left!)
That bug is inherent to how terminal scrollback works and can’t be fixed as long as you use it. Your only choice is that jump or stale backscroll. Or you use Pi new fullscreen mode which gives up on terminal and use alternate screens and implement own scrolling without terminal scrollback.
How do I know? I implemented my own terminal harness and faced the same issue.
> Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.
It's so interesting because Claude Code used to have this bug a long time ago, but it was eventually fixed. Strange that they both had/have the same issue.
I don't think it's necessarily a bug per se, but a central tradeoff in system prompt length between well-documenting the environment (harness specifics, exposed tools, tool use instructions etc) to the llm, vs the initial prompt stage ("prefill") growing so large that it results in an unpleasant lag to first response, and reduced available context, which is most noticeable with open models on resource-constrained consumer hardware.
You can use llm to optimize some of this, I condensed the tool descriptions of some larger LM Studio plugins to shrink prefill by almost 10k tokens. But there's a soft limit to this, if you don't want to under-document available tools and let the model guess (/behave unsafely).
One optimization around this is called "smart tool selection", which only sends tool descriptions when the model indicates need for a certain tool (suite), not all of them upfront.
FacelessJim · · focus · HN ↗
Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.
Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.
rsync · · focus · HN ↗
I am also using pi exclusively after having had decent success with openhands but begrudging all of the docker infrastructure ... and all of the emojis.
My only pain point is that in my extremely common and boring workflow, which is pi inside of gnu screen inside of OSX terminal.app ... all reasoning/thinking text is blinking ... like old fashioned ANSI blink on a BBS.
I cannot figure out how to disable the blinking thought/reasoning text ...
cbsks · · focus · HN ↗
If I remember correctly, the reasoning text was being output using the italic ANSI code, which was being formatted funny on my terminal. I fixed it by adding a font that supports italics. I recommend taking a look at the ansi codes.
rsync · · focus · HN ↗
This seems like an obvious configuration option - I can imagine someone disliking the italics as well…
nine_k · · focus · HN ↗
stpedgwdgfhgdd · · focus · HN ↗
Perhaps Pi should ask after x days of installation; is there anything I can do to make the interaction better?
lionkor · · focus · HN ↗
jasonjayr · · focus · HN ↗
That line is in the startup message everytime...
throwawayblahbl · · focus · HN ↗
[dead]
sothatsit · · focus · HN ↗
rpdillon · · focus · HN ↗
rmunn · · focus · HN ↗
funcDropShadow · · focus · HN ↗
alxhslm · · focus · HN ↗
Came from iTerm2, and Ghostty is much faster and minimal.
computershit · · focus · HN ↗
Tangentially I also kind of feel like there is some level of deification of Mitchell's products but that's a diff topic.
nine_k · · focus · HN ↗
huijzer · · focus · HN ↗
dvergeylen · · focus · HN ↗
kalleboo · · focus · HN ↗
jonwinstanley · · focus · HN ↗
Geezus_42 · · focus · HN ↗
lionkor · · focus · HN ↗
RickS · · focus · HN ↗
Couple skills to integrate with an obsidian MD task tracker, small chat interface on the phone made public via tailscale, and bam, a reminder bot you can text from the grocery store.
AgentMasterRace · · focus · HN ↗
vergessenmir · · focus · HN ↗
davedx · · focus · HN ↗
vorticalbox · · focus · HN ↗
This is what I did and then just wrote small skills so now pi can read and send messages for me.
Medea · · focus · HN ↗
berofeev · · focus · HN ↗
wilt_ · · focus · HN ↗
[0] <a href="https://github.com/wilt00/windows-terminal/releases" rel="nofollow">https://github.com/wilt00/windows-terminal/releases
simpaticoder · · focus · HN ↗
sejje · · focus · HN ↗
Nothing, really. Might be coming soon, but no.
You probably want to try bonsai, I guess, but don't expect good results.
aktenlage · · focus · HN ↗
rablackburn · · focus · HN ↗
Running smaller 4B-7B models entirely on the GPU VRAM will get you fast inference, but you will need to scope and define the tasks well. eg, using it the model as a classifier and just feeding it from a queue.
The best performing "agent"-like model to plug into a harness that I have found so far has been Qwen3.6-35B-A3B (mixture of experts) model as I can park most of it in system RAM and CPU, while the VRAM holds the attention/shared weights.
It's definitely workable as a local AI homelab. But expect homelab levels of tuning/fiddling with it.
With the improved support for AMD GPUs I'm finally considering getting a modern 16GB card (and maybe a second one in a few years assuming prices come down)
what · · focus · HN ↗
You could install this yourself for free? I get $0.50 isn’t all that much, but still?
simpaticoder · · focus · HN ↗
8n4vidtmkvmk · · focus · HN ↗
But AI for installing tricky opensource software is indeed a good use case. I do that too.
AgentMasterRace · · focus · HN ↗
Kurtz79 · · focus · HN ↗
Even if you spend just 10 minutes of it, I would say $0.50 it's not a bad deal.
what · · focus · HN ↗
One of the dumbest sayings ever. Unless you spend all of your time doing something that makes money, the time is worth $0. You could say that you prefer to do something else during that time and would happily pay to free it up.
whatshisface · · focus · HN ↗
simpaticoder · · focus · HN ↗
lemontheme · · focus · HN ↗
Also, I’d love to use Deepseek directly (or any of the Chinese providers, at that). Seems only fair to pay the lab that built the model. Unfortunately, any requests to Chinese servers is deeply frowned upon here (Belgium, EU). For personal use: sure. As a token intelligence strategy for the company: absolutely fucking not.
miek · · focus · HN ↗
[dead]
gigatexal · · focus · HN ↗
RussianCow · · focus · HN ↗
cellularmitosis · · focus · HN ↗
aktenlage · · focus · HN ↗
Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.
It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.
The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.
GCUMstlyHarmls · · focus · HN ↗
What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?
Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?
aktenlage · · focus · HN ↗
The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there's no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.
GCUMstlyHarmls · · focus · HN ↗
aktenlage · · focus · HN ↗
Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think I sloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if we'll done) the least important information. But with every step you gain less and need to shave of more important things.
> And also be effected by who did the quantization?
My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.
Otterly99 · · focus · HN ↗
- <a href="https://huggingface.co/unsloth/Qwen3.5-9B-GGUF" rel="nofollow">https://huggingface.co/unsloth/Qwen3.5-9B-GGUF - <a href="https://huggingface.co/empero-ai/Qwen3.8-9B-Distill-GGUF" rel="nofollow">https://huggingface.co/empero-ai/Qwen3.8-9B-Distill-GGUF (unofficial Qwen 3.8-9B) - <a href="https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF" rel="nofollow">https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF (my personnal favorite)
Godsend69 · · focus · HN ↗
[dead]
crossroadsguy · · focus · HN ↗
You should go to huggingface and maybe create an a/c with a throwaway email and enter your hardware details and that will filter the models for you.
gchamonlive · · focus · HN ↗
(1) <a href="https://news.ycombinator.com/item?id=49913854">https://news.ycombinator.com/item?id=49913854
gigatexal · · focus · HN ↗
andriy_koval · · focus · HN ↗
I use agy and codex, and don't see strong difference in ui quality.
gchamonlive · · focus · HN ↗
But I agree, it's definitely a matter of workflow, because I can see omp doing untold damage in the hands of the uninitiated.
ziphyrien · · focus · HN ↗
However, they have now set full-screen mode as the default.
syrusakbary · · focus · HN ↗
We launched Pi support in Wasmer a few days ago and reception has been great (so you can run pi in your iPhone or browser, or even embedded)
We have set up this demo, if you want to try Pi 1.0 online: <a href="https://wasmer.sh/?example=pi" rel="nofollow">https://wasmer.sh/?example=pi
(for an easter egg click on the Pi logo on the top left!)
dmarchand90 · · focus · HN ↗
Oh and that will be the new default 'Full-screen mode by default'
miroljub · · focus · HN ↗
How do I know? I implemented my own terminal harness and faced the same issue.
kelnos · · focus · HN ↗
It's so interesting because Claude Code used to have this bug a long time ago, but it was eventually fixed. Strange that they both had/have the same issue.
smokel · · focus · HN ↗
Or is it a non-trivial bug that requires a lot of refactoring, and could introduce a lot of new bugs? That would require careful review from a human.
The latter may well be a reason why I don't see extreme productivity gains in larger brown-field projects.
(Disclaimer: I see enormous benefits in one-off greenfield projects.)
kekebo · · focus · HN ↗
You can use llm to optimize some of this, I condensed the tool descriptions of some larger LM Studio plugins to shrink prefill by almost 10k tokens. But there's a soft limit to this, if you don't want to under-document available tools and let the model guess (/behave unsafely).
One optimization around this is called "smart tool selection", which only sends tool descriptions when the model indicates need for a certain tool (suite), not all of them upfront.