Love pi. I tried to run some local models and pi was the only one that actually worked decently because it didn’t have a gargantuan system prompt that would take minutes to prefill on my scrawny ass laptop.
Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.
Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.
You inspired me to try Pi out - so far it's worked flawlessly. Plugged it into OpenRouter and ~$.50 of Deepseek later I've installed llama.cpp and Llama 3.1. The local model doesn't work with Pi yet (and I know it will be bad and slow even if it does) but I'm curious to see what you can do on an 8GB consumer GPU these days...
> I'm curious to see what you can do on an 8GB consumer GPU these days
Running smaller 4B-7B models entirely on the GPU VRAM will get you fast inference, but you will need to scope and define the tasks well. eg, using it the model as a classifier and just feeding it from a queue.
The best performing "agent"-like model to plug into a harness that I have found so far has been Qwen3.6-35B-A3B (mixture of experts) model as I can park most of it in system RAM and CPU, while the VRAM holds the attention/shared weights.
It's definitely workable as a local AI homelab. But expect homelab levels of tuning/fiddling with it.
With the improved support for AMD GPUs I'm finally considering getting a modern 16GB card (and maybe a second one in a few years assuming prices come down)
FacelessJim · · focus · HN ↗
Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.
Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.
simpaticoder · · focus · HN ↗
rablackburn · · focus · HN ↗
Running smaller 4B-7B models entirely on the GPU VRAM will get you fast inference, but you will need to scope and define the tasks well. eg, using it the model as a classifier and just feeding it from a queue.
The best performing "agent"-like model to plug into a harness that I have found so far has been Qwen3.6-35B-A3B (mixture of experts) model as I can park most of it in system RAM and CPU, while the VRAM holds the attention/shared weights.
It's definitely workable as a local AI homelab. But expect homelab levels of tuning/fiddling with it.
With the improved support for AMD GPUs I'm finally considering getting a modern 16GB card (and maybe a second one in a few years assuming prices come down)