‹ BackHN Continuity

Thread

Pi 1.0

1684 points · 602 comments · sergiotapia

  1. FacelessJim · · focus · HN ↗
    Love pi. I tried to run some local models and pi was the only one that actually worked decently because it didn’t have a gargantuan system prompt that would take minutes to prefill on my scrawny ass laptop.

    Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.

    Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.

    1. simpaticoder · · focus · HN ↗
      You inspired me to try Pi out - so far it's worked flawlessly. Plugged it into OpenRouter and ~$.50 of Deepseek later I've installed llama.cpp and Llama 3.1. The local model doesn't work with Pi yet (and I know it will be bad and slow even if it does) but I'm curious to see what you can do on an 8GB consumer GPU these days...
      1. aktenlage · · focus · HN ↗
        I'd chime in with @rablackburn: mixture of experts is the way to go. I have a laptop with 6GB VRAM and I'm running KDE with a 4k display on the same machine, so there's only about 4 to 4.5GB actually available.

        Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.

        It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.

        The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.

        1. GCUMstlyHarmls · · focus · HN ↗
          > Qwen3.6-35B-A3B or gemma4-26B-A4B

          What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?

          Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?

          1. aktenlage · · focus · HN ↗
            I am using the unsloth 4bit quants for both, with quantization aware training for gemma. I haven't tried other quants with these models. I also use a q4 quantized KV cache.

            The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there's no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.

            1. GCUMstlyHarmls · · focus · HN ↗
              Thanks that's helpful.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.