‹ BackHN Continuity

Thread

Pi 1.0

1684 points · 602 comments · sergiotapia

  1. FacelessJim · · focus · HN ↗
    Love pi. I tried to run some local models and pi was the only one that actually worked decently because it didn’t have a gargantuan system prompt that would take minutes to prefill on my scrawny ass laptop.

    Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.

    Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.

    1. simpaticoder · · focus · HN ↗
      You inspired me to try Pi out - so far it's worked flawlessly. Plugged it into OpenRouter and ~$.50 of Deepseek later I've installed llama.cpp and Llama 3.1. The local model doesn't work with Pi yet (and I know it will be bad and slow even if it does) but I'm curious to see what you can do on an 8GB consumer GPU these days...
      1. sejje · · focus · HN ↗
        > I'm curious to see what you can do on an 8GB consumer GPU these days.

        Nothing, really. Might be coming soon, but no.

        You probably want to try bonsai, I guess, but don't expect good results.

        1. aktenlage · · focus · HN ↗
          Not my experience. Limited, but definitely not nothing.
      2. rablackburn · · focus · HN ↗
        > I'm curious to see what you can do on an 8GB consumer GPU these days

        Running smaller 4B-7B models entirely on the GPU VRAM will get you fast inference, but you will need to scope and define the tasks well. eg, using it the model as a classifier and just feeding it from a queue.

        The best performing "agent"-like model to plug into a harness that I have found so far has been Qwen3.6-35B-A3B (mixture of experts) model as I can park most of it in system RAM and CPU, while the VRAM holds the attention/shared weights.

        It's definitely workable as a local AI homelab. But expect homelab levels of tuning/fiddling with it.

        With the improved support for AMD GPUs I'm finally considering getting a modern 16GB card (and maybe a second one in a few years assuming prices come down)

      3. what · · focus · HN ↗
        > ~$.50 of Deepseek later I've installed llama.cpp and Llama 3.1

        You could install this yourself for free? I get $0.50 isn’t all that much, but still?

        1. simpaticoder · · focus · HN ↗
          Sure, but I'm not interested in learning about running cpp, installing CUDA, finding the right URLs for downloading llama weights. It's the best 50 cents I've spent in 20 years.
          1. 8n4vidtmkvmk · · focus · HN ↗
            Even so, I'm surprised it cost that much. I thought deepseek was cheaper.

            But AI for installing tricky opensource software is indeed a good use case. I do that too.

        2. AgentMasterRace · · focus · HN ↗
          you're living in the 2020s bro
        3. Kurtz79 · · focus · HN ↗
          Except it's not "free", you are using a fraction of your time, arguably your most precious finite resource.

          Even if you spend just 10 minutes of it, I would say $0.50 it's not a bad deal.

          1. what · · focus · HN ↗
            >time is money

            One of the dumbest sayings ever. Unless you spend all of your time doing something that makes money, the time is worth $0. You could say that you prefer to do something else during that time and would happily pay to free it up.

      4. whatshisface · · focus · HN ↗
        If you paid DeepSeek directly, that would have been 1 to 10 cents. OpenRouter has a huge overhead due to their cache logic, I'm surprised they keep business coming in the door for tasks other than system prompt - output pairs.
        1. simpaticoder · · focus · HN ↗
          I think it was actually less than that. I was doing something else too in another agent.
        2. lemontheme · · focus · HN ↗
          I thought openrouter just routes you to the same provider for the rest of the session, so that you keep hitting the same cache. Is that not the case?

          Also, I’d love to use Deepseek directly (or any of the Chinese providers, at that). Seems only fair to pay the lab that built the model. Unfortunately, any requests to Chinese servers is deeply frowned upon here (Belgium, EU). For personal use: sure. As a token intelligence strategy for the company: absolutely fucking not.

          1. miek · · focus · HN ↗

            [dead]

          2. gigatexal · · focus · HN ↗
            It does. They claim that anyway to just route you to the api endpoints for whatever you choose.
          3. RussianCow · · focus · HN ↗
            [delayed]
      5. cellularmitosis · · focus · HN ↗
        A YouTuber by the name of Codacus has been pushing the envelope in this area.
      6. aktenlage · · focus · HN ↗
        I'd chime in with @rablackburn: mixture of experts is the way to go. I have a laptop with 6GB VRAM and I'm running KDE with a 4k display on the same machine, so there's only about 4 to 4.5GB actually available.

        Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.

        It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.

        The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.

        1. GCUMstlyHarmls · · focus · HN ↗
          > Qwen3.6-35B-A3B or gemma4-26B-A4B

          What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?

          Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?

          1. aktenlage · · focus · HN ↗
            I am using the unsloth 4bit quants for both, with quantization aware training for gemma. I haven't tried other quants with these models. I also use a q4 quantized KV cache.

            The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there's no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.

            1. GCUMstlyHarmls · · focus · HN ↗
              Thanks that's helpful.
          2. aktenlage · · focus · HN ↗
            > Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc?

            Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think I sloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if we'll done) the least important information. But with every step you gain less and need to shave of more important things.

            > And also be effected by who did the quantization?

            My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.

      7. Otterly99 · · focus · HN ↗
        With 8GB I would recommend quantized versions of 9B models such as these:

        - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Qwen3.5-9B-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Qwen3.5-9B-GGUF - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;empero-ai&#x2F;Qwen3.8-9B-Distill-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;empero-ai&#x2F;Qwen3.8-9B-Distill-GGUF (unofficial Qwen 3.8-9B) - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;ornith-ai&#x2F;Ornith-1.5-9B-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;ornith-ai&#x2F;Ornith-1.5-9B-GGUF (my personnal favorite)

        1. Godsend69 · · focus · HN ↗

          [dead]

      8. crossroadsguy · · focus · HN ↗
        If that&#x27;s 8GB is available dedicatedly for the model then a lot but if it&#x27;s the sad story like my M1 Pro where even wtih 16GB unififed I&#x27;ve barely anything left for myself.

        You should go to huggingface and maybe create an a&#x2F;c with a throwaway email and enter your hardware details and that will filter the models for you.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.