‹ BackHN Continuity

Thread

Pi 1.0

1684 points · 602 comments · sergiotapia

  1. FacelessJim · · focus · HN ↗
    Love pi. I tried to run some local models and pi was the only one that actually worked decently because it didn’t have a gargantuan system prompt that would take minutes to prefill on my scrawny ass laptop.

    Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills.

    Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.

    1. rsync · · focus · HN ↗
      "Been running it almost barebones vanilla for a couple of months. Just a bunch of basic extensions and some skills."

      I am also using pi exclusively after having had decent success with openhands but begrudging all of the docker infrastructure ... and all of the emojis.

      My only pain point is that in my extremely common and boring workflow, which is pi inside of gnu screen inside of OSX terminal.app ... all reasoning/thinking text is blinking ... like old fashioned ANSI blink on a BBS.

      I cannot figure out how to disable the blinking thought/reasoning text ...

      1. cbsks · · focus · HN ↗
        I just fixed something very similar in my setup. Except in my case the reasoning text was shown in dark grey on a light gray background. Very ugly and hard to read.

        If I remember correctly, the reasoning text was being output using the italic ANSI code, which was being formatted funny on my terminal. I fixed it by adding a font that supports italics. I recommend taking a look at the ansi codes.

        1. rsync · · focus · HN ↗
          Thanks.

          This seems like an obvious configuration option - I can imagine someone disliking the italics as well…

          1. nine_k · · focus · HN ↗
            If only we had a way to tell the machine to locate and fix this problem to our liking...
            1. stpedgwdgfhgdd · · focus · HN ↗
              Good joke, but I wonder how many people get it based on the reactions below.

              Perhaps Pi should ask after x days of installation; is there anything I can do to make the interaction better?

              1. lionkor · · focus · HN ↗
                For a second I was hoping that Pi users would be the kind of people to not enjoy nags about improvements.
              2. jasonjayr · · focus · HN ↗
                > Pi can explain its own features and look up its docs. Ask it how to use or extend Pi.

                That line is in the startup message everytime...

      2. throwawayblahbl · · focus · HN ↗

        [dead]

      3. sothatsit · · focus · HN ↗
        I was running into similar issues where italics text was blinking. I traced it back to a bug in screen, which I patched in my own screen fork. Not sure if exactly the same bug but could be? <a href="https:&#x2F;&#x2F;github.com&#x2F;Sothatsit&#x2F;screen" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Sothatsit&#x2F;screen
      4. rpdillon · · focus · HN ↗
        See if you can replicate this inside of tmux. It might be screen&#x27;s escape code handling.
      5. rmunn · · focus · HN ↗
        Have you tried a different terminal app, such as Ghostty? <a href="https:&#x2F;&#x2F;ghostty.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;ghostty.org&#x2F; has a Mac build, and handles italics properly. That might solve your issue without having to edit any configuration files. Plus, as a side benefit, Ghostty ignores the ANSI color codes for blinking text, so you won&#x27;t ever see blinking text again.
        1. funcDropShadow · · focus · HN ↗
          I am wondering why people are so hyped about Ghostty? I gave it recently a try coming from kitty. And I had to configure stuff that worked out of the box with kitty, like Ctrl-Enter support and other key combos for agent harnesses. I like the development model of ghostt. And the developer really cares about creating great building blocks, e.g libghostty. But I am a bit underwhelmed.
          1. alxhslm · · focus · HN ↗
            Might depend on what you&#x27;re comparing to.

            Came from iTerm2, and Ghostty is much faster and minimal.

          2. computershit · · focus · HN ↗
            I feel like the ghostty hype make more sense as excitement around the direction of terminal infra than a claim it is the one true terminal all others suck. Kitty&#x27;s fine.

            Tangentially I also kind of feel like there is some level of deification of Mitchell&#x27;s products but that&#x27;s a diff topic.

      6. nine_k · · focus · HN ↗
        If you love your terminal app and won&#x27;t switch to Ghostty or WezTerm, at least try tmux instead of the venerable but ancient GNU screen.
        1. huijzer · · focus · HN ↗
          Alacritty plus Zellij works great for me. Much easier to use than tmux
          1. dvergeylen · · focus · HN ↗
            Didn&#x27;t know about Zellij, seems very good, thank you!
      7. kalleboo · · focus · HN ↗
        Terminal.app Settings&#x2F;Profiles has a checkbox &quot;Allow blinking text&quot; you can uncheck
      8. jonwinstanley · · focus · HN ↗
        Afaik iTerm2 is the usual upgrade for the macOS terminal isn’t it?
        1. Geezus_42 · · focus · HN ↗
          It&#x27;s the goto, but there are better options IMO.
      9. lionkor · · focus · HN ↗
        Apart from a different terminal app, try `&#x2F;settings`, find the &quot;TUI mode&quot; or whatever, and set it to fullscreen. It fixed all my flickering and other issues.
    2. RickS · · focus · HN ↗
      Couldn&#x27;t agree more. The vanilla openclaw install was this byzantine mess of MD files talking about souls and identities and such, it really put me off. Stripping back to a bare install of the underlying pi, it was delightfully minimal and easy to reason about. Excellent starting point for building an assistant agent without having to read or fight with a bunch of cruft on top.

      Couple skills to integrate with an obsidian MD task tracker, small chat interface on the phone made public via tailscale, and bam, a reminder bot you can text from the grocery store.

      1. AgentMasterRace · · focus · HN ↗
        you&#x27;re comparing apples to oranges here...
        1. vergessenmir · · focus · HN ↗
          They&#x27;re both harnesses, pi + cron&#x2F;event trigger is 90% openclaw
          1. davedx · · focus · HN ↗
            Don&#x27;t forget that all important whatsapp gateway
            1. vorticalbox · · focus · HN ↗
              There is a docker image that has an api to interact with WhatsApp and you can use the signal cli.

              This is what I did and then just wrote small skills so now pi can read and send messages for me.

        2. Medea · · focus · HN ↗
          More like comparing apple pie to apples. Turns out I just wanted the apples.
    3. berofeev · · focus · HN ↗
      You&#x27;re looking for fullscreen TUI mode: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;PiCodingAgent&#x2F;comments&#x2F;1vh5pys&#x2F;thank_you_for_full_screen_tui_mode&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;PiCodingAgent&#x2F;comments&#x2F;1vh5pys&#x2F;than...
    4. wilt_ · · focus · HN ↗
      Are you using Windows Terminal by any chance? I&#x27;m building a personal fork [0] with a patch for this exact bug (plus a few other open PRs from the upstream repo that seemed cool). Haven&#x27;t tried contributing it upstream, since the patch is fully vibe-coded and I&#x27;ve spent almost no time trying to understand how it works, but the bug hasn&#x27;t recurred since I&#x27;ve been using it.

      [0] <a href="https:&#x2F;&#x2F;github.com&#x2F;wilt00&#x2F;windows-terminal&#x2F;releases" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;wilt00&#x2F;windows-terminal&#x2F;releases

    5. simpaticoder · · focus · HN ↗
      You inspired me to try Pi out - so far it&#x27;s worked flawlessly. Plugged it into OpenRouter and ~$.50 of Deepseek later I&#x27;ve installed llama.cpp and Llama 3.1. The local model doesn&#x27;t work with Pi yet (and I know it will be bad and slow even if it does) but I&#x27;m curious to see what you can do on an 8GB consumer GPU these days...
      1. sejje · · focus · HN ↗
        &gt; I&#x27;m curious to see what you can do on an 8GB consumer GPU these days.

        Nothing, really. Might be coming soon, but no.

        You probably want to try bonsai, I guess, but don&#x27;t expect good results.

        1. aktenlage · · focus · HN ↗
          Not my experience. Limited, but definitely not nothing.
      2. rablackburn · · focus · HN ↗
        &gt; I&#x27;m curious to see what you can do on an 8GB consumer GPU these days

        Running smaller 4B-7B models entirely on the GPU VRAM will get you fast inference, but you will need to scope and define the tasks well. eg, using it the model as a classifier and just feeding it from a queue.

        The best performing &quot;agent&quot;-like model to plug into a harness that I have found so far has been Qwen3.6-35B-A3B (mixture of experts) model as I can park most of it in system RAM and CPU, while the VRAM holds the attention&#x2F;shared weights.

        It&#x27;s definitely workable as a local AI homelab. But expect homelab levels of tuning&#x2F;fiddling with it.

        With the improved support for AMD GPUs I&#x27;m finally considering getting a modern 16GB card (and maybe a second one in a few years assuming prices come down)

      3. what · · focus · HN ↗
        &gt; ~$.50 of Deepseek later I&#x27;ve installed llama.cpp and Llama 3.1

        You could install this yourself for free? I get $0.50 isn’t all that much, but still?

        1. simpaticoder · · focus · HN ↗
          Sure, but I&#x27;m not interested in learning about running cpp, installing CUDA, finding the right URLs for downloading llama weights. It&#x27;s the best 50 cents I&#x27;ve spent in 20 years.
          1. 8n4vidtmkvmk · · focus · HN ↗
            Even so, I&#x27;m surprised it cost that much. I thought deepseek was cheaper.

            But AI for installing tricky opensource software is indeed a good use case. I do that too.

        2. AgentMasterRace · · focus · HN ↗
          you&#x27;re living in the 2020s bro
        3. Kurtz79 · · focus · HN ↗
          Except it&#x27;s not &quot;free&quot;, you are using a fraction of your time, arguably your most precious finite resource.

          Even if you spend just 10 minutes of it, I would say $0.50 it&#x27;s not a bad deal.

          1. what · · focus · HN ↗
            &gt;time is money

            One of the dumbest sayings ever. Unless you spend all of your time doing something that makes money, the time is worth $0. You could say that you prefer to do something else during that time and would happily pay to free it up.

      4. whatshisface · · focus · HN ↗
        If you paid DeepSeek directly, that would have been 1 to 10 cents. OpenRouter has a huge overhead due to their cache logic, I&#x27;m surprised they keep business coming in the door for tasks other than system prompt - output pairs.
        1. simpaticoder · · focus · HN ↗
          I think it was actually less than that. I was doing something else too in another agent.
        2. lemontheme · · focus · HN ↗
          I thought openrouter just routes you to the same provider for the rest of the session, so that you keep hitting the same cache. Is that not the case?

          Also, I’d love to use Deepseek directly (or any of the Chinese providers, at that). Seems only fair to pay the lab that built the model. Unfortunately, any requests to Chinese servers is deeply frowned upon here (Belgium, EU). For personal use: sure. As a token intelligence strategy for the company: absolutely fucking not.

          1. miek · · focus · HN ↗

            [dead]

          2. gigatexal · · focus · HN ↗
            It does. They claim that anyway to just route you to the api endpoints for whatever you choose.
          3. RussianCow · · focus · HN ↗
            [delayed]
      5. cellularmitosis · · focus · HN ↗
        A YouTuber by the name of Codacus has been pushing the envelope in this area.
      6. aktenlage · · focus · HN ↗
        I&#x27;d chime in with @rablackburn: mixture of experts is the way to go. I have a laptop with 6GB VRAM and I&#x27;m running KDE with a 4k display on the same machine, so there&#x27;s only about 4 to 4.5GB actually available.

        Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens&#x2F;s on prompt processing and 20-40 t&#x2F;s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.

        It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.

        The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function&#x2F;file, etc), but not good enough for complex things.

        1. GCUMstlyHarmls · · focus · HN ↗
          &gt; Qwen3.6-35B-A3B or gemma4-26B-A4B

          What quantization are you running for these? Like, you cant just run the &quot;real&quot; ones on your laptop right?

          Wouldn&#x27;t the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?

          1. aktenlage · · focus · HN ↗
            I am using the unsloth 4bit quants for both, with quantization aware training for gemma. I haven&#x27;t tried other quants with these models. I also use a q4 quantized KV cache.

            The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there&#x27;s no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.

            1. GCUMstlyHarmls · · focus · HN ↗
              Thanks that&#x27;s helpful.
          2. aktenlage · · focus · HN ↗
            &gt; Wouldn&#x27;t the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc?

            Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think I sloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it&#x27;s about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if we&#x27;ll done) the least important information. But with every step you gain less and need to shave of more important things.

            &gt; And also be effected by who did the quantization?

            My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.

      7. Otterly99 · · focus · HN ↗
        With 8GB I would recommend quantized versions of 9B models such as these:

        - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Qwen3.5-9B-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;unsloth&#x2F;Qwen3.5-9B-GGUF - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;empero-ai&#x2F;Qwen3.8-9B-Distill-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;empero-ai&#x2F;Qwen3.8-9B-Distill-GGUF (unofficial Qwen 3.8-9B) - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;ornith-ai&#x2F;Ornith-1.5-9B-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;ornith-ai&#x2F;Ornith-1.5-9B-GGUF (my personnal favorite)

        1. Godsend69 · · focus · HN ↗

          [dead]

      8. crossroadsguy · · focus · HN ↗
        If that&#x27;s 8GB is available dedicatedly for the model then a lot but if it&#x27;s the sad story like my M1 Pro where even wtih 16GB unififed I&#x27;ve barely anything left for myself.

        You should go to huggingface and maybe create an a&#x2F;c with a throwaway email and enter your hardware details and that will filter the models for you.

    6. gchamonlive · · focus · HN ↗
      I use oh-my-pi, not sure how it compares, but I say someone praising antigravity for being a good harness(1), and for the love of good people settle for such low standards of user experience it&#x27;s almost pitiful.

      (1) <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49913854">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49913854

      1. gigatexal · · focus · HN ↗
        Yup! I am using the same having heard about it here.
      2. andriy_koval · · focus · HN ↗
        It maybe depend on workflow, you hit some corner cases for agy which I don&#x27;t hit.

        I use agy and codex, and don&#x27;t see strong difference in ui quality.

        1. gchamonlive · · focus · HN ↗
          Pardon me, but authorizing ls, grep, ps, find... doesn&#x27;t seem like corner cases to me. And codex and agy are close to the kind of mainstream audience that they aim for.

          But I agree, it&#x27;s definitely a matter of workflow, because I can see omp doing untold damage in the hands of the uninitiated.

    7. ziphyrien · · focus · HN ↗
      This bug was fixed several months ago; you just need to switch to full-screen mode, though you hadn’t done so previously.

      However, they have now set full-screen mode as the default.

    8. syrusakbary · · focus · HN ↗
      Same! Pi is incredibly exciting.

      We launched Pi support in Wasmer a few days ago and reception has been great (so you can run pi in your iPhone or browser, or even embedded)

      We have set up this demo, if you want to try Pi 1.0 online: <a href="https:&#x2F;&#x2F;wasmer.sh&#x2F;?example=pi" rel="nofollow">https:&#x2F;&#x2F;wasmer.sh&#x2F;?example=pi

      (for an easter egg click on the Pi logo on the top left!)

    9. dmarchand90 · · focus · HN ↗
      You can use full screen mode. (I asked my pi agent about this and it told me about it) That fixed the flicker issue for me

      Oh and that will be the new default &#x27;Full-screen mode by default&#x27;

    10. miroljub · · focus · HN ↗
      That bug is inherent to how terminal scrollback works and can’t be fixed as long as you use it. Your only choice is that jump or stale backscroll. Or you use Pi new fullscreen mode which gives up on terminal and use alternate screens and implement own scrolling without terminal scrollback.

      How do I know? I implemented my own terminal harness and faced the same issue.

    11. kelnos · · focus · HN ↗
      &gt; Now, if only they could fix the very annoying bug of the history jumping back at the beginning if I am not a the end while the model is reasoning that would great.

      It&#x27;s so interesting because Claude Code used to have this bug a long time ago, but it was eventually fixed. Strange that they both had&#x2F;have the same issue.

      1. smokel · · focus · HN ↗
        Why do these tools have such bugs? With AI it should be trivial to fix, no?

        Or is it a non-trivial bug that requires a lot of refactoring, and could introduce a lot of new bugs? That would require careful review from a human.

        The latter may well be a reason why I don&#x27;t see extreme productivity gains in larger brown-field projects.

        (Disclaimer: I see enormous benefits in one-off greenfield projects.)

        1. kekebo · · focus · HN ↗
          I don&#x27;t think it&#x27;s necessarily a bug per se, but a central tradeoff in system prompt length between well-documenting the environment (harness specifics, exposed tools, tool use instructions etc) to the llm, vs the initial prompt stage (&quot;prefill&quot;) growing so large that it results in an unpleasant lag to first response, and reduced available context, which is most noticeable with open models on resource-constrained consumer hardware.

          You can use llm to optimize some of this, I condensed the tool descriptions of some larger LM Studio plugins to shrink prefill by almost 10k tokens. But there&#x27;s a soft limit to this, if you don&#x27;t want to under-document available tools and let the model guess (&#x2F;behave unsafely).

          One optimization around this is called &quot;smart tool selection&quot;, which only sends tool descriptions when the model indicates need for a certain tool (suite), not all of them upfront.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.