‹ BackHN Continuity

Thread

Pi 1.0

1684 points · 602 comments · sergiotapia

  1. hhh · · focus · HN ↗
    I don't really understand the criteria for when something is 'proven' to the Pi team. Jev and the like took off less than a month ago, but MCP has been growing for nearly 2 years, and it only gets support now?

    Pi felt nice when I used it, and I do value keeping things minimal, but I just find the criteria very uneven.

    1. rsalus · · focus · HN ↗
      the latest 07-28 MCP spec is quite different than the previous iterations of MCP, so I understand the delay there tbh.
    2. BeetleB · · focus · HN ↗
      > but MCP has been growing for nearly 2 years, and it only gets support now?

      I don't know if you were aware, but not shipping with MCP was one of its "features":

      <a href="https:&#x2F;&#x2F;mariozechner.at&#x2F;posts&#x2F;2025-11-02-what-if-you-dont-need-mcp&#x2F;" rel="nofollow">https:&#x2F;&#x2F;mariozechner.at&#x2F;posts&#x2F;2025-11-02-what-if-you-dont-ne...

      They let you have it via a plugin&#x2F;extension.

      1. the_mitsuhiko · · focus · HN ↗
        A year later, some things have changed: <a href="https:&#x2F;&#x2F;earendil.com&#x2F;posts&#x2F;you-said-no-mcp&#x2F;" rel="nofollow">https:&#x2F;&#x2F;earendil.com&#x2F;posts&#x2F;you-said-no-mcp&#x2F;
    3. the_mitsuhiko · · focus · HN ↗
      Armin from Earendil here. I think the question is fair, and quite frankly the answer is pretty disappointing: we look at what the models are doing. They are trained on their respective harnesses and we&#x27;re not here to fight their behavior.

      Codex in particular is using responses lite internally and relies on codemode for parallel tool calling. So codemode was a given.

      Jev on the other hand is new but it&#x27;s not the first type of model we had troubles with supporting in Pi and we looked at how to make that make sense. The internal pi-ai SDK supports image generation and classifier models, but without building an extension it was never possible for you to utilize it.

      So there was a while functionality of Pi that few people used, because there were no obvious ways to hook it up with the coding agent. Codemode also allows us to close that gap.

      And once you have codemode, modern MCP can work quite well if the servers cooperate.

      1. octoberfranklin · · focus · HN ↗
        and we&#x27;re not here to influence their behavior

        That is totally disappointing.

    4. Zambyte · · focus · HN ↗
      Classification models have been around for literally almost a century at this point. I think it&#x27;s safe to say they are a proven technology.

      The only thing that makes Jev and the likes particularly interesting is that it is a general purpose classifier. In the past, classification tasks meant training a new model to solve your problem. Now you can just use an off the shelf general purpose model and hit the ground running.

      1. charcircuit · · focus · HN ↗
        But why does it need to be integrated with a minimal coding agent? Trying to support every possible thing that exists goes against being minimal.
        1. the_mitsuhiko · · focus · HN ↗
          It is in that sense not integrated with the coding agent. It&#x27;s just that some things cannot be done with bash alone, at least not as trivially. So if you were asking Pi to utilize Jev, it would not really have the right tools available to make sense of it, even though pi-ai, the underlying library, can make requests to it.

          Codemode as a mechanism can expose non LLM functionality to the coding agent. In that sense, Pi does not have a tool for Jev or other classifiers. It just now makes it easier for the agent to utilize it in the same way as it&#x27;s otherwise quite creative in using bash.

          1. charcircuit · · focus · HN ↗
            &gt;it would not really have the right tools available

            The point of Pi is that the user can tell the agent to improve itself and give it the tools it does need. The minimalism comes from the user creating what they need instead of the maintainers trying to support everything for the users.

            1. pkulak · · focus · HN ↗
              Sure, but some things are too low-level to be skills or extensions. Code mode seems like that to me.
              1. charcircuit · · focus · HN ↗
                With Pi the agent edits agent itself. That&#x27;s one of the reasons it&#x27;s written in typescript, to make such iteration fast. Going even lower, into the language runtime or operating system shouldn&#x27;t be necessary but technically also possible.
            2. the_mitsuhiko · · focus · HN ↗
              &gt; The point of Pi is that the user can tell the agent to improve itself and give it the tools it does need.

              The point of Pi is to be minimal but also follow what the models need. We were pretty outspoken that models need code execution, and that&#x27;s why Pi to this day has a very small set of tools available. However as more and more training with these models abstracts even over toolcalls themselves with code mode and similar things, it requires changes to Pi.

              Mario and I talked about this last week if you want to know our thinking: <a href="https:&#x2F;&#x2F;x.com&#x2F;pidotdev&#x2F;status&#x2F;2104510506627121451" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;pidotdev&#x2F;status&#x2F;2104510506627121451

              And yes, that&#x27;s why there is no Jev tool in Pi either.

        2. coldtea · · focus · HN ↗
          The agent code is minimal. What it supports doesn&#x27;t have to be, when that support doesn&#x27;t require much of it.
      2. alex7o · · focus · HN ↗
        Even jev is not truly novel, but it&#x27;s latency is, you can use a reranker and get the same things but not the same speed.
        1. Foobar8568 · · focus · HN ↗
          Well, according to claude and Jevbench, Qwen 3.6 35b with ninfer on a RTX 5090@480W is like 3-5 time slower but 10%-15% better performance on the public set, I could see prefill &gt; 15k for 700-800decode. Latency against what and which hardware? I don&#x27;t really get jev...
          1. alex7o · · focus · HN ↗
            Look I can convince my boss to pay for jev, but I won&#x27;t convince him to run our prod stuff on a rented vast.ai 5090. And the pricing wouldn&#x27;t be worth it. If you have ideas I would be glad to hear them
            1. luipugs · · focus · HN ↗
              Stage a coup to usurp your boss.
        2. mrkn1 · · focus · HN ↗
          If you want something even lighter than Jev to compare against, there&#x27;s also gutsy (<a href="https:&#x2F;&#x2F;github.com&#x2F;kouhxp&#x2F;gutsy" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;kouhxp&#x2F;gutsy) runs on CPU
      3. prometheus1992 · · focus · HN ↗
        &gt;&gt;The only thing that makes Jev and the likes particularly interesting is that it is a general purpose classifier.

        General purpose classifiers have existed and proven useful for quite a while now. We used these last year. for vision and text both.

      4. bathtub365 · · focus · HN ↗
        What’s the earliest classification model you know of?
        1. doormatt · · focus · HN ↗
          Frank Rosenblatt introduced the Perceptron in 1957–1958.
        2. strangecasts · · focus · HN ↗

          [dead]

        3. tel · · focus · HN ↗
          [delayed]
      5. peab · · focus · HN ↗
        the only thing novel about Jev is the incredible PR&#x2F;Marketing push that they achieved
      6. vinhnx · · focus · HN ↗
        I built VT Code in Rust for similar reasons. Single binary, low runtime overhead, with extensibility mostly handled through MCP.

        <a href="https:&#x2F;&#x2F;github.com&#x2F;vinhnx&#x2F;vtcode" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vinhnx&#x2F;vtcode

    5. extr · · focus · HN ↗
      I agree. I don&#x27;t necessarily &quot;trust&quot; Anthropic and OpenAI when it comes to CC&#x2F;Codex respectively, but I respect that they have immense internal resources and telemetry to be able to understand what features move the needle and nudge traces in the right direction. I don&#x27;t understand how non-labs judge feature inclusion? Just vibes?
      1. Aperocky · · focus · HN ↗
        What makes you think that labs don&#x27;t operate on &quot;vibes&quot;?

        If there&#x27;s anything that I can conclude about Anthropics idea of how a LLM should speak. Vibes would have been an euphemism

    6. alexhans · · focus · HN ↗
      It was already very good and has been used&#x2F;battle tested by many us for a long time.

      Some tools used to be 0.x for ages and, in this case, the 1.0 signals they&#x27;re happy enough and allows them to promote things in a better way.

      This (edit the durable part) is I guess the natural evolution of playing around building temporal like things for a need that many have.

    7. ryanisnan · · focus · HN ↗
      Human judgement is a thing.
      1. andix · · focus · HN ↗
        Yeah, they are a small team, they just take a decision. Done.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.