I don't really understand the criteria for when something is 'proven' to the Pi team. Jev and the like took off less than a month ago, but MCP has been growing for nearly 2 years, and it only gets support now?
Pi felt nice when I used it, and I do value keeping things minimal, but I just find the criteria very uneven.
A year later, some things have changed: <a href="https://earendil.com/posts/you-said-no-mcp/" rel="nofollow">https://earendil.com/posts/you-said-no-mcp/
Armin from Earendil here. I think the question is fair, and quite frankly the answer is pretty disappointing: we look at what the models are doing. They are trained on their respective harnesses and we're not here to fight their behavior.
Codex in particular is using responses lite internally and relies on codemode for parallel tool calling. So codemode was a given.
Jev on the other hand is new but it's not the first type of model we had troubles with supporting in Pi and we looked at how to make that make sense. The internal pi-ai SDK supports image generation and classifier models, but without building an extension it was never possible for you to utilize it.
So there was a while functionality of Pi that few people used, because there were no obvious ways to hook it up with the coding agent. Codemode also allows us to close that gap.
And once you have codemode, modern MCP can work quite well if the servers cooperate.
Classification models have been around for literally almost a century at this point. I think it's safe to say they are a proven technology.
The only thing that makes Jev and the likes particularly interesting is that it is a general purpose classifier. In the past, classification tasks meant training a new model to solve your problem. Now you can just use an off the shelf general purpose model and hit the ground running.
It is in that sense not integrated with the coding agent. It's just that some things cannot be done with bash alone, at least not as trivially. So if you were asking Pi to utilize Jev, it would not really have the right tools available to make sense of it, even though pi-ai, the underlying library, can make requests to it.
Codemode as a mechanism can expose non LLM functionality to the coding agent. In that sense, Pi does not have a tool for Jev or other classifiers. It just now makes it easier for the agent to utilize it in the same way as it's otherwise quite creative in using bash.
>it would not really have the right tools available
The point of Pi is that the user can tell the agent to improve itself and give it the tools it does need. The minimalism comes from the user creating what they need instead of the maintainers trying to support everything for the users.
With Pi the agent edits agent itself. That's one of the reasons it's written in typescript, to make such iteration fast. Going even lower, into the language runtime or operating system shouldn't be necessary but technically also possible.
> The point of Pi is that the user can tell the agent to improve itself and give it the tools it does need.
The point of Pi is to be minimal but also follow what the models need. We were pretty outspoken that models need code execution, and that's why Pi to this day has a very small set of tools available. However as more and more training with these models abstracts even over toolcalls themselves with code mode and similar things, it requires changes to Pi.
Mario and I talked about this last week if you want to know our thinking: <a href="https://x.com/pidotdev/status/2104510506627121451" rel="nofollow">https://x.com/pidotdev/status/2104510506627121451
And yes, that's why there is no Jev tool in Pi either.
Well, according to claude and Jevbench, Qwen 3.6 35b with ninfer on a RTX 5090@480W is like 3-5 time slower but 10%-15% better performance on the public set, I could see prefill > 15k for 700-800decode.
Latency against what and which hardware? I don't really get jev...
Look I can convince my boss to pay for jev, but I won't convince him to run our prod stuff on a rented vast.ai 5090. And the pricing wouldn't be worth it. If you have ideas I would be glad to hear them
If you want something even lighter than Jev to compare against, there's also gutsy (<a href="https://github.com/kouhxp/gutsy" rel="nofollow">https://github.com/kouhxp/gutsy) runs on CPU
I agree. I don't necessarily "trust" Anthropic and OpenAI when it comes to CC/Codex respectively, but I respect that they have immense internal resources and telemetry to be able to understand what features move the needle and nudge traces in the right direction. I don't understand how non-labs judge feature inclusion? Just vibes?
hhh · · focus · HN ↗
Pi felt nice when I used it, and I do value keeping things minimal, but I just find the criteria very uneven.
rsalus · · focus · HN ↗
BeetleB · · focus · HN ↗
I don't know if you were aware, but not shipping with MCP was one of its "features":
<a href="https://mariozechner.at/posts/2025-11-02-what-if-you-dont-need-mcp/" rel="nofollow">https://mariozechner.at/posts/2025-11-02-what-if-you-dont-ne...
They let you have it via a plugin/extension.
the_mitsuhiko · · focus · HN ↗
the_mitsuhiko · · focus · HN ↗
Codex in particular is using responses lite internally and relies on codemode for parallel tool calling. So codemode was a given.
Jev on the other hand is new but it's not the first type of model we had troubles with supporting in Pi and we looked at how to make that make sense. The internal pi-ai SDK supports image generation and classifier models, but without building an extension it was never possible for you to utilize it.
So there was a while functionality of Pi that few people used, because there were no obvious ways to hook it up with the coding agent. Codemode also allows us to close that gap.
And once you have codemode, modern MCP can work quite well if the servers cooperate.
octoberfranklin · · focus · HN ↗
That is totally disappointing.
Zambyte · · focus · HN ↗
The only thing that makes Jev and the likes particularly interesting is that it is a general purpose classifier. In the past, classification tasks meant training a new model to solve your problem. Now you can just use an off the shelf general purpose model and hit the ground running.
charcircuit · · focus · HN ↗
the_mitsuhiko · · focus · HN ↗
Codemode as a mechanism can expose non LLM functionality to the coding agent. In that sense, Pi does not have a tool for Jev or other classifiers. It just now makes it easier for the agent to utilize it in the same way as it's otherwise quite creative in using bash.
charcircuit · · focus · HN ↗
The point of Pi is that the user can tell the agent to improve itself and give it the tools it does need. The minimalism comes from the user creating what they need instead of the maintainers trying to support everything for the users.
pkulak · · focus · HN ↗
charcircuit · · focus · HN ↗
the_mitsuhiko · · focus · HN ↗
The point of Pi is to be minimal but also follow what the models need. We were pretty outspoken that models need code execution, and that's why Pi to this day has a very small set of tools available. However as more and more training with these models abstracts even over toolcalls themselves with code mode and similar things, it requires changes to Pi.
Mario and I talked about this last week if you want to know our thinking: <a href="https://x.com/pidotdev/status/2104510506627121451" rel="nofollow">https://x.com/pidotdev/status/2104510506627121451
And yes, that's why there is no Jev tool in Pi either.
coldtea · · focus · HN ↗
alex7o · · focus · HN ↗
Foobar8568 · · focus · HN ↗
alex7o · · focus · HN ↗
luipugs · · focus · HN ↗
mrkn1 · · focus · HN ↗
prometheus1992 · · focus · HN ↗
General purpose classifiers have existed and proven useful for quite a while now. We used these last year. for vision and text both.
bathtub365 · · focus · HN ↗
doormatt · · focus · HN ↗
strangecasts · · focus · HN ↗
[dead]
tel · · focus · HN ↗
peab · · focus · HN ↗
vinhnx · · focus · HN ↗
<a href="https://github.com/vinhnx/vtcode" rel="nofollow">https://github.com/vinhnx/vtcode
extr · · focus · HN ↗
Aperocky · · focus · HN ↗
If there's anything that I can conclude about Anthropics idea of how a LLM should speak. Vibes would have been an euphemism
alexhans · · focus · HN ↗
Some tools used to be 0.x for ages and, in this case, the 1.0 signals they're happy enough and allows them to promote things in a better way.
This (edit the durable part) is I guess the natural evolution of playing around building temporal like things for a need that many have.
ryanisnan · · focus · HN ↗
andix · · focus · HN ↗