I'm not exactly following through with the claim, can someone explain how the built-in classification would not necessitate more tokens used, or be much different from turning on reasoning? Not that I don't see the difference, I just doing see how OpenAI would do it well.
AFAIK Jev is nothing special technically so it's easy to embed it as an another tool for the LLM? For many batch tasks it can still be quite a token saver I think.
Or they can even offer it as a standalone API if deemed worth it.
1) It's very cheap and fast - you provide one input and many potential classifications, and the compute to ingest the input is shared.
2) It generates structured output natively - guaranteed to be correct
3) It's output probabilities are calibrated to actually mean something
OpenAI, or anyone else, could certainly replicate it - there are already articles guessing how Jev achieves its "parallel" classifications, but it seems the AI companies need to decide are they in the business of providing intelligence/tokens, or are they in the application business trying to compete with all their customers (not that Jev uses OpenAI).
the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
I did, in the first days Jev came out, when people were bringing it up. Another assumption. Please review the HN commenting guidelines, the one which starts with "Please don't comment on whether someone read an article." is relevant here.
Nothing in that paper changes that ML algorithms are dependent on the training data. We can step back from Jev and algos to consider Bayes Theorem. If your sample is not representative of the population, your resulting statistics will be off. The same is true here. If the data you train a model like Jev with is not representative, the probabilities and confidences it outputs will not be representative.
What makes Jev interesting is that it works well out of the box across domains. What people who are well known in the field believe is that this is the result of Typesafe having a really good training data set. People are saying similar of MiMo-2.6 today.
tolugenius · · focus · HN ↗
mnicky · · focus · HN ↗
Or they can even offer it as a standalone API if deemed worth it.
HarHarVeryFunny · · focus · HN ↗
1) It's very cheap and fast - you provide one input and many potential classifications, and the compute to ingest the input is shared.
2) It generates structured output natively - guaranteed to be correct
3) It's output probabilities are calibrated to actually mean something
OpenAI, or anyone else, could certainly replicate it - there are already articles guessing how Jev achieves its "parallel" classifications, but it seems the AI companies need to decide are they in the business of providing intelligence/tokens, or are they in the application business trying to compete with all their customers (not that Jev uses OpenAI).
verdverm · · focus · HN ↗
danielmarkbruce · · focus · HN ↗
<a href="https://arxiv.org/pdf/2507.16806" rel="nofollow">https://arxiv.org/pdf/2507.16806
verdverm · · focus · HN ↗
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
the underlying data set needs to be representative
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
and then you are going to ignore all the research and results that clearly show otherwise? why?
what might we infer about the importance of data from a learning algorithm like decision trees?
danielmarkbruce · · focus · HN ↗
Existing datasets, different reward function.
verdverm · · focus · HN ↗
I did, in the first days Jev came out, when people were bringing it up. Another assumption. Please review the HN commenting guidelines, the one which starts with "Please don't comment on whether someone read an article." is relevant here.
Nothing in that paper changes that ML algorithms are dependent on the training data. We can step back from Jev and algos to consider Bayes Theorem. If your sample is not representative of the population, your resulting statistics will be off. The same is true here. If the data you train a model like Jev with is not representative, the probabilities and confidences it outputs will not be representative.
What makes Jev interesting is that it works well out of the box across domains. What people who are well known in the field believe is that this is the result of Typesafe having a really good training data set. People are saying similar of MiMo-2.6 today.
verdverm · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=49816899">https://news.ycombinator.com/item?id=49816899
<a href="https://www.alexmolas.com/2026/09/23/jev-cant-be-calibrated.html" rel="nofollow">https://www.alexmolas.com/2026/09/23/jev-cant-be-calibrated....
[deleted] · · focus · HN ↗
[deleted]