I'm not exactly following through with the claim, can someone explain how the built-in classification would not necessitate more tokens used, or be much different from turning on reasoning? Not that I don't see the difference, I just doing see how OpenAI would do it well.
AFAIK Jev is nothing special technically so it's easy to embed it as an another tool for the LLM? For many batch tasks it can still be quite a token saver I think.
Or they can even offer it as a standalone API if deemed worth it.
1) It's very cheap and fast - you provide one input and many potential classifications, and the compute to ingest the input is shared.
2) It generates structured output natively - guaranteed to be correct
3) It's output probabilities are calibrated to actually mean something
OpenAI, or anyone else, could certainly replicate it - there are already articles guessing how Jev achieves its "parallel" classifications, but it seems the AI companies need to decide are they in the business of providing intelligence/tokens, or are they in the application business trying to compete with all their customers (not that Jev uses OpenAI).
the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
I did, in the first days Jev came out, when people were bringing it up. Another assumption. Please review the HN commenting guidelines, the one which starts with "Please don't comment on whether someone read an article." is relevant here.
Nothing in that paper changes that ML algorithms are dependent on the training data. We can step back from Jev and algos to consider Bayes Theorem. If your sample is not representative of the population, your resulting statistics will be off. The same is true here. If the data you train a model like Jev with is not representative, the probabilities and confidences it outputs will not be representative.
What makes Jev interesting is that it works well out of the box across domains. What people who are well known in the field believe is that this is the result of Typesafe having a really good training data set. People are saying similar of MiMo-2.6 today.
"Did you read the article" doesn't apply to a link someone put in a comment. If you are going to be a hall monitor, at least do it properly. You are just acting in bad faith at this point.
The relevant data is the reasoning trace. Doesn't need user data. You can learn from people's detailed reasoning steps how confident they are, even outside your domain.
We were talking about Jev and probability, now you're changing the problem, a rhetorical trick some people try to employ.
Another that uses dice rolling, coin flips, and an inventory level example to drive home the point that Jev's output are not real probabilities for outcomes.
I taught it (ML course; a day on RL, at a university), you should really stop making assumptions friend. Data quality and coverage matters in learning algorithms.
Here's one of the books used in that course <a href="https://amlbook.com/" rel="nofollow">https://amlbook.com/
Thinking blocks are not a place you can derive real confidence scores in LLMs
My initial comment and every one following is about RLCR and that paper. You don't appear to grasp the basics of that paper, it's reward function or how the optimizer is updating weights.
> You are out of your depth and grasping at straws.
Do you have any credentials or evidence that others can use to determine if this statement is not more accurately describing the author who wrote it?
Perhaps a PhD in ML, research output like published papers, or teaching/professional experience - all things I have
We could debate the merits of the paper contents, but I suspect you have intentionally moved on to personal attacks. Regardless, nothing you have said (nor can be found in this paper) has been a counter argument that learning algorithms are sensitive to training data, where the measured output difference is used by the optimization algorithm when updating the parameters. Garbage in, garbage out is a saying for a reason. No algorithm fixes non-representative data.
>you have to have training data with accurate probabilities
This was your claim. If you can't read and understand that paper in relation to your claim, you are out of your depth. You haven't made a single claim relevant to that paper - just hand wavy comments about data.
I clarified multiple times that I meant "accurate, representative data" to "derive accurate probabilities"
you are still employing underhanded techniques in an attempt "win an internet debate" (my impression)
try being more accommodating and flexible over repeating the same lame things
it's not hard to say, "ah I see what you were trying to say..." and move towards a more constructive conversation
RLCR / Jev et al. can only give as accurate predictions and probabilities as the underlying data they are trained on represents. Biased data results in biased probabilities, no algorithm fixes this. Can we agree on this point?
They post train on big math and improve calibration across five different non math benchmarks so the new claim "can only give as accurate predictions and probabilities as the underlying data they are trained on represents" is also off base. RLCR generates its own calibration examples from ordinary questions and answer keys. RL usually isn't trying to represent a data set, it's closer to search.
I'm not expecting you to human RL here, but I do hope next time you remember there is another imperfect human behind the screen and are less too online in future interactions
> 2) It generates structured output natively - guaranteed to be correct
It's not guaranteed to be correct: it's guaranteed to be _formatted in a particular way_. You can get the same thing with grammars on any LLM.
Jev and Jev-like models have other advantages, but I feel like people forget grammars exist for LLMs.
Grammars do risk pushing models off distribution in a way that impacts their output quality in a way Jev allegedly does not suffer from. Additionally, Jev's ability to answer questions independently is also exciting. Using an LLM to answer multiple questions in one generation has the property of earlier answers influencing later ones. TBD how many of TypeSafe's claims stand up, but my testing so far is promising. I hope they author some papers on their methods as well, but that might destroy their moat.
If you really know what grammers did, grammer is a filter to mask out option llm provided but you don't like.
It does not change potential distribution in any means. It DROPS part of answer model returned directly.
The text generation model go wild because model relies on previous section it answered to continue later section. And because now it contain item model have no idea, it is completely screwed.
In the case you only require model to answer one of a,b,c,d and don't care about later segment at all. It don't really matter.
What I mean is that, in general, constrained decoding can push model output off into less probable regimes. This is well studied; see for example <a href="https://arxiv.org/pdf/2606.21619" rel="nofollow">https://arxiv.org/pdf/2606.21619. The mask may only retain very improbable logits. In pathological cases, the constrained output may be little better than noise filtered through the constraint. When using existing structured output APIs, it may not be possible to even know.
You don't even bother text after the [a] at first place in this case
Your question is something like
anwser only a,b,c,d for following question
a. b. c. d....
the model output possibility of next character
a: 0.8 b: 0.7 c: 0.3 f: 0.2 d: 0.1
If the list contains option you did not provide.
The model is confused anyway, it don't matter if you use grammer to filter out the bad option or not, the answer is screwed already.
Yes, agreed. I was speaking in general, of course. This particular topic is of interest to me, so thinking of the edge cases and confounds vs Jev.
In your example, I would expect an LLM to do fine and if you have access to the raw logits you can measure whether or not it was confused and assign a confidence to the answer it gave.
I do think that Jev handles more than this though and, in my early testing, does things that are not easily accomplished with guided decoding techniques.
The way jev actually internally work could be interesting though. I believe most llm are only tuned to return the first or second logits(or a few more) correctly as that is what the sampler would choose anyway. Do they alter existing model for better behavior across all options? Or they distilled one to have the proper behavior? We can only guess without the actual implementation.
Yes! I really hope they release some papers on their techniques. I am very curious.
I ran it through MMLU a few days ago and it scored ~90% so seems to have a lot of general world knowledge trained in. Makes me think your speculation is right. I have some credits left, might try and think of an experiment. I saw a gist where someone was asking it which model it was and it was picking qwen a lot, but who knows...
Although the underlying model is unknown. If it expose input token count, the tokenizer may be probable though. Most tokenizer segemnts wildly different in CJK inputs. It can probably be used to fingerprint the tokenizer based on token count if it is using existing tokenizer.
Not an expert at all here, but I saw a comment on the jev post saying that it you constrain an LLM suck that it outputs a valid structure, if the token with the highest probability is not the one that you expected because of the structure (and so you pick the valid lower one) this means the LLM was already confused and your answer is less likely to be correct anyways.
> It's output probabilities are calibrated to actually mean something
Don't fall for marketing BS so easily.
Jev can output drastically different probabilities if you simply reorder the list of choices. And Jev's "confidence" output is fake/redundant - it's just a formula applied to probabilities, it conveys no additional information.
I bet they will eventually "fix" (read hide under the rug) the ordering problem by ordering the list on the backend before feeding to the model.
It seems that anyway most of the value is in the speed and cost.
If it really matters to you whether whether some business-specific classification confidence is above/below some specific threshold (vs just relative order), then you'd be better off training or fine tuning a custom model for that. Maybe that is something that TypeSafe are planning to also provide?
tolugenius · · focus · HN ↗
mnicky · · focus · HN ↗
Or they can even offer it as a standalone API if deemed worth it.
HarHarVeryFunny · · focus · HN ↗
1) It's very cheap and fast - you provide one input and many potential classifications, and the compute to ingest the input is shared.
2) It generates structured output natively - guaranteed to be correct
3) It's output probabilities are calibrated to actually mean something
OpenAI, or anyone else, could certainly replicate it - there are already articles guessing how Jev achieves its "parallel" classifications, but it seems the AI companies need to decide are they in the business of providing intelligence/tokens, or are they in the application business trying to compete with all their customers (not that Jev uses OpenAI).
verdverm · · focus · HN ↗
danielmarkbruce · · focus · HN ↗
<a href="https://arxiv.org/pdf/2507.16806" rel="nofollow">https://arxiv.org/pdf/2507.16806
verdverm · · focus · HN ↗
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
the underlying data set needs to be representative
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
and then you are going to ignore all the research and results that clearly show otherwise? why?
what might we infer about the importance of data from a learning algorithm like decision trees?
danielmarkbruce · · focus · HN ↗
Existing datasets, different reward function.
verdverm · · focus · HN ↗
I did, in the first days Jev came out, when people were bringing it up. Another assumption. Please review the HN commenting guidelines, the one which starts with "Please don't comment on whether someone read an article." is relevant here.
Nothing in that paper changes that ML algorithms are dependent on the training data. We can step back from Jev and algos to consider Bayes Theorem. If your sample is not representative of the population, your resulting statistics will be off. The same is true here. If the data you train a model like Jev with is not representative, the probabilities and confidences it outputs will not be representative.
What makes Jev interesting is that it works well out of the box across domains. What people who are well known in the field believe is that this is the result of Typesafe having a really good training data set. People are saying similar of MiMo-2.6 today.
verdverm · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=49816899">https://news.ycombinator.com/item?id=49816899
<a href="https://www.alexmolas.com/2026/09/23/jev-cant-be-calibrated.html" rel="nofollow">https://www.alexmolas.com/2026/09/23/jev-cant-be-calibrated....
[deleted] · · focus · HN ↗
[deleted]
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
We both know who is
> just acting in bad faith at this point.
danielmarkbruce · · focus · HN ↗
Take RL 101. This is a common pattern.
verdverm · · focus · HN ↗
Another that uses dice rolling, coin flips, and an inventory level example to drive home the point that Jev's output are not real probabilities for outcomes.
<a href="https://news.ycombinator.com/item?id=49830385">https://news.ycombinator.com/item?id=49830385
> Take RL 101
I taught it (ML course; a day on RL, at a university), you should really stop making assumptions friend. Data quality and coverage matters in learning algorithms.
Here's one of the books used in that course <a href="https://amlbook.com/" rel="nofollow">https://amlbook.com/
Thinking blocks are not a place you can derive real confidence scores in LLMs
danielmarkbruce · · focus · HN ↗
You are out of your depth and grasping at straws.
verdverm · · focus · HN ↗
Do you have any credentials or evidence that others can use to determine if this statement is not more accurately describing the author who wrote it?
Perhaps a PhD in ML, research output like published papers, or teaching/professional experience - all things I have
We could debate the merits of the paper contents, but I suspect you have intentionally moved on to personal attacks. Regardless, nothing you have said (nor can be found in this paper) has been a counter argument that learning algorithms are sensitive to training data, where the measured output difference is used by the optimization algorithm when updating the parameters. Garbage in, garbage out is a saying for a reason. No algorithm fixes non-representative data.
danielmarkbruce · · focus · HN ↗
This was your claim. If you can't read and understand that paper in relation to your claim, you are out of your depth. You haven't made a single claim relevant to that paper - just hand wavy comments about data.
verdverm · · focus · HN ↗
you are still employing underhanded techniques in an attempt "win an internet debate" (my impression)
try being more accommodating and flexible over repeating the same lame things
it's not hard to say, "ah I see what you were trying to say..." and move towards a more constructive conversation
RLCR / Jev et al. can only give as accurate predictions and probabilities as the underlying data they are trained on represents. Biased data results in biased probabilities, no algorithm fixes this. Can we agree on this point?
danielmarkbruce · · focus · HN ↗
verdverm · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=c1Fv1uKTd-w" rel="nofollow">https://www.youtube.com/watch?v=c1Fv1uKTd-w
oh-seven
danielmarkbruce · · focus · HN ↗
alex_sf · · focus · HN ↗
> 2) It generates structured output natively - guaranteed to be correct
It's not guaranteed to be correct: it's guaranteed to be _formatted in a particular way_. You can get the same thing with grammars on any LLM.
Jev and Jev-like models have other advantages, but I feel like people forget grammars exist for LLMs.
time0ut · · focus · HN ↗
mmis1000 · · focus · HN ↗
It does not change potential distribution in any means. It DROPS part of answer model returned directly.
The text generation model go wild because model relies on previous section it answered to continue later section. And because now it contain item model have no idea, it is completely screwed.
In the case you only require model to answer one of a,b,c,d and don't care about later segment at all. It don't really matter.
time0ut · · focus · HN ↗
mmis1000 · · focus · HN ↗
Your question is something like
anwser only a,b,c,d for following question a. b. c. d....
the model output possibility of next character a: 0.8 b: 0.7 c: 0.3 f: 0.2 d: 0.1
If the list contains option you did not provide. The model is confused anyway, it don't matter if you use grammer to filter out the bad option or not, the answer is screwed already.
time0ut · · focus · HN ↗
In your example, I would expect an LLM to do fine and if you have access to the raw logits you can measure whether or not it was confused and assign a confidence to the answer it gave.
I do think that Jev handles more than this though and, in my early testing, does things that are not easily accomplished with guided decoding techniques.
mmis1000 · · focus · HN ↗
time0ut · · focus · HN ↗
I ran it through MMLU a few days ago and it scored ~90% so seems to have a lot of general world knowledge trained in. Makes me think your speculation is right. I have some credits left, might try and think of an experiment. I saw a gist where someone was asking it which model it was and it was picking qwen a lot, but who knows...
Anyway, thank you for the interesting discussion!
mmis1000 · · focus · HN ↗
LelouBil · · focus · HN ↗
Is this actually true ?
hbrn · · focus · HN ↗
Don't fall for marketing BS so easily.
Jev can output drastically different probabilities if you simply reorder the list of choices. And Jev's "confidence" output is fake/redundant - it's just a formula applied to probabilities, it conveys no additional information.
I bet they will eventually "fix" (read hide under the rug) the ordering problem by ordering the list on the backend before feeding to the model.
HarHarVeryFunny · · focus · HN ↗
If it really matters to you whether whether some business-specific classification confidence is above/below some specific threshold (vs just relative order), then you'd be better off training or fine tuning a custom model for that. Maybe that is something that TypeSafe are planning to also provide?