First, congrats to the team on launching something genuinely interesting and new.
Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that.
This is interesting, but the speed comparison seems misleading? A generative model that can output code in a Turing-complete language can do anything a computer can do.
Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring, but it's nothing like the code generating models we're all using today for code and automation.
Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value. You can enforce structured output from an LLM too, with an appropriate harness, etc.
Assuming there's no funny business, the Doom demo is cool.
When they say "can't hallucinate" they mean they produce a confidence value for every result, so you could see for example it has 0.1 confidence, and you can disregard the result - that'd be different from hallucinating where it believes it's correct
that's right, but because these models are probabilistic, it's also possible to be confidently wrong (and all future models will be smarter still and still have that possibility)
Correct. Not to say we're getting into the weeds of probability here as well.
"What are the odds a thunder will strike in Paris at 1pm UTC of 2026-09-16" - that could be a 0.001 chance going from blind historical measurements; 0.01 if it's raining; or 1 or 1 after the date has passed.
The probability values don’t really represent confidence in modern LLMs though, especially after RLHF and RLVR.
System One says they use RLCD, Reinforcement Learning for Calibrated Decisions, which presumably has accurate probabilities as an explicit optimisation goal.
RLVR generally upweights tokens along the whole thinking trace that led to a correct answer, whether each token was "correct" or not. RLVR doesn't train a model to output an 80% likelihood, it just trains it to produce correct answers, and not to produce incorrect ones.
System One hasn't said how RLCD works, but they do say it is explicitly training models to output "calibrated" probabilities, which makes it distinct from RLVR. This is how they describe it:
> System One models are trained for calibrated decisions: their probabilities are optimized against outcomes to reflect uncertainty.
In RLCD, you basically massively negatively reward a distribution that is {yes: 0.9, no: 0.1} if the answer was no, and less negatively reward a {yes: 0.6, no: 0.4}. Various nuances when designing the details, but that is the rough idea
Nothing, but imagine using LLMs for a classification task
People out there are so resigned to the models being unreliable that they are really doing things like hallucinating deliberately, and then matching the hallucinations to embeddings -
I'm certainly not resigned to that, at least for classification.
Even non-frontier models are absurdly good at this in a broad sense.
Which would make it hard to judge "a model that will never produce unreliable outputs in the first place" against something that is already really, really good and exceptional in domain-specific areas with the tiniest amount of elbow grease.
It would be great to see benchmarks for Jev that demonstrate the value of calibrated uncertainty.
For example, one could set a confidence threshold over which we trust the model decision, and otherwise reject. This provides a lever to trade-off accuracy and automation %.
Then we can ask questions like "What % of decisions can we automate to achieve 90% accuracy"?
i ran some interesting experiments on jev today and wanted to share the results. i compared cost and latency of support ticket triage with different arrival rates for jev vs mainstream LLMs: <a href="https://suraj-website-eta.vercel.app/blog/what-a-correct-decision-costs" rel="nofollow">https://suraj-website-eta.vercel.app/blog/what-a-correct-dec...
I would be happy enough with: only produces what it can verify with sources.
If you eg try to remember a court case (ie produce the reference via LLM token generation only), it's easy enough to check with your data whether it really exists. Similar for following links and other references.
If your data or sources are wrong, obviously your report about them will be wrong. But I wouldn't call that a hallucination.
It's not a binary thing. You can get closer or further away from that standard.
And humans also behave differently in different contexts. A conversation at the pub has more such hallucinations than a formal deposit in court. For the latter, a good lawyer will look at her shoes, when you ask him what colour her laces are.
That's not true that incorrect sources means incorrect report. Often, LLMs have some sense of what is true, and due to that, they hallucinate plausible sources that appear to back that knowledge up.
No, I don't believe so. Hallucinations are not "high probability" in a real sense. They are an artifact of the random walk the inference algorithm takes, which causes it to latch on to and chase attractors in the noise. This random walk behavior is necessary for chat interfaces to be useful, but are less critical to typed output predictors. I'm guessing they found some optimization that is possible if you give up caring about chat.
What we would want to see if a confidence value that is in line with the actual correctness. If the value is 0.9 for 1000 different answers, then approximately 900 of those answers should be correct.
I read "hallucinations" as "generates novel output with no grounding/source". i.e. "it just made something completely up".
I believe their "accuracy" metric (sonnet 5 level) is where "right/wrong" is measured.
What about the LLM calls though that are done midchain? In the Home Assistant video the multi-intent prompt gets split using what looks like a traditional llm model, which I'm assuming is vulnerable to classical hallucinations.
jacobgold · · focus · HN ↗
Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that.
This is interesting, but the speed comparison seems misleading? A generative model that can output code in a Turing-complete language can do anything a computer can do.
Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring, but it's nothing like the code generating models we're all using today for code and automation.
Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value. You can enforce structured output from an LLM too, with an appropriate harness, etc.
Assuming there's no funny business, the Doom demo is cool.
dbbk · · focus · HN ↗
CompleteSkeptic · · focus · HN ↗
flockonus · · focus · HN ↗
"What are the odds a thunder will strike in Paris at 1pm UTC of 2026-09-16" - that could be a 0.001 chance going from blind historical measurements; 0.01 if it's raining; or 1 or 1 after the date has passed.
janalsncm · · focus · HN ↗
jiggawatts · · focus · HN ↗
bigglebear · · focus · HN ↗
sothatsit · · focus · HN ↗
System One says they use RLCD, Reinforcement Learning for Calibrated Decisions, which presumably has accurate probabilities as an explicit optimisation goal.
nkozyra · · focus · HN ↗
sothatsit · · focus · HN ↗
System One hasn't said how RLCD works, but they do say it is explicitly training models to output "calibrated" probabilities, which makes it distinct from RLVR. This is how they describe it:
> System One models are trained for calibrated decisions: their probabilities are optimized against outcomes to reflect uncertainty.
porridgeraisin · · focus · HN ↗
orbital-decay · · focus · HN ↗
zenlikethat · · focus · HN ↗
People out there are so resigned to the models being unreliable that they are really doing things like hallucinating deliberately, and then matching the hallucinations to embeddings -
<a href="https://softwaredoug.com/blog/2026/08/10/hypothetical-classifications" rel="nofollow">https://softwaredoug.com/blog/2026/08/10/hypothetical-classi...
You could do that or you could just... use a model that will never produce unreliable outputs in the first place.
threecheese · · focus · HN ↗
zenlikethat · · focus · HN ↗
threecheese · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
nkozyra · · focus · HN ↗
Even non-frontier models are absurdly good at this in a broad sense.
Which would make it hard to judge "a model that will never produce unreliable outputs in the first place" against something that is already really, really good and exceptional in domain-specific areas with the tiniest amount of elbow grease.
Speed and cost look good though (for now)!
ActivePattern · · focus · HN ↗
For example, one could set a confidence threshold over which we trust the model decision, and otherwise reject. This provides a lever to trade-off accuracy and automation %.
Then we can ask questions like "What % of decisions can we automate to achieve 90% accuracy"?
suraj_phanindra · · focus · HN ↗
[dead]
suraj_phanindra · · focus · HN ↗
TheWayWithin · · focus · HN ↗
[dead]
8note · · focus · HN ↗
llm hallucinations are high probability tokens that are incorrect vs the real world
dozerly · · focus · HN ↗
resonious · · focus · HN ↗
rpunkfu · · focus · HN ↗
jubilanti · · focus · HN ↗
spencerflem · · focus · HN ↗
A calculator either gets the right answer or doesn’t answer.
It wouldn’t have to be all knowing as long as it knew perfectly what it doesn’t know
baq · · focus · HN ↗
stpedgwdgfhgdd · · focus · HN ↗
baq · · focus · HN ↗
eru · · focus · HN ↗
I would be happy enough with: only produces what it can verify with sources.
If you eg try to remember a court case (ie produce the reference via LLM token generation only), it's easy enough to check with your data whether it really exists. Similar for following links and other references.
If your data or sources are wrong, obviously your report about them will be wrong. But I wouldn't call that a hallucination.
baq · · focus · HN ↗
eru · · focus · HN ↗
And humans also behave differently in different contexts. A conversation at the pub has more such hallucinations than a formal deposit in court. For the latter, a good lawyer will look at her shoes, when you ask him what colour her laces are.
hdjrudni · · focus · HN ↗
Humans are known to hallucinate a lot. Ask 10 different witnesses at a crime scene what they saw and they'll all report different things.
A good, non-hallucinating LLM would only report things for which it has evidence. It would consult the facts every single time.
It's a pain in the butt for humans to fact-check everything but LLMs can quickly look up all kinds of stuff. That's what makes them useful.
eru · · focus · HN ↗
So you can bolt the fact-check / source-check pass onto whatever other system you have, without having to redesign the underlying system.
tedbradley · · focus · HN ↗
adastra22 · · focus · HN ↗
elil17 · · focus · HN ↗
darylteo · · focus · HN ↗
I believe their "accuracy" metric (sonnet 5 level) is where "right/wrong" is measured.
bradly · · focus · HN ↗
csomar · · focus · HN ↗
tahaazizi · · focus · HN ↗
[dead]