I built non-autoregressive decision models with RL a year ago
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
I built non-autoregressive decision models with RL a year ago
Unofficial Hacker News client; not affiliated with Y Combinator.
prometheus1992 · · focus · HN ↗
"Breakthrough", "our research went in another direction" , "Two years in stealth", "System One thinking model", "Jev can't hallucinate", "RLCD","We are doing very cool stuff, but we will have to hire you to tell you", - these are some of the things that they said on their website on the launch blog.
I had used versions of bert to achieve the same functionality years ago. But to me it seems like they were able to trick the VCs with "can't hallucinate" etc.
To the above author, kudos for sharing your work and making it open. Something like this shouldn't be closed in the first place when it has been available for so many years
seizethecheese · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
And they’re acting like their probability isn’t as hallucinated as any other LLM guess.
seizethecheese · · focus · HN ↗
TeMPOraL · · focus · HN ↗
seizethecheese · · focus · HN ↗
dropofwill · · focus · HN ↗
That does align with my experience, though we’re not using anything close to frontier for these sort of tasks.
I am interested if it can actually improve on that. As an engineer i like the elegance of guaranteed output, but the retry works pretty well in practice.
fastball · · focus · HN ↗
For example, if you feed in some context to Jev and Claude Haiku and say "make the appropriate tool call based on this context", Claude (or any other frontier LLM) will hallucinate tool calls some percentage of the time. Jev will not. While yes, the "will not" is constrained by Jev's (lack of) capabilities in some sense, this is actually a very real need for a wide variety of use-cases people are currently using off-the-shelf LLMs for at the moment.
Probably the better example is the whole probability thing, where even if you use something like constrained decoding to ensure an LLM only outputs a certain schema, and therefore can't hallucinate a class, if you ask for probabilities, the probabilities output by the model are just hallucinations. Jev meanwhile is outputting calibrated probabilities for different choices based on the actual landscape.
refulgentis · · focus · HN ↗
fastball · · focus · HN ↗
adastra22 · · focus · HN ↗
That doesn’t mean the models outputs are correct, nor is TypeSafe claiming that afaict.