I built non-autoregressive decision models with RL a year ago
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
I built non-autoregressive decision models with RL a year ago
Unofficial Hacker News client; not affiliated with Y Combinator.
prometheus1992 · · focus · HN ↗
"Breakthrough", "our research went in another direction" , "Two years in stealth", "System One thinking model", "Jev can't hallucinate", "RLCD","We are doing very cool stuff, but we will have to hire you to tell you", - these are some of the things that they said on their website on the launch blog.
I had used versions of bert to achieve the same functionality years ago. But to me it seems like they were able to trick the VCs with "can't hallucinate" etc.
To the above author, kudos for sharing your work and making it open. Something like this shouldn't be closed in the first place when it has been available for so many years
seizethecheese · · focus · HN ↗
mrbonner · · focus · HN ↗
dojomouse · · focus · HN ↗
There’s a big difference between deterministic and smooth though. Typical LLMs certainly aren’t reliably smooth, so the small prompt change might product a large and unpredictable output change. I’m not sure if that’s any better with the typesafe approach.