‹ BackHN Continuity

Thread

I built non-autoregressive decision models with RL a year ago

1363 points · 319 comments · nandakishor_ml

  1. prometheus1992 · · focus · HN ↗
    I think the main gripe that people had with Jev and Typesafe was the language used when they launched. To me personally it seemed like a parody/con/shady at first.

    "Breakthrough", "our research went in another direction" , "Two years in stealth", "System One thinking model", "Jev can't hallucinate", "RLCD","We are doing very cool stuff, but we will have to hire you to tell you", - these are some of the things that they said on their website on the launch blog.

    I had used versions of bert to achieve the same functionality years ago. But to me it seems like they were able to trick the VCs with "can't hallucinate" etc.

    To the above author, kudos for sharing your work and making it open. Something like this shouldn't be closed in the first place when it has been available for so many years

    1. seizethecheese · · focus · HN ↗
      I was confused by the “can’t hallucinate” thing, because it sounded like BS but people were taking it seriously. I purposefully asked a stupid question sort of like “this can’t hallucinate because it only has one output and there’s a schema?”. Was disappointed to learn the answer was yes.
      1. MisterMunchkin · · focus · HN ↗
        Yeah it’s hilarious, it definitely can hallucinate. Just because it can only hallucinate “A” or “B” rather than a whole paragraph, doesn’t mean it is suddenly more accurate.

        And they’re acting like their probability isn’t as hallucinated as any other LLM guess.

        1. seizethecheese · · focus · HN ↗
          They’re definining hallucination as a property of iterative generation, which is fair enough, but then it’s sort of like selling a boat and saying it doesn’t need tire changes.
          1. TeMPOraL · · focus · HN ↗
            It does make some sense given they're positioning it as alternative to the normal way you'd implement such output shape, which is to slap a prompt on a frontier LLM and maybe run it in "constrained output" mode if you like things fancy. Against that use case, the "no hallucinations" and parallelism and cost claims all sound legitimate and useful -- and similarly, "but we could do that with BERT two years ago" does not.
            1. seizethecheese · · focus · HN ↗
              I mean, the constrained output mode also doesn’t hallucinate in this sense.
              1. dropofwill · · focus · HN ↗
                They do actually admit that about constrained decoding somewhere in the docs. They argue it’s useless in practice because when the constraints actually kick in it harms the output too much and that it’s better to just error and retry in those cases.

                That does align with my experience, though we’re not using anything close to frontier for these sort of tasks.

                I am interested if it can actually improve on that. As an engineer i like the elegance of guaranteed output, but the retry works pretty well in practice.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.