‹ BackHN Continuity

Thread

I built non-autoregressive decision models with RL a year ago

1363 points · 319 comments · nandakishor_ml

  1. prometheus1992 · · focus · HN ↗
    I think the main gripe that people had with Jev and Typesafe was the language used when they launched. To me personally it seemed like a parody/con/shady at first.

    "Breakthrough", "our research went in another direction" , "Two years in stealth", "System One thinking model", "Jev can't hallucinate", "RLCD","We are doing very cool stuff, but we will have to hire you to tell you", - these are some of the things that they said on their website on the launch blog.

    I had used versions of bert to achieve the same functionality years ago. But to me it seems like they were able to trick the VCs with "can't hallucinate" etc.

    To the above author, kudos for sharing your work and making it open. Something like this shouldn't be closed in the first place when it has been available for so many years

    1. yojo · · focus · HN ↗
      Is this equivalent though? The Laya article ends with “ Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.”

      I have a dozen different things at work that are currently using LLMs as classifiers for different questions. I don’t have the time, data, or resources to fine tune a model for each of them.

      I haven’t had a chance to plug in Jev yet (waiting on approvals), but if it has the general intelligence claimed in the press release, then Laya is in no way comparable for my use case, and whatever TypeSafe has done is a substantial innovation over the Laya paper.

      1. vasco · · focus · HN ↗
        In a world of agents, doing a BERT run takes about 2 hours from having an empty folder. Just a thought you could consider. Once you've done the first you can do the rest of them before the end of the work day.
        1. adastra22 · · focus · HN ↗
          BERT run on what? You would need training data, no? The things would use Jev for have no training data. Not that kind of problem.
          1. hamandcheese · · focus · HN ↗
            Presumably, if you are positioned to plug in Jev (or an LLM classifier), then you are also positioned to collect training data.
            1. yojo · · focus · HN ↗
              The domain is code analysis, all languages and frameworks. It’s b2b SaaS, so total volume is not incredibly high. And many customers have contract clauses that we don’t train on their data.

              I’m not convinced we could train easily here, or that it’s worth the investment compared to (previously) spending fractional cents on Luna, or now paying even less on Jev. Especially given that these numbers are not meaningful to our margins.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.