‹ BackHN Continuity

Thread

I built non-autoregressive decision models with RL a year ago

1363 points · 319 comments · nandakishor_ml

  1. Oras · · focus · HN ↗
    I played around with Jev last night and did it for classification tasks that I used Gemini 2.5 flash lite with.

    It’s a bit faster and bit cheaper, but this is compared to LLM. The consistency was nice to see, BUT, as someone who trained NLP models prior to LLMs, it’s just BERT with more data. I can see why people would want ready made one shot classifier, and I can see the value of sending multiple classifier in one call, but I wouldn’t call it breakthrough. And I believe many labs will replicate it in no time and might have it as part of their harness.

    I see it as a wake up call for the tech community to go back to basics for most tasks instead of relying solely on generic LLMs.

    1. lhl · · focus · HN ↗
      There have been other "universal"/general classifiers like GLiNER, GLiFormer, etc based on BERTs (Laya itself is based on ModernBERT!), but I do think there's something underrated about slapping classification on a "big" model like I've seen post-Jev announcement, lots of Qwen stuff, but the most interesting to me so far is razorback16/openjev using DiffusionGemma. There's a level of generalization that lots and lots of parameters get you that you can't really get out of small models.
      1. NitpickLawyer · · focus · HN ↗
        > using DiffusionGemma.

        That's an interesting choice. One question I had when looking at the jev copy on their blog is if one "line" in their output looks / attends to other lines. I think not, since they say it's parallel and not autoregressive. In that regard, it would be interesting to play with diffusion, and see if you'd get better results by playing with types, locking some, and so on.

        1. robrenaud · · focus · HN ↗
          > That's an interesting choice. One question I had when looking at the jev copy on their blog is if one "line" in their output looks / attends to other lines. I think not, since they say it's parallel and not autoregressive.

          I don't understand the connection between the lack of autoregression and options attending to each other.

          Non autoregressive models can attend to all the inputs simultanously.

          An autogregressive model can can attend to all the options in the context of each other by simply writing the options out twice. Autoregressive models actually requires this, since one of them will come later, and the earlier prefill inputs can't attend to the later ones.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.