‹ BackHN Continuity

Thread

I built non-autoregressive decision models with RL a year ago

1363 points · 319 comments · nandakishor_ml

  1. Oras · · focus · HN ↗
    I played around with Jev last night and did it for classification tasks that I used Gemini 2.5 flash lite with.

    It’s a bit faster and bit cheaper, but this is compared to LLM. The consistency was nice to see, BUT, as someone who trained NLP models prior to LLMs, it’s just BERT with more data. I can see why people would want ready made one shot classifier, and I can see the value of sending multiple classifier in one call, but I wouldn’t call it breakthrough. And I believe many labs will replicate it in no time and might have it as part of their harness.

    I see it as a wake up call for the tech community to go back to basics for most tasks instead of relying solely on generic LLMs.

    1. tchalla · · focus · HN ↗
      Anyone who has worked in ML for 10+ years would already know that the usage of LLMs for everything is lazy, wasteful and a high degree of marketing on it.
      1. dominotw · · focus · HN ↗
        why would you waste your time messing around with a team of expensive ml engineers and data scientists that produce vastly inferior to a llm.

        We ripped out custom homegrown ml models that were developed in last 10 yrs and put an llm in its place. Its the opposite of wasteful. Even local gemma models are vastly superior.

        1. tchalla · · focus · HN ↗
          There’s a middle option. Once you figure that out, you’d soon understand my point today or tomorrow. I’ve been in this field for 21 years and I use LLMs everyday. I also know when to not use them.
          1. ramses0 · · focus · HN ↗
            It's the transportation "mode shifting" difficulty. Per the AI, the term of art is "Pure Transfer Penalty". It's the "ick" when doing bike => bus => bike instead of "only bike" or "only car".

            Mode switching has a cost. Usually std::sort is good enough compared to picking the prime optimal algorithm for your expected shape. Just call the function and get on with your day.

          2. ithkuil · · focus · HN ↗
            I think both your arguments are true. It all depends on the velocity of the capability growth and the fact that opportunity cost is expensive.

            Once we get out of this hypergriwth phase the very same AI companies that now are giving you llms will provide a service that employed a rich mixture of optimized models that will reduce the operational costs to achieve the required results

          3. throwaway7783 · · focus · HN ↗
            Is the middle option asking LLM to generate a classic ML model? Or generate tons of them and pick the best?
          4. indymike · · focus · HN ↗
            I'm a little confused: LLMs were invented in 2018.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.