‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. AgentMasterRace · · focus · HN ↗
    I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.
    1. tbeseda · · focus · HN ↗
      For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

      I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.

      Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.

      1. nico · · focus · HN ↗
        Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier

        I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes

        The classifiers also run in <1ms, so they can be very fast and precise at the same time

        But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)

        For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data

        1. derefr · · focus · HN ↗
          What do you use to determine that a particular task in a heterogeneous pile of tasks requires reasoning? The logistic classifier itself is too dumb to recognize the details of the problem that make it reasoning-sensitive (IIRC recognizing the “fiddliness” of a given problem requires a recognizer at least as complex as the problem itself.) And if you’re using the lightweight LLM for that, then you may as well skip the classifier and just use the LLM all the time, since that eval step is already going to be dominating your response time anyway.

          My understanding of Jev is that it’s a replacement for the LLM you’d necessarily need to use to identify reasoning-sensitive workloads in a heterogeneous mix, where Jev will be cheaper than an actual LLM and so act as an actual optimization / de-bottlenecking change.

          1. nico · · focus · HN ↗
            I’m in the process of piecing together the different task/dataset-specific classifiers

            Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach

            For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.