‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. AgentMasterRace · · focus · HN ↗
    I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.
    1. tbeseda · · focus · HN ↗
      For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

      I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.

      Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.

      1. senko · · focus · HN ↗
        > The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

        You can already so that with classification models such as ModernBERT, at 0.4B.

        Jev's value is its zero shot performance without having to fine-tune.

        1. beepbooptheory · · focus · HN ↗
          I am sure I am missing something obvious here, but why is that valuable? Like, what kinds of projects are there where you need to classify stuff but are unable to make a bespoke model targeting the specific problem?
          1. shaewest · · focus · HN ↗
            For my org, it meant we could trial classifiers across various internal systems with little to no engineering effort. In one case we ended up building our own classifier instead of Jev, but in others we kept Jev because it was zero-effort for a great impact.
            1. beepbooptheory · · focus · HN ↗
              But like how often are you gonna do this in general? Why does dev time or effort really matter here when either way you are building something to just, you know, actually use going forward? Its not like one needs to build a new classifier everyday.
              1. tyre · · focus · HN ↗
                Because a lot of people employed as engineers can’t build classifiers. They’re able to glue libraries and tools together, build UIs, and write APIs, but don’t have the curiosity or creative problem solving to learn and master a new domain. Even when “master” is scoped to something like this.

                On top, most EMs wouldn’t take a risk on an exploration of something “unknown” (to them) and couldn’t get buy-in from a PM.

                I say this as an EM. Interview hundreds of people and, while, yes, some people don’t interview well, you might be shocked at the level of creative thinking. Even when “creative” is narrowly scoped to “this is a solved problem in a related domain.

                1. beepbooptheory · · focus · HN ↗
                  I guess this makes sense. I am no business guy, but if my product/company was focused on some sort of classification problem, my naive intuition would be to focus on hiring guys that can do it, rather than try to make the problem easier for them. But perhaps at the end of the day this is still cheaper? It's the same reason why we use Postgres instead of hire database experts to build something special?
                  1. senko · · focus · HN ↗
                    > if my product/company was focused on some sort of classification problem

                    It's more likely the product is focused on something else, but a classification model could come in handy...

          2. jmalicki · · focus · HN ↗
            > unable to make a bespoke model targeting the specific problem?

            There is a fixed cost (and some maintenance) to e.g. fine tuning ModernBERT.

            Maybe once you include all of that it might be a half-day to a day of engineering time to set everything up in a maintainable fashion.

            For Jev, it takes all of 30 seconds of prompting. And it's not that much more expensive to deploy vs. a BERT model.

            1. benterix · · focus · HN ↗
              The biggest disadvantage of Jev is that it's a proprietary product and you need to send them your data. A bespoke solution makes much more sense in many scenarios.
          3. vickychijwani · · focus · HN ↗
            Tons. For example, most web and mobile app developers won’t know where to start with making a bespoke model (and would likely have no interest in making one), but they will have lots of usecases for a classifier.
            1. retinaros · · focus · HN ↗
              its a 1 hour project on claude code on ur local machine. lol...
              1. vickychijwani · · focus · HN ↗
                Only for folks who already know what they’re doing.

                What I said will only make sense if you take yourself out of your current context and think entirely from the perspective of someone who knows little-to-nothing about ML.

                It’s the same mistake folks on HN made when Dropbox launched, drawing comparisons to rsync and other Unix tools as if they were somehow equivalent.

          4. xigoi · · focus · HN ↗
            Anything where you don’t have a decent amount of training data.
          5. senko · · focus · HN ↗
            You need to have enough high quality data to train with, knowledge how to do it, and developer time.

            In practice, that's enough of a barrier to not even try the approach on a number of cases where it might potentially be useful.

            I wouldn't be surprised if Jev turned out to be a "gateway drug" that validates approach on a use case, the team gathers experience and labeled data, and switches to an in house locally tuned model to minimize costs.

            1. davrosthedalek · · focus · HN ↗
              Exactly. And that labeled data could be collected by just recording what they feed jev and what the decision is.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.