‹ BackHN Continuity

Thread

OpenAI is well positioned to fast-follow Jev

328 points · 233 comments · JohnBerryman

  1. orbital-decay · · focus · HN ↗
    Every major AI shop has a ton of in-house classifiers already, big, small, generalist, specialized. Some are used in inference pipelines (e.g. safeguards), some are used in data preparation, training, analysis and investigation, research, various one-off and intermediate tasks etc. Offering them on a public API doesn't always make business sense. I don't see much substance to this buzz, looks like people that are new to all this are discovering that classifiers exist, they are more efficient at classification, and many tasks commonly done with generative models are classification in disguise. Which is not bad at all, a fresh look at their use is great to have.
    1. JohnBerryman · · focus · HN ↗
      For me, I think the big deal is that it promises to be general and broadly applicable and high quality. That's new and special. But we'll wait to see if the claims actually hold.
    2. EagnaIonat · · focus · HN ↗
      I fed into the hype at first. Testing Jev and Laya, they both suffer from the same issues as LLMs that stop them being useful beyond limited classifications.

      I can't see any benefits that a typical ML classifier would not be better at.

      1. edot · · focus · HN ↗
        Agreed. I tested Jev on OpenRouter this past weekend and it’s “okay” but a specific classifier is significantly better. It used to require skill to import sklearn (ok, not really), but now it’s literally one prompt and upload your Excel file or whatever and you can get your classifier out. It’ll run free, instant, more accurate.
        1. boostermodule · · focus · HN ↗
          This is predicated on you having training data already. I approach Jev more like Langchain -- you can prototype something new extremely fast and cheap, and if the use case works well enough, rip it out and build something bespoke. If it doesn't, you didn't spend a bunch of time curating a training dataset anyway.
          1. sanderjd · · focus · HN ↗
            Yeah I think that's right. It's actually nice to have a better-than-nothing placeholder that can be replaced if it becomes valuable to do so.
      2. tomrod · · focus · HN ↗
        Prompt ingestion is going to be the biggest differentiator.

        Being able to route prompt to features that then route to special models would be a really solid implementation.

        1. EagnaIonat · · focus · HN ↗
          It starts to break down once you go over 20 classifications. Which is very basic routing that can easily be done with typical ML models for cheaper and faster.
          1. tomrod · · focus · HN ↗
            Thanks for the breadcrumb!
      3. ainch · · focus · HN ↗
        I think the main argument would just be that because the model is general, you don't need to retrain it from scratch for a new problem - just tweak the input prompt. For a typical classifier there's a lot more hassle - collecting the data, training it yourself, retraining under distribution shift... In that sense Jev seems great for prototyping or small-scale use cases.
        1. firejake308 · · focus · HN ↗
          Counterargument: this works for quick prototyping, but for any serious business, you will eventually develop a benchmark/eval to track how well the general model is working, and once you have that dataset, you might as well train a specific model
          1. woah · · focus · HN ↗
            Jev's bet is that if it works well enough for random use cases that nobody complains, then management won't feel a need to develop a benchmark/eval, and they won't need to employ all those data science guys.
            1. momojo · · focus · HN ↗
              I'd also add that they're hoping Jevon's Paradox also leads to a whole new segment of users who would have never reached for a classifier in the first place, given the barrier to entry.
              1. woah · · focus · HN ↗
                And if you do get complaints or feedback on the classification, have a dev log into the user's account, tweak the Jev prompt a little until the issue goes away, and push it to production
                1. what · · focus · HN ↗
                  > tweak the Jev prompt a little until the issue goes away

                  But makes issues for someone (or everyone) else?

              2. sanderjd · · focus · HN ↗
                Yes this is what I'm interested in. I think they might be right. I'm already finding myself thinking "well maybe a classifier would be useful here now that it's so easy to do...".

                This probably just means that I could have been reaching for that tool more often already. But in practice I wasn't, and this has opened my eyes to the potential opportunities there.

          2. ACCount39 · · focus · HN ↗
            Or not. And replace the generalist with the next generalist that gets you +15% on that benchmark for the same price, or gives you the same benchmark performance for half the price.

            One advantage of using generalist models is that the generalists are improving - regardless of whether you're doing anything about it.

            1. firejake308 · · focus · HN ↗
              Yes, but the generalists are not routinely improving across all domains. The large labs are really focusing on agentic use, so I imagine that creative writing has deteriorated considering how distinctive Claude's writing style has become. Or I recently had an image-parsing task, and I was excited to try Qwen because I heard it had gotten a lot better at agentic tasks, but it failed my image-parsing benchmark.
              1. ACCount39 · · focus · HN ↗
                There are focus areas, but capabilities improve across all domains - some slower than others. "Agentic use" is in itself a very general thing - because many tasks benefit from being able to leverage adaptive model-driven workflows.

                Creative writing and Claude - amusing that you say that, given that Anthropic just went and tried to unfuck it in Opus 5.5 specifically. It is an example of a capability no one typically cares about, yes. No money in creative writing. But even there, we had gains in newer models.

        2. EagnaIonat · · focus · HN ↗
          Training a classification model is trivial these days, even for a number far bigger than what Jev can do.
        3. ygouzerh · · focus · HN ↗
          I think here it's mostly that for normal business cases, we doesn't need to build one.

          As a DevOps Engineer, I never once saw before the advantage of using a classifier. Now I see multiple parts of the stack where a better level of expressiveness will be useful (PR validations, Blue/Green validation, notification router for alerts, quick smoke tests, etc).

          Nobody will give us the time and budget to build a custom classifier for these use cases, but a simple API call yes.

      4. ricardobeat · · focus · HN ↗
        Using Jev as a plain classifier is the least interesting case. See robotic control, navigation, computer use examples, none of it possible with a classifier.
        1. orbital-decay · · focus · HN ↗
          That's the point, they're classification in disguise. Agentic game engines/mods started doing this long ago due to the latency requirements (although they're typically using small BERT-like models that need to be finetuned, or low TTFT generative models and structured outputs). New or newly discovered use cases are great, sure.
      5. sanderjd · · focus · HN ↗
        I guess I'm circling toward this view. The question is, are there things that are 1. worth doing, 2. for which jev (or jev-like systems) works well, and 3. are not worth the effort to train a custom classifier. Probably yes, but it seems like it might be a pretty narrow path. But a lot depends on #2. The trade-off between #1 and #3 is less stark the more successful one shot models are at handling use cases successfully.
        1. aDyslecticCrow · · focus · HN ↗
          Scripts and debugging, one-off log parsing or filtering.

          I saw an article about 2+ years ago of a researcher using a small local AI strapped into excel to evaluate the abstract and intro of 10000 papers for "papers that research X in domain of Y", and let it loose.

          jev is probably more capable avd faster than that workflow was, but saved one dude a few very grindy weeks for a litteratur review.

          It's amusing how long it took, and much hype it gets for someone releasing the least revolutionary ML architecture in a new package. But i can see a fair few uses.

          1. sanderjd · · focus · HN ↗
            Yeah this seems right to me.
      6. nextaccountic · · focus · HN ↗
        Jev is prompted with natural language, so it is flexible and a good fit to replace subagents for certain tasks

        A LLM agent could be trained to use jev effectively as a tool call, even (but even without specific RL they do a good job already)

    3. [deleted] · · focus · HN ↗

      [deleted]

    4. Razengan · · focus · HN ↗
      If your "master AI" is good enough, it should be able to find and learn about and use specialized tech AI like Jev if it suits your goals

      and then whatever tech it is will be absorbed/assimilated/Sherlocked into the leading products anyway

    5. andriy_koval · · focus · HN ↗
      > Every major AI shop has a ton of in-house classifiers already, big, small, generalist

      I think building generalist classifier is some open ended research task, where frontier labs can contribute: different internal reasoning, instruction tuning, building datasets and benchmarks, building and distilling super large models.

    6. bigmadshoe · · focus · HN ↗
      Correct me if I'm wrong, but a zero-shot classifier like Jev is fundamentally different to a classifier with a fixed task (e.g. for safeguards), unless they trained a general purpose system to complete the safeguard task, which seems unlikely.
      1. mmis1000 · · focus · HN ↗
        Fixed guard today is not very fixed. For ex, the safeguard qwen released is a full 4b llm model. It has no different to normal llm model arch except tuned for this specific purpose,
        1. bigmadshoe · · focus · HN ↗
          So it is tuned specifically to classify content for safeguarding? I'm not familiar with this particular model, but it most likely has a specific classifier head that is tuned for the safeguard task. This is completely different to zero-shot classification.
      2. janalsncm · · focus · HN ↗
        Correct, but zero-shot classifiers are also not new.
        1. BoorishBears · · focus · HN ↗
          But zero-shot classifiers with this level of intelligence, world knowledge, ergonomics, cost profile, and ease of use are new.

          I feel like good engineering doesn't just ignore those things, or at least it didn't before recently. Now I guess social media has added a pressure to reduce everything to a hot take.

          1. catlifeonmars · · focus · HN ↗
            They almost certainly would perform worse than more specialized classifiers trained with less data. It’s kind of a paradox of generalization. I think there’s an interesting space where you use generalized models to generate ad hoc specialized classifiers.
            1. aDyslecticCrow · · focus · HN ↗
              Depends what you man by "more specialised". You wont train very good language understanding without alot of data. It probably uses the core tranformer stack from an LLM.

              Classic classifiers are regularly just tuned general models; Training a CCN on ImageNet and tune it for cats and dogs gives better results than just training it on cats and dogs.

              There is likley a small network used to tranform model output vector to probabilities, but that wouldn't be massive. Retraining that small network for specific task may beat jev; but that's bairly considered training by modern standards.

            2. BoorishBears · · focus · HN ↗
              Expecting a strong zero-shot performer to perform worse in a low data regime?

              That only makes sense if you try to rope in data previously used to establish the model's priors, but that wouldn't make sense in this context. That same additional data is what enables things like...

              > use generalized models to generate ad hoc specialized classifiers.

            3. senordevnyc · · focus · HN ↗
              They almost certainly would perform worse than more specialized classifiers trained with less data. It’s kind of a paradox of generalization.

              Isn't this exactly what the bitter lesson is about?

          2. aDyslecticCrow · · focus · HN ↗
            > Ergonomics, cost profile, and ease of use are new.

            Following AI from the academic papers side; jev really feels silly. They one-pass the LLM tranformer stack and tune the output network for a probability value.

            (some clever pararellization optimisations to make it viable to offer as an api, since the normal kv cashing no longer works if you oneshot the tranformer)

            The largest change is the packaging; An api with a tolken based pricing, and a schema to define the output structure for quick setup.

            Previous projects would probably involve installing pytorch, running a converter script on Qwen, and write a fair bit of matrix math to change the output shape.

            I'm kinda amused that it took this long though.

          3. hodgehog11 · · focus · HN ↗
            An LLM is a zero-shot classifier with a large number of classes. All you need to do is establish what the output means and you can fine-tune an LLM final layer for this task if you like (and others have done). A student of mine did this as an exercise two years ago, and it was cool, but not publishable.

            I agree with you on the "ease of use" business though. No one thought to make this sort of thing commercially available.

            But there is no hot take here. Jev is not some new paradigm; engineering-wise, it is a trivial modification to the existing pipeline. That doesn't mean it isn't commercially viable.

            1. BoorishBears · · focus · HN ↗
              No, all they had to do was come up with a quality post-training recipe, production inference stack that wouldn't fall over, GTM, documentation, schemas, etc. etc.

              (also most signs point to this being LLaDA 2.0-adjacent so throw in solving some substantial mid-training)

              I think it's 100% a hot take to call what they built trivial. Or at least it used to be.

              There was a time when that kind of stuff was something between sour grapes and cluelessness about the gap between an idea and an actual commercial product deployed at scale, but now that's just weirdly normalized.

              In fact, if anything I'm the weirdo for repeatedly taking issue with the way people are trivializing it ¯\_(ツ)_/¯

              1. janalsncm · · focus · HN ↗
                Someone else posted the jevbench site which compares jev to a bunch of other models. If you look only at the accuracy dimension:

                <a href="https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models?w=100-0-0-0#jevc-weights" rel="nofollow">https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models?w=100-0-0-0#jevc-weig...

                Jev actually isn’t anywhere near the top. It even loses to open weight clones. This tells me that whatever their “calibration” dataset is, it doesn’t seem to be anything special.

                1. BoorishBears · · focus · HN ↗
                  ... why didn&#x27;t you link to the actual benchmark which does have Jev at the top?

                  <a href="https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models" rel="nofollow">https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models

                  You linked to some weird subtable that labeled: &quot; Not the default — not the JevBench Score&quot;, that can only be reached after you see what I just linked... lmao are you really this hard up about things?

                  Also every single question (even in the hard set) is single dimensional?: <a href="https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench&#x2F;blob&#x2F;main&#x2F;datasets&#x2F;public&#x2F;hard.jsonl" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench&#x2F;blob&#x2F;main&#x2F;datase...

                  Jeeze, this is getting sad. I guess after all the mass-psychoses where people thought pointless things are going to change the world, we were due for a mass-psychosis where something interesting just has to be pointless?

                  1. janalsncm · · focus · HN ↗
                    Because I was specifically responding to your claim that Jev’s training recipe would give it better accuracy than others. It doesn’t have better accuracy than others. You could do as well or better by distilling qwen for example.

                    Jev is ranked higher than others on the overall benchmark due to speed and&#x2F;or cost, not accuracy.

                    1. brokencode · · focus · HN ↗
                      Well yeah, that’s the whole idea. If speed and cost don’t matter, you could use Astra.

                      Obviously it’s the speed and cost that make it compelling. The tradeoff is accuracy.

                      Enough to matter? Maybe, maybe not. It’s not like it’s way down the chart. It’s probably good enough for a lot of tasks.

                2. ombansod · · focus · HN ↗

                  [dead]

              2. hodgehog11 · · focus · HN ↗
                Architecturally, it is trivial. That&#x27;s something the community would have consensus on, so not a hot take.

                I see your point, but Jev doesn&#x27;t exist in a vacuum. When one (like me) says &quot;trivial&quot;, they mean it relative to other attempts and developments in the field, all of which require everything you&#x27;ve mentioned at minimum. Commercialising any product, and doing it well, is hard. But the R&amp;D factor here is substantially more straightforward than almost any other product in its category, because there is no architectural breakthrough here.

                1. BoorishBears · · focus · HN ↗
                  &quot;all of which require everything you&#x27;ve mentioned at minimum&quot;

                  Sorry who else did everything I mentioned? I think the guy behind Laya tried after noticing Jev&#x27;s traction... but the site&#x27;s auth went down and has stayed down for a day now.

                  &quot;substantially more straightforward than almost any other product in its category&quot;

                  More straightforward than the spite projects based on constrained decoding? Or Laya with it&#x27;s couple of days post-training ModernBERT?

                  -

                  I have no doubt other teams can build models like this and I&#x27;ve love for a frontier lab to give us an even smarter model with these ergonomics... but in the rush to show Jev what&#x27;s up, we&#x27;re mostly getting slop.

                  PS: I don&#x27;t know anyone who&#x27;s done anything of note who uses trivial like that. The commentariat do, and the &quot;I could have done that&quot; crowd do, but I don&#x27;t pay much attention to them until they actually do the thing.

                  1. hodgehog11 · · focus · HN ↗
                    By category, I meant other language models in general. The point of others putting something up to beat Jev is to show that, to date, no one has bothered to produce something like Jev, because anyone with decent LLM experience can roll their own for purpose with little effort and have been doing so for years. And can beat it on any metric you choose.

                    Let me put it this way. OpenAI and Anthropic have a slight moat over the Chinese labs because they have strong training data and the most advanced RL strategies. It will take the Chinese labs significant R&amp;D effort to bridge that, especially in math (and there is a good chance they will, provided they want to).

                    Jev has no moat other than the fact that no one else has bothered to package a model in this way. Another lab could build a strong competitor very quickly if they want to put the effort in. That&#x27;s the point of this post. There is no uncertainty about what they have done, nothing to figure out. Someone just needs to do it. I&#x27;m not sure what to say if you can&#x27;t see the difference between the two. Jev is worth celebrating because of the idea to package it in this way. But it is not a paradigm shift and that is likely a problem for them.

            2. ozgung · · focus · HN ↗
              &gt; All you need to do is establish what the output means and you can fine-tune an LLM final layer for this task

              Yeah, that was the original idea with GPT, Generative Pre-trained Transformer, and earlier open pre-trained transformers.

              Today people use AI via APIs rather then fine-tuning models by themselves and when someone provides this as an API they got excited.

              1. hodgehog11 · · focus · HN ↗
                No it wasn&#x27;t. Those were models trained from scratch, required large scale data, and the nontrivial parts involved training at scale and the autoregressive task which no one expected to work as well as it does. It is the difference between developing a foundation model, and using one. I believe Jev falls in the latter category, because the task itself is no different, only the output.
      3. zxexz · · focus · HN ↗
        This whole thing reminds me of DeepMind’s Variational Bayesian Last Layers[0], which never gained much traction in the broader “AI” world, but is a remarkably useful tool. And a relatively obvious one that anyone with experience in SVI and with transformer pretraining, seems to independently rediscover (including me) before finding this paper.

        [0] <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2404.11599" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2404.11599

      4. BiteCode_dev · · focus · HN ↗
        The generalist aspect of jev is what is good about it. Old classifiers tended to be specialized, not good at ambiquity or limited.

        LLM were used as classifiers because they solved that.

        Jev have the flexibility of LLM and the perf and api of classifiers.

    7. bluejay2387 · · focus · HN ↗
      &quot;I don&#x27;t see much substance to this buzz...&quot;

      Agreed. This isn&#x27;t new. I led a research team at a Fortune 500 that used a transformer based classifier approach in a commercial product as far back as 2022 and we didn&#x27;t come up with it. It was already common enough that we found the inspiration for our implementation on some web forum. Models like RouteLLM have been around for a long time. The news here isn&#x27;t that a new model type came about, its that a large percentage of people messing around with this stuff that are new to AI just learned that not all transformer based implementations need to be autoregressive.

      1. zer00eyz · · focus · HN ↗
        &gt; Agreed. This isn&#x27;t new. ...

        It doesn&#x27;t have to be new, it just has to be consumable by devs.

        You could send text before Twilio. You could process credit cards before Stripe.

        Jev, at the end of the day is an easy to use API.

        Everyone seems to forget that usability is a thing.

        1. bluejay2387 · · focus · HN ↗
          It took us a few hours to implement that one we used in 2022. This isn&#x27;t about usability its about a huge population working on this stuff not really knowing what is available until it becomes a meme.
    8. gwern · · focus · HN ↗
      Entertainingly, OpenAI had a general purpose zero-shot classifier API built on GPT-3! Just no one ever cared that much about it, so I guess it got dropped somewhere along the way since 2020&#x2F;2021.
    9. 0x20cowboy · · focus · HN ↗
      This. It’s machine learning vs. “AI” for the uninitiated. Soon there will be a new ground breaking model that does k-means clustering and will get a billon dollar funding (but only if you&#x27;re young and live in SF)

      The good news is it’s fun to see people discover and get excited about things that I like as well.

      1. aDyslecticCrow · · focus · HN ↗
        Its in a modern and easily to deploy package. The hype is a bit wierd. &quot;0 cost output tolkens&quot; is such a silly phrasing.

        I would have never considered importing pytorch for filtering through log files before even knowing my way around it. But if i can type a filtering condition by text and hit enter; i may actually use that to save some time.

        Id want something local though, but thats hardly a difficult demand for what it is.

    10. d2ou · · focus · HN ↗
      People in my lab (sklearn people) developped something that I feel close to jev but focused on tabular data : <a href="https:&#x2F;&#x2F;tabicl.readthedocs.io&#x2F;en&#x2F;latest&#x2F;" rel="nofollow">https:&#x2F;&#x2F;tabicl.readthedocs.io&#x2F;en&#x2F;latest&#x2F;

      This is a transformer based classifier with massive pretraining on synthetic datasets and it outperforms boosting classifiers on many benchmarks without the need of more gradient descent steps (the forward pass on X_train, y_train IS the training).

      I understand that jev focus on text entry. But I feel that it is a similar kind of model but trained on text. Did someone test it on tabular data as well ?

    11. rf15 · · focus · HN ↗
      You may say that, but every time I hinted at this in the past I just got downvoted to oblivion. The average LLM enjoyer was not aware of this.
    12. dbbk · · focus · HN ↗
      Not many people know this but OpenAI even has a FREE separate moderation model and API endpoint that can classify user generated content.
    13. minhhai22091 · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.