Every major AI shop has a ton of in-house classifiers already, big, small, generalist, specialized. Some are used in inference pipelines (e.g. safeguards), some are used in data preparation, training, analysis and investigation, research, various one-off and intermediate tasks etc. Offering them on a public API doesn't always make business sense. I don't see much substance to this buzz, looks like people that are new to all this are discovering that classifiers exist, they are more efficient at classification, and many tasks commonly done with generative models are classification in disguise. Which is not bad at all, a fresh look at their use is great to have.
I fed into the hype at first. Testing Jev and Laya, they both suffer from the same issues as LLMs that stop them being useful beyond limited classifications.
I can't see any benefits that a typical ML classifier would not be better at.
I think the main argument would just be that because the model is general, you don't need to retrain it from scratch for a new problem - just tweak the input prompt. For a typical classifier there's a lot more hassle - collecting the data, training it yourself, retraining under distribution shift... In that sense Jev seems great for prototyping or small-scale use cases.
Counterargument: this works for quick prototyping, but for any serious business, you will eventually develop a benchmark/eval to track how well the general model is working, and once you have that dataset, you might as well train a specific model
Jev's bet is that if it works well enough for random use cases that nobody complains, then management won't feel a need to develop a benchmark/eval, and they won't need to employ all those data science guys.
I'd also add that they're hoping Jevon's Paradox also leads to a whole new segment of users who would have never reached for a classifier in the first place, given the barrier to entry.
And if you do get complaints or feedback on the classification, have a dev log into the user's account, tweak the Jev prompt a little until the issue goes away, and push it to production
Yes this is what I'm interested in. I think they might be right. I'm already finding myself thinking "well maybe a classifier would be useful here now that it's so easy to do...".
This probably just means that I could have been reaching for that tool more often already. But in practice I wasn't, and this has opened my eyes to the potential opportunities there.
Or not. And replace the generalist with the next generalist that gets you +15% on that benchmark for the same price, or gives you the same benchmark performance for half the price.
One advantage of using generalist models is that the generalists are improving - regardless of whether you're doing anything about it.
Yes, but the generalists are not routinely improving across all domains. The large labs are really focusing on agentic use, so I imagine that creative writing has deteriorated considering how distinctive Claude's writing style has become. Or I recently had an image-parsing task, and I was excited to try Qwen because I heard it had gotten a lot better at agentic tasks, but it failed my image-parsing benchmark.
There are focus areas, but capabilities improve across all domains - some slower than others. "Agentic use" is in itself a very general thing - because many tasks benefit from being able to leverage adaptive model-driven workflows.
Creative writing and Claude - amusing that you say that, given that Anthropic just went and tried to unfuck it in Opus 5.5 specifically. It is an example of a capability no one typically cares about, yes. No money in creative writing. But even there, we had gains in newer models.
orbital-decay · · focus · HN ↗
EagnaIonat · · focus · HN ↗
I can't see any benefits that a typical ML classifier would not be better at.
ainch · · focus · HN ↗
firejake308 · · focus · HN ↗
woah · · focus · HN ↗
momojo · · focus · HN ↗
woah · · focus · HN ↗
what · · focus · HN ↗
But makes issues for someone (or everyone) else?
sanderjd · · focus · HN ↗
This probably just means that I could have been reaching for that tool more often already. But in practice I wasn't, and this has opened my eyes to the potential opportunities there.
ACCount39 · · focus · HN ↗
One advantage of using generalist models is that the generalists are improving - regardless of whether you're doing anything about it.
firejake308 · · focus · HN ↗
ACCount39 · · focus · HN ↗
Creative writing and Claude - amusing that you say that, given that Anthropic just went and tried to unfuck it in Opus 5.5 specifically. It is an example of a capability no one typically cares about, yes. No money in creative writing. But even there, we had gains in newer models.