Every major AI shop has a ton of in-house classifiers already, big, small, generalist, specialized. Some are used in inference pipelines (e.g. safeguards), some are used in data preparation, training, analysis and investigation, research, various one-off and intermediate tasks etc. Offering them on a public API doesn't always make business sense. I don't see much substance to this buzz, looks like people that are new to all this are discovering that classifiers exist, they are more efficient at classification, and many tasks commonly done with generative models are classification in disguise. Which is not bad at all, a fresh look at their use is great to have.
Correct me if I'm wrong, but a zero-shot classifier like Jev is fundamentally different to a classifier with a fixed task (e.g. for safeguards), unless they trained a general purpose system to complete the safeguard task, which seems unlikely.
But zero-shot classifiers with this level of intelligence, world knowledge, ergonomics, cost profile, and ease of use are new.
I feel like good engineering doesn't just ignore those things, or at least it didn't before recently. Now I guess social media has added a pressure to reduce everything to a hot take.
An LLM is a zero-shot classifier with a large number of classes. All you need to do is establish what the output means and you can fine-tune an LLM final layer for this task if you like (and others have done). A student of mine did this as an exercise two years ago, and it was cool, but not publishable.
I agree with you on the "ease of use" business though. No one thought to make this sort of thing commercially available.
But there is no hot take here. Jev is not some new paradigm; engineering-wise, it is a trivial modification to the existing pipeline. That doesn't mean it isn't commercially viable.
No, all they had to do was come up with a quality post-training recipe, production inference stack that wouldn't fall over, GTM, documentation, schemas, etc. etc.
(also most signs point to this being LLaDA 2.0-adjacent so throw in solving some substantial mid-training)
I think it's 100% a hot take to call what they built trivial. Or at least it used to be.
There was a time when that kind of stuff was something between sour grapes and cluelessness about the gap between an idea and an actual commercial product deployed at scale, but now that's just weirdly normalized.
In fact, if anything I'm the weirdo for repeatedly taking issue with the way people are trivializing it ¯\_(ツ)_/¯
Jev actually isn’t anywhere near the top. It even loses to open weight clones. This tells me that whatever their “calibration” dataset is, it doesn’t seem to be anything special.
You linked to some weird subtable that labeled: " Not the default — not the JevBench Score", that can only be reached after you see what I just linked... lmao are you really this hard up about things?
Also every single question (even in the hard set) is single dimensional?: <a href="https://github.com/fstandhartinger/jevbench/blob/main/datasets/public/hard.jsonl" rel="nofollow">https://github.com/fstandhartinger/jevbench/blob/main/datase...
Jeeze, this is getting sad. I guess after all the mass-psychoses where people thought pointless things are going to change the world, we were due for a mass-psychosis where something interesting just has to be pointless?
Because I was specifically responding to your claim that Jev’s training recipe would give it better accuracy than others. It doesn’t have better accuracy than others. You could do as well or better by distilling qwen for example.
Jev is ranked higher than others on the overall benchmark due to speed and/or cost, not accuracy.
orbital-decay · · focus · HN ↗
bigmadshoe · · focus · HN ↗
janalsncm · · focus · HN ↗
BoorishBears · · focus · HN ↗
I feel like good engineering doesn't just ignore those things, or at least it didn't before recently. Now I guess social media has added a pressure to reduce everything to a hot take.
hodgehog11 · · focus · HN ↗
I agree with you on the "ease of use" business though. No one thought to make this sort of thing commercially available.
But there is no hot take here. Jev is not some new paradigm; engineering-wise, it is a trivial modification to the existing pipeline. That doesn't mean it isn't commercially viable.
BoorishBears · · focus · HN ↗
(also most signs point to this being LLaDA 2.0-adjacent so throw in solving some substantial mid-training)
I think it's 100% a hot take to call what they built trivial. Or at least it used to be.
There was a time when that kind of stuff was something between sour grapes and cluelessness about the gap between an idea and an actual commercial product deployed at scale, but now that's just weirdly normalized.
In fact, if anything I'm the weirdo for repeatedly taking issue with the way people are trivializing it ¯\_(ツ)_/¯
janalsncm · · focus · HN ↗
<a href="https://benchmarkheaven.com/jev-models?w=100-0-0-0#jevc-weights" rel="nofollow">https://benchmarkheaven.com/jev-models?w=100-0-0-0#jevc-weig...
Jev actually isn’t anywhere near the top. It even loses to open weight clones. This tells me that whatever their “calibration” dataset is, it doesn’t seem to be anything special.
BoorishBears · · focus · HN ↗
<a href="https://benchmarkheaven.com/jev-models" rel="nofollow">https://benchmarkheaven.com/jev-models
You linked to some weird subtable that labeled: " Not the default — not the JevBench Score", that can only be reached after you see what I just linked... lmao are you really this hard up about things?
Also every single question (even in the hard set) is single dimensional?: <a href="https://github.com/fstandhartinger/jevbench/blob/main/datasets/public/hard.jsonl" rel="nofollow">https://github.com/fstandhartinger/jevbench/blob/main/datase...
Jeeze, this is getting sad. I guess after all the mass-psychoses where people thought pointless things are going to change the world, we were due for a mass-psychosis where something interesting just has to be pointless?
janalsncm · · focus · HN ↗
Jev is ranked higher than others on the overall benchmark due to speed and/or cost, not accuracy.
brokencode · · focus · HN ↗
Obviously it’s the speed and cost that make it compelling. The tradeoff is accuracy.
Enough to matter? Maybe, maybe not. It’s not like it’s way down the chart. It’s probably good enough for a lot of tasks.