Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Unofficial Hacker News client; not affiliated with Y Combinator.
AgentMasterRace · · focus · HN ↗
tbeseda · · focus · HN ↗
I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.
Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.
clhodapp · · focus · HN ↗
nico · · focus · HN ↗
I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes
The classifiers also run in <1ms, so they can be very fast and precise at the same time
But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)
For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data
rahimnathwani · · focus · HN ↗
Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?
If so, I'm curious whether you compared that approach (A) with:
B) Jev only, with a single output.
C) Jev with multiple outputs fed into a logistic classifier.
Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)
nico · · focus · HN ↗
Here's a gist with code you can use to test the Banking77 dataset: <a href="https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecdad5029557" rel="nofollow">https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller
I haven’t compared different ways of sending requests to Jev
The data to train the classifiers comes from the datasets used to test them (not from Jev)
derefr · · focus · HN ↗
My understanding of Jev is that it’s a replacement for the LLM you’d necessarily need to use to identify reasoning-sensitive workloads in a heterogeneous mix, where Jev will be cheaper than an actual LLM and so act as an actual optimization / de-bottlenecking change.
nico · · focus · HN ↗
Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach
For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time
jmalicki · · focus · HN ↗
dcl · · focus · HN ↗
senko · · focus · HN ↗
You can already so that with classification models such as ModernBERT, at 0.4B.
Jev's value is its zero shot performance without having to fine-tune.
beepbooptheory · · focus · HN ↗
shaewest · · focus · HN ↗
beepbooptheory · · focus · HN ↗
tyre · · focus · HN ↗
On top, most EMs wouldn’t take a risk on an exploration of something “unknown” (to them) and couldn’t get buy-in from a PM.
I say this as an EM. Interview hundreds of people and, while, yes, some people don’t interview well, you might be shocked at the level of creative thinking. Even when “creative” is narrowly scoped to “this is a solved problem in a related domain.
beepbooptheory · · focus · HN ↗
senko · · focus · HN ↗
It's more likely the product is focused on something else, but a classification model could come in handy...
jmalicki · · focus · HN ↗
There is a fixed cost (and some maintenance) to e.g. fine tuning ModernBERT.
Maybe once you include all of that it might be a half-day to a day of engineering time to set everything up in a maintainable fashion.
For Jev, it takes all of 30 seconds of prompting. And it's not that much more expensive to deploy vs. a BERT model.
benterix · · focus · HN ↗
vickychijwani · · focus · HN ↗
retinaros · · focus · HN ↗
vickychijwani · · focus · HN ↗
What I said will only make sense if you take yourself out of your current context and think entirely from the perspective of someone who knows little-to-nothing about ML.
It’s the same mistake folks on HN made when Dropbox launched, drawing comparisons to rsync and other Unix tools as if they were somehow equivalent.
xigoi · · focus · HN ↗
senko · · focus · HN ↗
In practice, that's enough of a barrier to not even try the approach on a number of cases where it might potentially be useful.
I wouldn't be surprised if Jev turned out to be a "gateway drug" that validates approach on a use case, the team gathers experience and labeled data, and switches to an in house locally tuned model to minimize costs.
davrosthedalek · · focus · HN ↗
gjs278 · · focus · HN ↗
[dead]
baobabKoodaa · · focus · HN ↗
Oras · · focus · HN ↗
I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff
I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.
I suppose with larger model, this could be useful, but would require more ram and will be slower.
zergrush · · focus · HN ↗
they all suck
sharih · · focus · HN ↗