Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Unofficial Hacker News client; not affiliated with Y Combinator.
nico · · focus · HN ↗
For emails, I get 95% accuracy with this method, with only 50-100 examples for training
Training the model takes less than 5 minutes on a CPU
The resulting model is <1MB, and inference is sub 100ms
Some other cool things about this approach:
* the model doesn’t train on some “ideal” or general classification, instead it learns your preferences
* the model runs on pretty much any mobile device and can be retrained online on the device
* privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model)
Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).
rgbrgb · · focus · HN ↗
nico · · focus · HN ↗
The type of task in which it does really well, especially against Laya, is classification with >50 classes
But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)
For the latter cases, you could probably enhance the architecture with a lightweight LLM, something like a Gemma model. Or even some basic MLP
samuel · · focus · HN ↗
This is the same route but WAAAY faster and cheaper. And you can modify it like you do with code or prompts. It's really appealing, TBH.
0x457 · · focus · HN ↗
Originally it was so I can label data to fine-tune a VLM, but now a few tiny classifiers that run in milliseconds on cpu.
Now its collecting data to make a domain specific BERT and do what Jev does.
nico · · focus · HN ↗
Also curious about if you plan on doing some sort of routing for the requests. Like detecting the type of task to decide which model to route the request to
0x457 · · focus · HN ↗
This whole thing started because I wanted something to help me play Dune Imperium. Even relatively large models with vision encoders couldn't reliably extract the full state of the board. Now that I have ~2k labeled screenshots, I want to train heads on top of SigLIP2 to extract all of that data in one go.
That's how it started. Now the thing supports multiple kinds of datasets:
A model router isn't planned because I'm trying to gear everything toward self-hosting, and there just isn't that much to route between. I’ll probably build something Jev-like for smart-home control, though.The FastApply dataset is already ~20k entries, with the majority of outputs being 8k–16k tokens. The STT dataset is roughly 30 hours and growing.
Basically, the whole thing has turned into a Collect -> Distill -> Train pipeline for whatever I happen to need.
nico · · focus · HN ↗
Amazing, thank you for sharing your setup. Very cool applications
exographicskip · · focus · HN ↗
I didn't know I wanted this! Can just imagine asking how to make a Trinidad Sour and being called an insect lmao
0x457 · · focus · HN ↗
I had a version that on validation dataset nearly 1:1 across entire spectrum, but I ended up linking this revision more.
Also had qwen3-tts version, but no matter what I do it sounded more like brat Cortana. The way piper-tts work under the hood make it sound more machine-like.
andai · · focus · HN ↗
swyx · · focus · HN ↗
this misses the point of jev somewhat - the point is that this is a foundational, general purpose classifier model - see some good sources <a href="https://x.com/mparakhin/status/2101683565520199887?s=12" rel="nofollow">https://x.com/mparakhin/status/2101683565520199887?s=12
0x457 · · focus · HN ↗
mjyoke1111 · · focus · HN ↗
[dead]
constantlm · · focus · HN ↗
tk90 · · focus · HN ↗
30KB model, 40-50ms inference. Pretty happy with the results so far!
I can see an entire industry of tiny models like this, now that we have AI to help us do the grunt setup work (validation/training data creation, data cleaning, etc). Or just use a general classifier like Jev/Kev ha
ozozozd · · focus · HN ↗
What’s the model architecture?
bicepjai · · focus · HN ↗
Flere-Imsaho · · focus · HN ↗
e12e · · focus · HN ↗
I'm assuming you only need to consider a single language for your emails?
nico · · focus · HN ↗
Haven’t tried with more
Do you have a specific use case?
e12e · · focus · HN ↗
In my case Norwegian, English and Japanese.
dotancohen · · focus · HN ↗
psadri · · focus · HN ↗
These days, I tend to start my coding sessions by the high level problem I'm trying to solve vs the prescriptive, specific solution I may have in mind. It often surfaces ideas and approaches that I did not know about.
dotancohen · · focus · HN ↗
roseway4 · · focus · HN ↗
Our resulting RBF models are tiny and fit in L1 cache, with microsecond inference latency.
foofoobar · · focus · HN ↗
nico · · focus · HN ↗
<a href="https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecdad5029557" rel="nofollow">https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
The <10MB does not include the embeddings encoder
The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller