Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Unofficial Hacker News client; not affiliated with Y Combinator.
adrithmetiqa · · focus · HN ↗
jubilanti · · focus · HN ↗
Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?
edit: Where I'm coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don't need anything cheaper or faster than that.
est · · focus · HN ↗
As a closed source chat-API provider, you just need to find a way speak JSON correctly at API output.
ed_mercer · · focus · HN ↗
fzysingularity · · focus · HN ↗
But both LLMs and Jev-like models would need to prefill, the only optimization Jev does differently is the decode which can be emulated by reading off logprobs.
We don’t know the param size of Jev, to determine the most comparable model, but if I had to guess it’s sub-100B.
sampullman · · focus · HN ↗
fooker · · focus · HN ↗
Orders of magnitude faster and cheaper answers.
Sure a MacBook Pro can control a servo motor but maybe an arduino or a cheaper microcontroller for deployment?
bitpush · · focus · HN ↗
vengadanathan · · focus · HN ↗
not sure what made you think you cant go cheaper than jev. it is being done for a long time.
fooker · · focus · HN ↗
vengadanathan · · focus · HN ↗
fooker · · focus · HN ↗
Once you discover what the useful problem to solve is and how to utilize it in production, you can of course replicate it.
This is not alchemy taught by some wizard in secret.
Very often, the innovation is in identifying what to build, what has interesting use cases, what people will pay for.
Once this is established, there is further research optimizing it further, replicating it locally, etc.
The fact that anyone could have come up with it is irrelevant.
sheepscreek · · focus · HN ↗
nextaccountic · · focus · HN ↗
The same reason I want diffusion language models to be mainstream
fzysingularity · · focus · HN ↗
The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.
tclancy · · focus · HN ↗
fzysingularity · · focus · HN ↗
This looks pretty decent: <a href="https://www.aidancooper.co.uk/constrained-decoding/" rel="nofollow">https://www.aidancooper.co.uk/constrained-decoding/
phoghed · · focus · HN ↗
No comment on this model itself, might be over fitted to the Jev benchmarks just to beat it.
<a href="https://github.com/Mushroom-Systems/lichen" rel="nofollow">https://github.com/Mushroom-Systems/lichen
sroussey · · focus · HN ↗
Name Entity Recognition (NER) is one example.
So many of them... <a href="https://huggingface.co/models?language=ner&sort=trending" rel="nofollow">https://huggingface.co/models?language=ner&sort=trending
Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don't want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).
The nice thing about Jev is that people started taking about models that are not LLM text streams again.
anvuong · · focus · HN ↗
This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it's always stressed not to treat these sigmoid'ed or softmax'ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.
Lerc · · focus · HN ↗
That said, I think the advantage of Jev style approaches is not their capabilities, but rather the capabilities that they have for a much lower resource requirement.
peab · · focus · HN ↗
Nowadays any LLM and any harness you use will just do this for you.
But there are helper libraries like Instructor that have been around since like gpt3, which abstract away retries and stuff to make this super easy.
jrop · · focus · HN ↗