Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Unofficial Hacker News client; not affiliated with Y Combinator.
adrithmetiqa · · focus · HN ↗
k__ · · focus · HN ↗
make3 · · focus · HN ↗
bigyabai · · focus · HN ↗
alanwreath · · focus · HN ↗
prometheus1992 · · focus · HN ↗
yfontana · · focus · HN ↗
fra · · focus · HN ↗
latentsea · · focus · HN ↗
make3 · · focus · HN ↗
Jev is basically a kind of FLAN-BERT, if you want, where it has built-in multi-task ability, but doesn't generate text. It only generates 255 floats all at once, making it much faster, and what those floats mean (if anything) depends on the prompt.
Eg, the following query is put in the encoder model:
{"question": "Rank these 5 things by increasing order of how big they are", "choices": ["truck", "cow", "mouse", "ant", "building"] }
The model returns [3., 2., 1., 0., 4.], and 249 other meaningless floats that are hidden from you by the UI.
The UI stitches the first 5 floats with the choices and returns something like:
{"rank": ["ant", "mouse", "cow", "truck", "tower"]}
dannyw · · focus · HN ↗
make3 · · focus · HN ↗
vlovich123 · · focus · HN ↗
jmalicki · · focus · HN ↗
vlovich123 · · focus · HN ↗
make3 · · focus · HN ↗
vlovich123 · · focus · HN ↗
pests · · focus · HN ↗
[dead]
make3 · · focus · HN ↗
aftbit · · focus · HN ↗
Rzor · · focus · HN ↗
aftbit · · focus · HN ↗
_menelaus · · focus · HN ↗
tbeseda · · focus · HN ↗
seizethecheese · · focus · HN ↗
bigyabai · · focus · HN ↗
lucideer · · focus · HN ↗
8note · · focus · HN ↗
lucideer · · focus · HN ↗
throwaway27448 · · focus · HN ↗
slashdev · · focus · HN ↗
mcmcmc · · focus · HN ↗
angry_octet · · focus · HN ↗
AtlasBarfed · · focus · HN ↗
robflynn · · focus · HN ↗
QuantumNomad_ · · focus · HN ↗
robflynn · · focus · HN ↗
There was a github repo but I have not checked it.
edit I see, thats an unaffiliated site that latched onto that, my bad, here's the GH that I should've linked: <a href="https://github.com/vosen/ZLUDA" rel="nofollow">https://github.com/vosen/ZLUDA
[deleted] · · focus · HN ↗
[deleted]
bobmarleybiceps · · focus · HN ↗
(Though it could turn into "nvidia pay lots of people to use LLMs to make non-portable, tightly coupled backends to _even more_ open source projects")
trollbridge · · focus · HN ↗
latentsea · · focus · HN ↗
flyinglizard · · focus · HN ↗
It’s just that Nvidia’s stuff works, and available at scale, and includes the full stack with networking, cooling and such.
angry_octet · · focus · HN ↗
trollbridge · · focus · HN ↗
api · · focus · HN ↗
If you can make them work Arc cards are a massive bargain. On raw compute the silicon is not bad.
[deleted] · · focus · HN ↗
[deleted]
AndrewKemendo · · focus · HN ↗
In fact NVIDIA wasn't even a close competitor to 3DFX in the graphics card game at that point
bigyabai · · focus · HN ↗
It's not like 3DFX was the first GPU vendor. Nvidia saw the opportunity to be the first true GPGPU vendor, and they beat their competitors.
trollbridge · · focus · HN ↗
ATI (AMD) in 1989 copied the unpatentable parts of the 8514/A, improved it, and went on to dominate 2D accelerators.
nVidia’s first 3D card was a complete flop. They did not achieve success for years.
Nvidia is in the right place at the right time.
bigyabai · · focus · HN ↗
trollbridge · · focus · HN ↗
Keyframe · · focus · HN ↗
trollbridge · · focus · HN ↗
You should look at the IA/A - it had its own C like compiler, CPU, etc which did things reminiscent of a modern GPU or SIMD.
PunchyHamster · · focus · HN ↗
bigyabai · · focus · HN ↗
hgoel · · focus · HN ↗
The momentum is with CUDA because CUDA is the most broadly usable one. Especially with AI-driven optimization loops and similar APIs, competitors can more easily pick up momentum, if they'd actually try.
michaelrwolfe1 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
Ohentis · · focus · HN ↗
tbeseda · · focus · HN ↗
1M x $0.50 == 1B x $0.0005
transitorykris · · focus · HN ↗
Ohentis · · focus · HN ↗
[dead]
onlyrealcuzzo · · focus · HN ↗
For everyone else who is conscious of cost, you're already seeing this being built into harnesses.
Almost certainly, you'll see versions of this from all the Chinese labs as fast as humanly possible.
If I had to guess, Cursor/Grok or Google/Antigravity will be the first major players to natively support something like this to drive down cost, as they're primarily the budget conscious choices.
I would be astounded if Anthropic leads the way on a cost reduction.
themitchelli · · focus · HN ↗
wgd · · focus · HN ↗
OpenAI and Anthropic don't want to give out logprobs these days but could trivially add a dedicated classification API to their existing models if there was enough demand.
Renaud · · focus · HN ↗
wgd · · focus · HN ↗
I mean, Jev is also probably cheaper because it's a rather small model (or at least, I suspect it is based on the overall level of intelligence it demonstrates) so that helps make it cheap too.
psyphy2 · · focus · HN ↗
Thats also why it feels weird they say they don't "charge for output tokens" since its literally generating a single (or at most very few tokens).
selcuka · · focus · HN ↗
It also returns confidence scores for all choices.
Granted, they are not stable. They fluctuate even when you reorder choices, but it still counts as an additional feature.
wonnage · · focus · HN ↗
jubilanti · · focus · HN ↗
Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?
edit: Where I'm coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don't need anything cheaper or faster than that.
est · · focus · HN ↗
As a closed source chat-API provider, you just need to find a way speak JSON correctly at API output.
ed_mercer · · focus · HN ↗
fzysingularity · · focus · HN ↗
But both LLMs and Jev-like models would need to prefill, the only optimization Jev does differently is the decode which can be emulated by reading off logprobs.
We don’t know the param size of Jev, to determine the most comparable model, but if I had to guess it’s sub-100B.
sampullman · · focus · HN ↗
fooker · · focus · HN ↗
Orders of magnitude faster and cheaper answers.
Sure a MacBook Pro can control a servo motor but maybe an arduino or a cheaper microcontroller for deployment?
bitpush · · focus · HN ↗
vengadanathan · · focus · HN ↗
not sure what made you think you cant go cheaper than jev. it is being done for a long time.
fooker · · focus · HN ↗
vengadanathan · · focus · HN ↗
fooker · · focus · HN ↗
Once you discover what the useful problem to solve is and how to utilize it in production, you can of course replicate it.
This is not alchemy taught by some wizard in secret.
Very often, the innovation is in identifying what to build, what has interesting use cases, what people will pay for.
Once this is established, there is further research optimizing it further, replicating it locally, etc.
The fact that anyone could have come up with it is irrelevant.
sheepscreek · · focus · HN ↗
nextaccountic · · focus · HN ↗
The same reason I want diffusion language models to be mainstream
fzysingularity · · focus · HN ↗
The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.
tclancy · · focus · HN ↗
fzysingularity · · focus · HN ↗
This looks pretty decent: <a href="https://www.aidancooper.co.uk/constrained-decoding/" rel="nofollow">https://www.aidancooper.co.uk/constrained-decoding/
phoghed · · focus · HN ↗
No comment on this model itself, might be over fitted to the Jev benchmarks just to beat it.
<a href="https://github.com/Mushroom-Systems/lichen" rel="nofollow">https://github.com/Mushroom-Systems/lichen
sroussey · · focus · HN ↗
Name Entity Recognition (NER) is one example.
So many of them... <a href="https://huggingface.co/models?language=ner&sort=trending" rel="nofollow">https://huggingface.co/models?language=ner&sort=trending
Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don't want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).
The nice thing about Jev is that people started taking about models that are not LLM text streams again.
anvuong · · focus · HN ↗
This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it's always stressed not to treat these sigmoid'ed or softmax'ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.
Lerc · · focus · HN ↗
That said, I think the advantage of Jev style approaches is not their capabilities, but rather the capabilities that they have for a much lower resource requirement.
peab · · focus · HN ↗
Nowadays any LLM and any harness you use will just do this for you.
But there are helper libraries like Instructor that have been around since like gpt3, which abstract away retries and stuff to make this super easy.
jrop · · focus · HN ↗