Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Unofficial Hacker News client; not affiliated with Y Combinator.
adrithmetiqa · · focus · HN ↗
jubilanti · · focus · HN ↗
Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?
edit: Where I'm coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don't need anything cheaper or faster than that.
ed_mercer · · focus · HN ↗
fzysingularity · · focus · HN ↗
But both LLMs and Jev-like models would need to prefill, the only optimization Jev does differently is the decode which can be emulated by reading off logprobs.
We don’t know the param size of Jev, to determine the most comparable model, but if I had to guess it’s sub-100B.