‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. adrithmetiqa · · focus · HN ↗
    Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
    1. jubilanti · · focus · HN ↗
      Do people really have zero awareness that Structured Outputs with a constrained schema has been a thing for a while now, and open weight models that give you logprobs can give you distributions per key?

      Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?

      edit: Where I'm coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don't need anything cheaper or faster than that.

      1. est · · focus · HN ↗
        You don't even need a constrained decoding.

        As a closed source chat-API provider, you just need to find a way speak JSON correctly at API output.

      2. ed_mercer · · focus · HN ↗
        jev is optimized for it, a standard LLM isn't and it also costlier and slower.
        1. fzysingularity · · focus · HN ↗
          When you say Jev is optimized for it, I get that there’s no need for unnecessary decodes in an autoregressive fashion.

          But both LLMs and Jev-like models would need to prefill, the only optimization Jev does differently is the decode which can be emulated by reading off logprobs.

          We don’t know the param size of Jev, to determine the most comparable model, but if I had to guess it’s sub-100B.

      3. sampullman · · focus · HN ↗
        Price and speed are the difference. But deepseek flash is usually fast and cheap enough, so the use cases are somewhat limited.
      4. fooker · · focus · HN ↗
        > Like what am I missing?

        Orders of magnitude faster and cheaper answers.

        Sure a MacBook Pro can control a servo motor but maybe an arduino or a cheaper microcontroller for deployment?

        1. bitpush · · focus · HN ↗
          Excellent analogy
        2. vengadanathan · · focus · HN ↗
          No necessarily, we are getting similar result gemma 4 12b, predicting one output token for label. it is just prefill and with nvfp4 caliberated , we are getting 70-80 ms p90 for the classification workload we have.

          not sure what made you think you cant go cheaper than jev. it is being done for a long time.

          1. fooker · · focus · HN ↗
            Can you host files cheaper than Dropbox? Of course you can. :)
            1. vengadanathan · · focus · HN ↗
              Well if there is compliance / data privacy need to it, it makes absolute sense. We operate in vocie ai industry and run in regulated environment. Jev etc doesnt support languages we want (indic languages) and for our usecase and it is not self hosted either, so running fine tuned LLM based classifier is more accurate and low latent (no external network hop as it runs in VPC). Idea is to re-use as much as possible but if compliance, accuracy becomes bottle neck it makes more sense to do it yourself.
              1. fooker · · focus · HN ↗
                Right.

                Once you discover what the useful problem to solve is and how to utilize it in production, you can of course replicate it.

                This is not alchemy taught by some wizard in secret.

                Very often, the innovation is in identifying what to build, what has interesting use cases, what people will pay for.

                Once this is established, there is further research optimizing it further, replicating it locally, etc.

                The fact that anyone could have come up with it is irrelevant.

      5. sheepscreek · · focus · HN ↗
        Many people yes, but it’s probably for the best. Without fine-tuning on such style the results would be unreliable. It wouldn’t be the probability of the outcome of whatever you intend, but just the next token probability - which could alter if you add a space or punctuation to the prompt. Very flaky.
      6. nextaccountic · · focus · HN ↗
        Pricing

        The same reason I want diffusion language models to be mainstream

      7. fzysingularity · · focus · HN ↗
        FWIW you’re absolutely correct on how most developers are unaware of this as they’re mostly operating on the API level and not aware of server side (vLLM/SGLang) capabilities.

        The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.

        1. tclancy · · focus · HN ↗
          Can you link to some info on this? I am just starting to poke around in the space and would love a leg up.
          1. fzysingularity · · focus · HN ↗
            I don’t know of one definitive resource but “constrained decoding”, regex/grammar-based decoding, json-schema decoding will give you a bunch of hits. Look up vLLM, SGLang and outlines’ implementations for more technical details.

            This looks pretty decent: <a href="https:&#x2F;&#x2F;www.aidancooper.co.uk&#x2F;constrained-decoding&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.aidancooper.co.uk&#x2F;constrained-decoding&#x2F;

          2. phoghed · · focus · HN ↗
            This is an example implementation that was making the rounds recently. If you point whatever model you prefer at it, it’ll do a good job of explaining it.

            No comment on this model itself, might be over fitted to the Jev benchmarks just to beat it.

            <a href="https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen

        2. sroussey · · focus · HN ↗
          Classifiers (instead of LLMs) return results with confidence scores.

          Name Entity Recognition (NER) is one example.

          So many of them... <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;models?language=ner&amp;sort=trending" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;models?language=ner&amp;sort=trending

          Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don&#x27;t want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).

          The nice thing about Jev is that people started taking about models that are not LLM text streams again.

          1. anvuong · · focus · HN ↗
            &gt; Classifiers (instead of LLMs) return results with confidence scores.

            This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it&#x27;s always stressed not to treat these sigmoid&#x27;ed or softmax&#x27;ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.

      8. Lerc · · focus · HN ↗
        I remember reading a paper about a model that used a parser bound to the end of a LLM where it squashed all logprobs for outputs that were parsing errors. I have not seen this used in the manner I thought I would. I thought it would have been entirely possible for a model to construct it&#x27;s own grammar for how it would prefer to respond and then opt to generate tokens that matched (with a metacode to turn it off obviously).

        That said, I think the advantage of Jev style approaches is not their capabilities, but rather the capabilities that they have for a much lower resource requirement.

      9. peab · · focus · HN ↗
        Yeah, I&#x27;m with you.

        Nowadays any LLM and any harness you use will just do this for you.

        But there are helper libraries like Instructor that have been around since like gpt3, which abstract away retries and stuff to make this super easy.

      10. jrop · · focus · HN ↗
        Yeah I though llama.cpp had this during the very early days, if my memory serves me correctly.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.