‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. adrithmetiqa · · focus · HN ↗
    Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
    1. jubilanti · · focus · HN ↗
      Do people really have zero awareness that Structured Outputs with a constrained schema has been a thing for a while now, and open weight models that give you logprobs can give you distributions per key?

      Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?

      edit: Where I'm coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don't need anything cheaper or faster than that.

      1. fzysingularity · · focus · HN ↗
        FWIW you’re absolutely correct on how most developers are unaware of this as they’re mostly operating on the API level and not aware of server side (vLLM/SGLang) capabilities.

        The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.

        1. tclancy · · focus · HN ↗
          Can you link to some info on this? I am just starting to poke around in the space and would love a leg up.
          1. fzysingularity · · focus · HN ↗
            I don’t know of one definitive resource but “constrained decoding”, regex/grammar-based decoding, json-schema decoding will give you a bunch of hits. Look up vLLM, SGLang and outlines’ implementations for more technical details.

            This looks pretty decent: <a href="https:&#x2F;&#x2F;www.aidancooper.co.uk&#x2F;constrained-decoding&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.aidancooper.co.uk&#x2F;constrained-decoding&#x2F;

          2. phoghed · · focus · HN ↗
            This is an example implementation that was making the rounds recently. If you point whatever model you prefer at it, it’ll do a good job of explaining it.

            No comment on this model itself, might be over fitted to the Jev benchmarks just to beat it.

            <a href="https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen

        2. sroussey · · focus · HN ↗
          Classifiers (instead of LLMs) return results with confidence scores.

          Name Entity Recognition (NER) is one example.

          So many of them... <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;models?language=ner&amp;sort=trending" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;models?language=ner&amp;sort=trending

          Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don&#x27;t want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).

          The nice thing about Jev is that people started taking about models that are not LLM text streams again.

          1. anvuong · · focus · HN ↗
            &gt; Classifiers (instead of LLMs) return results with confidence scores.

            This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it&#x27;s always stressed not to treat these sigmoid&#x27;ed or softmax&#x27;ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.