‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. adrithmetiqa · · focus · HN ↗
    Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
    1. k__ · · focus · HN ↗
      No need, as that functionality can run locally no problem.
      1. make3 · · focus · HN ↗
        the appeal would be if they can deliver it at a much higher performance and similar speed, which is plausible
    2. bigyabai · · focus · HN ↗
      BeRT and FLAN-T5 were used as classifiers 5-7 years ago, they were technically "frontier" for their time.
      1. alanwreath · · focus · HN ↗
        This is the exact comment I’ve been waiting for, what is the difference between classifiers and jev?
        1. prometheus1992 · · focus · HN ↗
          nothing in utility. we used various bert variants to satisfy our usecases and still in uses. free and they run locally.
        2. yfontana · · focus · HN ↗
          Jev is a classifier. The big thing about it is that it has high accuracy on domains it wasn't fine-tuned for, like an LLM, but with speed and cost comparable to traditional classifiers.
        3. fra · · focus · HN ↗
          BERT need to be fine tuned for your use case, Jev generalizes. It’s a pretty big difference!
          1. latentsea · · focus · HN ↗
            So what you're saying is it's artificial... general... intelligence? /s
        4. make3 · · focus · HN ↗
          FLAN-T5 generated text (Jev does not generate text), and BERT wasn't able to do tasks without fine-tuning.

          Jev is basically a kind of FLAN-BERT, if you want, where it has built-in multi-task ability, but doesn't generate text. It only generates 255 floats all at once, making it much faster, and what those floats mean (if anything) depends on the prompt.

          Eg, the following query is put in the encoder model:

          {"question": "Rank these 5 things by increasing order of how big they are", "choices": ["truck", "cow", "mouse", "ant", "building"] }

          The model returns [3., 2., 1., 0., 4.], and 249 other meaningless floats that are hidden from you by the UI.

          The UI stitches the first 5 floats with the choices and returns something like:

          {"rank": ["ant", "mouse", "cow", "truck", "tower"]}

          1. dannyw · · focus · HN ↗
            So it generates logits in a 255 token output space? ;)
            1. make3 · · focus · HN ↗
              logits assumed some form of softmax or logistic, which may not be the case
          2. vlovich123 · · focus · HN ↗
            How does it know that you’re asking for “rank” instead of something else if it’s not generating text?
            1. jmalicki · · focus · HN ↗
              By reading, not generating, text.
              1. vlovich123 · · focus · HN ↗
                It has to output “rank” - that requires generation unless I’m mistaken
                1. make3 · · focus · HN ↗
                  no, 255 numbers come out at once all the time, the order is determined by the inputs
                  1. vlovich123 · · focus · HN ↗
                    That’s not what I’m asking about - the text prompted for rank and it output “rank” in the response object. How did it do that? Like if I’d asked it to group similar items, how would it know how to structure that output and know that the key should be “groupings”
                    1. pests · · focus · HN ↗

                      [dead]

                    2. make3 · · focus · HN ↗
                      The output key is determined through code by the harness from the inputs, and the model generates 255 floats in order that follow that schema by reading the expected return type "rank" in the input. The harness then programmatically uses the floats of the 255 floats that are useful. The harness can assume that the correct float will be in the correct position as the model is trained to follow the schemas
          3. aftbit · · focus · HN ↗
            Is Jev's architecture public somewhere? I'd love to read more about it.
            1. Rzor · · focus · HN ↗
              <a href="https:&#x2F;&#x2F;laya.convaiinnovations.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;laya.convaiinnovations.com&#x2F;
              1. aftbit · · focus · HN ↗
                That&#x27;s one of the copycats. They reproduced the shape of System One decisions, but who knows if they did it in the same way.
        5. _menelaus · · focus · HN ↗
          The fact that its not narrow and stupid is what&#x27;s different
    3. tbeseda · · focus · HN ↗
      If I had to guess, it&#x27;s already built and is just waiting on Product&#x27;s&#x2F;Marketing&#x27;s desk. How do you position this without looking like your roadmap is being determined by newcomers? Probably don&#x27;t want to adopt the same verbiage+acronyms - but also can&#x27;t be seen to be just sherlocking features.
      1. seizethecheese · · focus · HN ↗
        I think Apple has demonstrated that shipping second has essentially no negative impact if your product is seen as higher quality.
        1. bigyabai · · focus · HN ↗
          I think Nvidia has demonstrated that shipping first is a multi-trillion dollar opportunity if you don&#x27;t shy away from a challenge.
          1. lucideer · · focus · HN ↗
            I think it depends on the model: Nvidia ship products with APIs, apple&#x2F;jev&#x2F;etc. ship end user products. The former is much more subject to lock-in, increasing the value of early market adoption because there&#x27;s a 3P Nvidia ecosystem sprung up in response. Apple&#x2F;jev&#x2F;etc. do have APIs &amp; corresponding 3P ecosystems but those are usually a smaller component of market capture than direct product end users, so the space ends up more competitive.
            1. 8note · · focus · HN ↗
              jev ships an api, does it not?
              1. lucideer · · focus · HN ↗
                yup. I said that in my comment...
          2. throwaway27448 · · focus · HN ↗
            Chip manufacturing intrinsically comes with one hell of a moat. There&#x27;s not much parallel in software.
            1. slashdev · · focus · HN ↗
              That’s kind of funny because Nvidia’s biggest moat is arguably CUDA, the software ecosystem around their chips
              1. mcmcmc · · focus · HN ↗
                CUDA is a lock-in moat, the infrastructure needed for chip manufacturing is a barrier-to-entry moat. Two different things.
                1. angry_octet · · focus · HN ↗
                  There are many microarchitecture patents used in NVIDIA chips. I&#x27;m sure they have a team that rips apart AMD chips looking for infringement. The way CUDA works is tied to many GPU architecture decisions and it would be hard to decouple them efficiently. Obviously a huge effort was made to get PyTorch decoupled from CUDA.
                2. AtlasBarfed · · focus · HN ↗
                  If cuda is an API, and llms make apis effortless, then how big of a moat is cuda?
                  1. robflynn · · focus · HN ↗
                    I ran across this a few days ago: <a href="https:&#x2F;&#x2F;zluda.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;zluda.org&#x2F;
                    1. QuantumNomad_ · · focus · HN ↗
                      All of the buttons and links on that page redirect to spam pages. Most of the times I clicked, it brings me to some site that wants to sell me a VPN.
                      1. robflynn · · focus · HN ↗
                        Oh, yikes, I should&#x27;ve checked those links before posting it here. That&#x27;s certainly not a good look for them.

                        There was a github repo but I have not checked it.

                        edit I see, thats an unaffiliated site that latched onto that, my bad, here&#x27;s the GH that I should&#x27;ve linked: <a href="https:&#x2F;&#x2F;github.com&#x2F;vosen&#x2F;ZLUDA" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vosen&#x2F;ZLUDA

                    2. [deleted] · · focus · HN ↗

                      [deleted]

                  2. bobmarleybiceps · · focus · HN ↗
                    IMO, yes a lot of the &quot;nvidia pays lots of people to make non-portable, tightly coupled backends to open source projects&quot; is potentially going to be less of a moat?

                    (Though it could turn into &quot;nvidia pay lots of people to use LLMs to make non-portable, tightly coupled backends to _even more_ open source projects&quot;)

                  3. trollbridge · · focus · HN ↗
                    It turns out the moat is “writing drivers that work”; Nvidia drivers simply work, and Intel’s are poor quality. So if I want stuff that works I need to buy Nvidia gear.
                  4. latentsea · · focus · HN ↗
                    It&#x27;s becoming less of one. Previously I would have shied away from buying an AMD card because of CUDA, but with local LLMs getting good enough to be usable and frontier models becoming as good as they have, I bit the bullet and got an R9700 for local inference. Dealing with working around CUDA used to be more of a manual process, but when you can point an agent at it and get stuff working, it&#x27;s dramatically less painful and scary than it used to be. Plus, at least in ComfyUI and local LLMs I&#x27;m finding support for AMD has gotten really good. Lately I&#x27;ve been witnessing a lot of people using agents to write custom kernels for RDNA4 and improving performance dramatically.
                3. flyinglizard · · focus · HN ↗
                  I’m totally guessing, but I can’t imagine CUDA has any significance at the frontier lab scale. The operational and capex costs are so massive that the convenience of the platform becomes a minuscule consideration.

                  It’s just that Nvidia’s stuff works, and available at scale, and includes the full stack with networking, cooling and such.

              2. angry_octet · · focus · HN ↗
                Historically there was a big patent moat in (Graphics) GPU design. This continues with CUDA, but obviously Intel and AMD could find ways to support eg PyTorch. What we don&#x27;t know is how much effort that cost them, or why they decided they couldn&#x27;t make a CUDA API compatible competitor.
                1. trollbridge · · focus · HN ↗
                  Intel could have beat the pants off Nvidia a long time ago with Arc if they’d bothered to ship usable drivers. But they refuse to, and simply can’t figure it out, so Arc cards remain cheap because they’re so #%£€ing hard to get working well, and everyone is nervous they’ll lay off the driver team again.
                  1. api · · focus · HN ↗
                    AMD has always had software problems too. Hardware companies often devalue software and suck at it.

                    If you can make them work Arc cards are a massive bargain. On raw compute the silicon is not bad.

          3. [deleted] · · focus · HN ↗

            [deleted]

          4. AndrewKemendo · · focus · HN ↗
            NVIDIA literally hired all of the 3DFX team and patents after they already proved the graphics accelleration hardware market was massive with the Voodoo card line

            In fact NVIDIA wasn&#x27;t even a close competitor to 3DFX in the graphics card game at that point

            1. bigyabai · · focus · HN ↗
              It panned out great. 3DFX filed for bankruptcy less than 18 months later, and Nvidia could pivot from designing raster chips to considering CUDA&#x27;s architecture.

              It&#x27;s not like 3DFX was the first GPU vendor. Nvidia saw the opportunity to be the first true GPGPU vendor, and they beat their competitors.

              1. trollbridge · · focus · HN ↗
                IBM shipped the first PC GPU (the Image Adapter&#x2F;A, 1989) following on the first PC 2D accelerator (the 8514&#x2F;A, 1987). There was zero benefit to being first. 3dfx had the first mass market, cheap GPU in 1996.

                ATI (AMD) in 1989 copied the unpatentable parts of the 8514&#x2F;A, improved it, and went on to dominate 2D accelerators.

                nVidia’s first 3D card was a complete flop. They did not achieve success for years.

                Nvidia is in the right place at the right time.

                1. bigyabai · · focus · HN ↗
                  We&#x27;re talking about GPGPU products, not raster GPUs. I cleared that up pretty well in my last comment.
                  1. trollbridge · · focus · HN ↗
                    The IA&#x2F;A was general purpose.
                2. Keyframe · · focus · HN ↗
                  not to be _that_ guy, but S3 was the dominant force in 3D. ATI was nowhere to be seen with their Mach&#x2F;Rage crap. It was an S3 era, and then seemingly out of nowhere 3dfx swept in with Voodoo (not seemingly though - 3dfx came out from imploding SGI). Only a bit later Nvidia, after it recovered from NV1 fiasco, brute forced relentlessly from Riva 128 to TNT to TNT2 to Geforce 256 which broke 3dfx (along with their lack of business sense) and had Nvidia bought them. ATI did their 2D schtick only during that timeline while S3 stumbled with Virge and ATI didn&#x27;t come as a threat until Radeon; After they bought ArtX (which did Gamecube graphics) which turned into Radeon. S3 died off in that race with Savage3D and the only other major players were Matrox and 3dlabs. Matrox kind of found temp refuge in video segment, and 3dlabs infamously pushed for OpenGL 2 and survived for a bit on 3d workstations. The most surprising (IMO) was the downfall of E&amp;S which kickstarted most of the things mentioned. E&amp;S and Real3D are both a great story in themselves how first movers can become absolutely forgotten and obscure real quick (with Intel740 ended up in ATI).
                  1. trollbridge · · focus · HN ↗
                    S3’s stuff was no more “3D” than the Image Adapter&#x2F;A was.

                    You should look at the IA&#x2F;A - it had its own C like compiler, CPU, etc which did things reminiscent of a modern GPU or SIMD.

          5. PunchyHamster · · focus · HN ↗
            It&#x27;s more due to ineptitude of competition. If AMD shipped second, but better product it would be another thing but it is still a bit of a mess of an ecosystem on AMD side
            1. bigyabai · · focus · HN ↗
              It&#x27;s not AMD&#x27;s responsibility to dethrone Nvidia any more than it is Apple&#x27;s. AMD sells CDNA, but the momentum is with CUDA and Khronos can&#x27;t get anyone to sit at the same table anymore.
              1. hgoel · · focus · HN ↗
                It isn&#x27;t their responsibility, but it doesn&#x27;t make sense to argue that being first got NVIDIA a multitrillion dollar market if the others aren&#x27;t even trying to compete. There is no &quot;first&quot; if it&#x27;s really just &quot;only one even trying&quot;.

                The momentum is with CUDA because CUDA is the most broadly usable one. Especially with AI-driven optimization loops and similar APIs, competitors can more easily pick up momentum, if they&#x27;d actually try.

      2. michaelrwolfe1 · · focus · HN ↗
        There is zero stigma to shipping second. If anything, the labs’ customers are probably begging for them to add these features natively so they don’t have to deal with the hassle of adding another provider to their stack.
    4. [deleted] · · focus · HN ↗

      [deleted]

    5. Ohentis · · focus · HN ↗
      I don&#x27;t think there would be any utility for that. Anything jev can do, a frontier model can also do. Just not as quickly or as cheaply.
      1. tbeseda · · focus · HN ↗
        I think these products (Jev and the inevitable offerings from Anthropic, OpenAI, etc) want to become more than end-user output machines. They&#x27;d benefit from being in the hotpath of other services. Not backgrounded generation but in-band, request-time work.

        1M x $0.50 == 1B x $0.0005

        1. transitorykris · · focus · HN ↗
          To expand a bit for my current use cases. Inline routing of work to heavy task specific models, and prompt&#x2F;context generation (user is asking something, what and how much should we prompt the expensive LLM with). Latency or time to first token does matter for some applications.
        2. Ohentis · · focus · HN ↗

          [dead]

    6. onlyrealcuzzo · · focus · HN ↗
      Probably at the frontier stage - you will only see it where Jev is better regardless of cost.

      For everyone else who is conscious of cost, you&#x27;re already seeing this being built into harnesses.

      Almost certainly, you&#x27;ll see versions of this from all the Chinese labs as fast as humanly possible.

      If I had to guess, Cursor&#x2F;Grok or Google&#x2F;Antigravity will be the first major players to natively support something like this to drive down cost, as they&#x27;re primarily the budget conscious choices.

      I would be astounded if Anthropic leads the way on a cost reduction.

    7. themitchelli · · focus · HN ↗
      But for those of us that prefer open source and self hosting, JEV alternative LAYA will beat anything the frontier models package up.
    8. wgd · · focus · HN ↗
      Negative three years, give or take. Although recent Anthropic and OpenAI models no longer expose the capability. But for any open model you just tell it to respond with a single token &quot;Y&#x2F;N&quot; and take the logit difference. If you want multiple distinct questions answered you just ask them independently and put the shared context first so it gets cached.

      OpenAI and Anthropic don&#x27;t want to give out logprobs these days but could trivially add a dedicated classification API to their existing models if there was enough demand.

      1. Renaud · · focus · HN ↗
        I think the main differentiator offered by Jev is not the ability to answer questions, most models can be coerced into that function if they don’t already have a dedicated pipeline for it, rather it’s the extreme speed of the evaluation, and very low cost that opens new possibilities.
        1. wgd · · focus · HN ↗
          The evaluation is fast because it&#x27;s all prefill computation with only a single token of inference. Ditto cost, you&#x27;re paying 100% input costs and nearly zero output. There really isn&#x27;t any architectural magic to Jev, it&#x27;s just a straightforward application of normal LLM tech with some good marketing.

          I mean, Jev is also probably cheaper because it&#x27;s a rather small model (or at least, I suspect it is based on the overall level of intelligence it demonstrates) so that helps make it cheap too.

          1. psyphy2 · · focus · HN ↗
            Yes this is exactly my assumption as well. I think ppl forgot before agents it was expected slo to have a ttft in range of a few hundard ms, which is what jev is also achieving.

            Thats also why it feels weird they say they don&#x27;t &quot;charge for output tokens&quot; since its literally generating a single (or at most very few tokens).

        2. selcuka · · focus · HN ↗
          &gt; it’s the extreme speed of the evaluation, and very low cost that opens new possibilities.

          It also returns confidence scores for all choices.

          Granted, they are not stable. They fluctuate even when you reorder choices, but it still counts as an additional feature.

          1. wonnage · · focus · HN ↗
            They’re useable if you’re trying to rank autosuggestions or something, not great at implying actual understanding
    9. jubilanti · · focus · HN ↗
      Do people really have zero awareness that Structured Outputs with a constrained schema has been a thing for a while now, and open weight models that give you logprobs can give you distributions per key?

      Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?

      edit: Where I&#x27;m coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don&#x27;t need anything cheaper or faster than that.

      1. est · · focus · HN ↗
        You don&#x27;t even need a constrained decoding.

        As a closed source chat-API provider, you just need to find a way speak JSON correctly at API output.

      2. ed_mercer · · focus · HN ↗
        jev is optimized for it, a standard LLM isn&#x27;t and it also costlier and slower.
        1. fzysingularity · · focus · HN ↗
          When you say Jev is optimized for it, I get that there’s no need for unnecessary decodes in an autoregressive fashion.

          But both LLMs and Jev-like models would need to prefill, the only optimization Jev does differently is the decode which can be emulated by reading off logprobs.

          We don’t know the param size of Jev, to determine the most comparable model, but if I had to guess it’s sub-100B.

      3. sampullman · · focus · HN ↗
        Price and speed are the difference. But deepseek flash is usually fast and cheap enough, so the use cases are somewhat limited.
      4. fooker · · focus · HN ↗
        &gt; Like what am I missing?

        Orders of magnitude faster and cheaper answers.

        Sure a MacBook Pro can control a servo motor but maybe an arduino or a cheaper microcontroller for deployment?

        1. bitpush · · focus · HN ↗
          Excellent analogy
        2. vengadanathan · · focus · HN ↗
          No necessarily, we are getting similar result gemma 4 12b, predicting one output token for label. it is just prefill and with nvfp4 caliberated , we are getting 70-80 ms p90 for the classification workload we have.

          not sure what made you think you cant go cheaper than jev. it is being done for a long time.

          1. fooker · · focus · HN ↗
            Can you host files cheaper than Dropbox? Of course you can. :)
            1. vengadanathan · · focus · HN ↗
              Well if there is compliance &#x2F; data privacy need to it, it makes absolute sense. We operate in vocie ai industry and run in regulated environment. Jev etc doesnt support languages we want (indic languages) and for our usecase and it is not self hosted either, so running fine tuned LLM based classifier is more accurate and low latent (no external network hop as it runs in VPC). Idea is to re-use as much as possible but if compliance, accuracy becomes bottle neck it makes more sense to do it yourself.
              1. fooker · · focus · HN ↗
                Right.

                Once you discover what the useful problem to solve is and how to utilize it in production, you can of course replicate it.

                This is not alchemy taught by some wizard in secret.

                Very often, the innovation is in identifying what to build, what has interesting use cases, what people will pay for.

                Once this is established, there is further research optimizing it further, replicating it locally, etc.

                The fact that anyone could have come up with it is irrelevant.

      5. sheepscreek · · focus · HN ↗
        Many people yes, but it’s probably for the best. Without fine-tuning on such style the results would be unreliable. It wouldn’t be the probability of the outcome of whatever you intend, but just the next token probability - which could alter if you add a space or punctuation to the prompt. Very flaky.
      6. nextaccountic · · focus · HN ↗
        Pricing

        The same reason I want diffusion language models to be mainstream

      7. fzysingularity · · focus · HN ↗
        FWIW you’re absolutely correct on how most developers are unaware of this as they’re mostly operating on the API level and not aware of server side (vLLM&#x2F;SGLang) capabilities.

        The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.

        1. tclancy · · focus · HN ↗
          Can you link to some info on this? I am just starting to poke around in the space and would love a leg up.
          1. fzysingularity · · focus · HN ↗
            I don’t know of one definitive resource but “constrained decoding”, regex&#x2F;grammar-based decoding, json-schema decoding will give you a bunch of hits. Look up vLLM, SGLang and outlines’ implementations for more technical details.

            This looks pretty decent: <a href="https:&#x2F;&#x2F;www.aidancooper.co.uk&#x2F;constrained-decoding&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.aidancooper.co.uk&#x2F;constrained-decoding&#x2F;

          2. phoghed · · focus · HN ↗
            This is an example implementation that was making the rounds recently. If you point whatever model you prefer at it, it’ll do a good job of explaining it.

            No comment on this model itself, might be over fitted to the Jev benchmarks just to beat it.

            <a href="https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen

        2. sroussey · · focus · HN ↗
          Classifiers (instead of LLMs) return results with confidence scores.

          Name Entity Recognition (NER) is one example.

          So many of them... <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;models?language=ner&amp;sort=trending" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;models?language=ner&amp;sort=trending

          Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don&#x27;t want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).

          The nice thing about Jev is that people started taking about models that are not LLM text streams again.

          1. anvuong · · focus · HN ↗
            &gt; Classifiers (instead of LLMs) return results with confidence scores.

            This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it&#x27;s always stressed not to treat these sigmoid&#x27;ed or softmax&#x27;ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.

      8. Lerc · · focus · HN ↗
        I remember reading a paper about a model that used a parser bound to the end of a LLM where it squashed all logprobs for outputs that were parsing errors. I have not seen this used in the manner I thought I would. I thought it would have been entirely possible for a model to construct it&#x27;s own grammar for how it would prefer to respond and then opt to generate tokens that matched (with a metacode to turn it off obviously).

        That said, I think the advantage of Jev style approaches is not their capabilities, but rather the capabilities that they have for a much lower resource requirement.

      9. peab · · focus · HN ↗
        Yeah, I&#x27;m with you.

        Nowadays any LLM and any harness you use will just do this for you.

        But there are helper libraries like Instructor that have been around since like gpt3, which abstract away retries and stuff to make this super easy.

      10. jrop · · focus · HN ↗
        Yeah I though llama.cpp had this during the very early days, if my memory serves me correctly.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.