‹ BackHN Continuity

Thread

OpenJev

722 points · 296 comments · ilreb

  1. tecleandor · · focus · HN ↗
    I'm confused... This has no relation with the Jev team, isn't it?

    It's trying to "emulate" Jev behavior using a regular small LLM model (Qwen3 0.6B or MiniCPM5 2B). And with the smallest model it takes like between half to two seconds to run in my M2 Max, so it's not super fast.

    I mean, it's faster than asking to a regular LLM, but I think that's not proper to have Jev on the name (also legally...)

    Edit: no shade, and I'll give it a try for some ideas. I'd also like to have an open weights Jev but I think the naming is misguiding. I also have to try Jev that, BTW, got access pretty quickly, less than a day I think...

    1. CharlieDigital · · focus · HN ↗
      OP's point here is that the overall approach of restricting output token space and using parallel prompts to produce concurrent results and taking the most relevant ones isn't something novel to Jev (not saying there's nothing novel, but a facsimile can be created at the application layer using any small, fast model)
      1. Foobar8568 · · focus · HN ↗
        I still don't get the point of jev....it's basically an optimized models/runner on really short context and output?
        1. orbital-decay · · focus · HN ↗
          It's a specialized classifier model. It classifies input text into categories with a confidence score. Usually those classifiers are small like in the OP but jev is supposedly big, smart, and fast enough to play DOOM by having the scene described in text and classifying it into button presses.
          1. Foobar8568 · · focus · HN ↗
            Well the Doom demo is again passing a textual structure....I am not really convinced on how it&#x27;s different than any other llm that execute small context within 100ms. On a MBP M3Max with LFM 2.5B, I get about 500ms -600ms on &quot;source_text&quot;: &quot;Invoice #4471 issued March 3, 2026 to Beaver Dam Logistics for $12,840.00, net 30.&quot; with a 4 property structure output <a href="https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;primitives&#x2F;advanced" rel="nofollow">https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;primitives&#x2F;advanced

            I can&#x27;t test it on a better model &#x2F; my main workstation, but sub 1sec for short prompts is not impressive? I am sure that we can get something like 100ms-300ms with a Qwen 3.8 27b model for a similar query on a 5090 class GPU.

            edit: 203ms wall clock on a somewhat busy workstation with <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;LilaRest&#x2F;gemma-4-31B-it-NVFP4-turbo" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;LilaRest&#x2F;gemma-4-31B-it-NVFP4-turbo

          2. MrYanMYN · · focus · HN ↗
            It is mostly Harness hype. People actually explore the capabilities of classifier models which up until this point weren&#x27;t touched. You can recreate most of those with LFM 2.5 classifier locally
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.