‹ BackHN Continuity

Thread

OpenJev

722 points · 296 comments · ilreb

  1. mmastrac · · focus · HN ↗
    If you want to try a _legit_ Jev implementation that matches (at least in my evals), the vLLM patch to turn DiffusionGemma into Jev is available.

    On my DGX Spark I get very similar latency numbers, and it matches my evals + or - a few points on each test (DG wins some, Jev wins some, both show low confidence when wrong).

    I ran the same evals against a Qwen36 and it clearly lost to both of them, so you are leaving both knowledge and instinctual reasoning on the table with any smaller models, FWIW.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250

    1. cmrdporcupine · · focus · HN ↗
      This PR is interesting but it&#x27;s making the assumption that what Jev has done is based on a diffusion model or that a diffusion model is superior for this work. Which may or may not be the case.

      If I understand it though it does mean you can evaluate a bunch of questions simultaneously, which is an advantage.

      Also: While I think it&#x27;s expected&#x2F;normal to see LLM-generated programs... there&#x27;s a lot of LLM written comments in that PR, which is sad to see. Auto-human.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.