‹ BackHN Continuity

Thread

OpenJev

722 points · 296 comments · ilreb

  1. mmastrac · · focus · HN ↗
    If you want to try a _legit_ Jev implementation that matches (at least in my evals), the vLLM patch to turn DiffusionGemma into Jev is available.

    On my DGX Spark I get very similar latency numbers, and it matches my evals + or - a few points on each test (DG wins some, Jev wins some, both show low confidence when wrong).

    I ran the same evals against a Qwen36 and it clearly lost to both of them, so you are leaving both knowledge and instinctual reasoning on the table with any smaller models, FWIW.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250

    1. mungoman2 · · focus · HN ↗
      This is very interesting! Seems like a promising direction.

      I wonder though if it supports the same claims as Jev: answers are not impacted by other answers to the same questions, nor the existance of other questions? It seems by sharing KV cache all questions will be visible. And I think the diffusion causes the answers to attend to eachother?

      Also I think without fine-tuning we can not say that the probabilities it output are actually probabilities. Maybe fine tuning using Brier scoring would do the trick?

      Maybe some kind of hierarchical structure of the KV cache can make questions independent, and with smaller diffusion canvas’ generated in batch can be a way to make answers generate independently?

      1. cmrdporcupine · · focus · HN ↗
        &gt; It seems by sharing KV cache all questions will be visible

        Yeah, this is partially why in my approach I&#x27;ve done this instead, and not used diffusion model:

        1. Convert the state into one shared prompt.

        2. Run that shared prompt through the model once.

        3. Fork the model’s internal state once per question.

        4. Add a different question to each fork.

        5. Ask each fork for its next-token scores.

        6. Calculate only 64 possible label scores—not the whole vocabulary.

        Basically ... skip decode.

        Won&#x27;t be as fast as doing diffusion model parallel across a pile of questions at once, but:

        a) let&#x27;s you use pretty much any existing text model (with some modifications). I&#x27;ve got qwen3.6 moe and qwen3.8 flash next running, am getting gemma4 working now

        b) the problem you identified

        It&#x27;s possible I&#x27;m getting high on my own supply and misunderstand entirely the whole thing, but it seems to work?

        <a href="https:&#x2F;&#x2F;github.com&#x2F;rdaum&#x2F;eider&#x2F;commit&#x2F;9c2d5c049068c33da2a48bc9f131f4b4dd868b92" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;rdaum&#x2F;eider&#x2F;commit&#x2F;9c2d5c049068c33da2a48b...

        I don&#x27;t have the chutzpah to go creating PRs for vLLM to do the same.

        1. aaquibahm · · focus · HN ↗
          why not mask attention and do it all in one forward pass ? tokens belonging to a question can just see that question and the main prompt
          1. cmrdporcupine · · focus · HN ↗
            Ok I tried it. And it&#x27;s faster. Though I&#x27;m still in verifying quality phase.

            But only works for Gemma4, or other attention-only models. (i.e. not the Qwen3.8 models I had working with the other way)

            I&#x27;m getting 63ms per answer for 1 question, 79 for 3, 93.7ms for 5. So scales nicely, too. ~55ms fixed cost + ~7-8ms per additional question.

            Much better. Thanks. I&#x27;ll commit and share work after testing it some more.

            EDIT: how can I credit you in the commit body?

            1. aaquibahm · · focus · HN ↗
              cool! im @Aaquib111 on github if you would like to credit, but no pressure
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.