‹ BackHN Continuity

Thread

Reverse-engineered Jev-like model

169 points · 24 comments · rochansinha

  1. mmastrac · · focus · HN ↗
    Any diffusion model is potentially a Jev in disguise: <a href="https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250

    Runs ~0.2s per decision on my DGX Spark.

      10&#x2F;10 programming language detection
      9&#x2F;10 human language detection
      10&#x2F;12 unit magnitude comparison
    
    All incorrect answers are marked with low-P.

    It (DiffusionGemma with the Jev mode) can also solve an ASCII maze.

    1. budro · · focus · HN ↗
      Your maze demo gave me the idea to try out the &quot;speculative fan-out&quot; pattern [1]. It seemed interesting to try solving mazes in one-shot. Unfortunately it seems like Jev can&#x27;t reliably solve basic mazes even with a step count of 1! [2] I was very surprised. Could you point me in the direction of your maze solving code so I can see if it&#x27;s a skill issue? The only other explanation I can come up with is that Jev was not trained on spatial reasoning tasks at all, and on the other hand DiffusionGemma has a vision tower and significantly more spatial data in its training set.

      [1] <a href="https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;patterns&#x2F;fan-out" rel="nofollow">https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;patterns&#x2F;fan-out [2] <a href="https:&#x2F;&#x2F;github.com&#x2F;Bud-ro&#x2F;jev-demos&#x2F;tree&#x2F;master&#x2F;packages&#x2F;maze_lookahead" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Bud-ro&#x2F;jev-demos&#x2F;tree&#x2F;master&#x2F;packages&#x2F;maz...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.