Any diffusion model is potentially a Jev in disguise: <a href="https://github.com/vllm-project/vllm/pull/57250" rel="nofollow">https://github.com/vllm-project/vllm/pull/57250
Runs ~0.2s per decision on my DGX Spark.
10/10 programming language detection
9/10 human language detection
10/12 unit magnitude comparison
All incorrect answers are marked with low-P.
It (DiffusionGemma with the Jev mode) can also solve an ASCII maze.
Your maze demo gave me the idea to try out the "speculative fan-out" pattern [1]. It seemed interesting to try solving mazes in one-shot. Unfortunately it seems like Jev can't reliably solve basic mazes even with a step count of 1! [2] I was very surprised. Could you point me in the direction of your maze solving code so I can see if it's a skill issue? The only other explanation I can come up with is that Jev was not trained on spatial reasoning tasks at all, and on the other hand DiffusionGemma has a vision tower and significantly more spatial data in its training set.
mmastrac · · focus · HN ↗
Runs ~0.2s per decision on my DGX Spark.
All incorrect answers are marked with low-P.It (DiffusionGemma with the Jev mode) can also solve an ASCII maze.
budro · · focus · HN ↗
[1] <a href="https://docs.typesafe.ai/patterns/fan-out" rel="nofollow">https://docs.typesafe.ai/patterns/fan-out [2] <a href="https://github.com/Bud-ro/jev-demos/tree/master/packages/maze_lookahead" rel="nofollow">https://github.com/Bud-ro/jev-demos/tree/master/packages/maz...
[deleted] · · focus · HN ↗
[deleted]