‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. velominati · · focus · HN ↗
    Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
    1. odo1242 · · focus · HN ↗
      Most likely: it does still have attention layers (the O(n^2) part), but it’s not autoregressive (which makes it O(n^3) because you have to run the whole model again for each predicted token)
      1. Ohentis · · focus · HN ↗
        I feel like the most likely is that it largely works like how all the recent copycats work: take an existing llm, modify the decoder, do RL training. I think the main reason that jev works better is that they spent more time on that post training step.
        1. odo1242 · · focus · HN ↗
          My guess as well
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.