‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. velominati · · focus · HN ↗
    Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
    1. odo1242 · · focus · HN ↗
      Most likely: it does still have attention layers (the O(n^2) part), but it’s not autoregressive (which makes it O(n^3) because you have to run the whole model again for each predicted token)
      1. k__ · · focus · HN ↗
        Then again, it's only a very small number and fixed set of tokens for the output.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.