‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. velominati · · focus · HN ↗
    Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
    1. hellohello2 · · focus · HN ↗
      The main difficulty with fast (low-latency) inference, is not actually computation but loading parameters from memory. The problem with generating 1 token at a time isn't that that its expensive computationally (it is, but so is training), but that you need to stream your entire model from memory for every single token (and also the KV cache but that's besides the point). So the strategy usually taken is to process big batches of user requests, letting you share the memory loads across users. This means individual answers aren't that fast, but you can do lots at once. This is why local inference isn't cost-effective; its because the model weren't designed for it in the first place.

      My hypothesis for Jev is that they simply generate many answers independently in parallel from your prompt, and then discard the duplicates (or train to avoid duplication with attention between the ). In that way the entire batch is 1 user's prompt.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.