I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.
It seems people are just guessing at the architecture behind Jev. Obviously the functionality itself is easy to replicate, but why Jev seems to be making such a splash (beyond the doh! factor of it's huge applicability) is the ultra-low cost and speed, which may be due to architecture.
The Laya model compared in TFA shows one way Jev may be getting it's speed and low cost - by using a BERT-like bidirectional model rather than an auto-regressive one (LLM).
ricardobeat · · focus · HN ↗
I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.
HarHarVeryFunny · · focus · HN ↗
The Laya model compared in TFA shows one way Jev may be getting it's speed and low cost - by using a BERT-like bidirectional model rather than an auto-regressive one (LLM).