‹ BackHN Continuity

Thread

A single function Jev-like wrapper for LLMs, including vision models

157 points · 45 comments · allanrbo

  1. bicsi · · focus · HN ↗
    Of course it works, Jev is nothing but an API breakthrough
    1. dist-epoch · · focus · HN ↗
      Jev is rumored to be a 30B model, and it's input price is MUCH cheaper than similarly sized models. The maker is also heavily focused on having a profitable product, so it's unlikely to be subsidizing the cost, especially since they say they have more demand than what they can serve.
      1. imtringued · · focus · HN ↗
        They can't overturn the economics of attention by restricting themselves to a single token output.

        Sure they are no longer memory bandwidth bound thanks to that but someone could add a similar projector to a conventional model, train with a Jev style dataset and call it a day.

        Whatever they are doing on inputs must either mean they intentionally chose a Mamba successor or they suffer from the same compute costs as everyone else.

        1. dist-epoch · · focus · HN ↗
          Jev claims 70-500 ms latency, including for the first request. This requires some clever engineering at least, which will take a little to duplicate.

          Maybe first request is unbatched, to have fast prefill, and the subsequent ones are batched.

          They also don't restrict your prompt. You can have a dumb one, where you put the variable data at the front, and the details on how to process it at the back, thus you bust the user-part of the KV cache every request.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.