A single function Jev-like wrapper for LLMs, including vision models
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
A single function Jev-like wrapper for LLMs, including vision models
Unofficial Hacker News client; not affiliated with Y Combinator.
bicsi · · focus · HN ↗
dist-epoch · · focus · HN ↗
imtringued · · focus · HN ↗
Sure they are no longer memory bandwidth bound thanks to that but someone could add a similar projector to a conventional model, train with a Jev style dataset and call it a day.
Whatever they are doing on inputs must either mean they intentionally chose a Mamba successor or they suffer from the same compute costs as everyone else.
dist-epoch · · focus · HN ↗
Maybe first request is unbatched, to have fast prefill, and the subsequent ones are batched.
They also don't restrict your prompt. You can have a dumb one, where you put the variable data at the front, and the details on how to process it at the back, thus you bust the user-part of the KV cache every request.