Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Unofficial Hacker News client; not affiliated with Y Combinator.
velominati · · focus · HN ↗
odo1242 · · focus · HN ↗
microtonal · · focus · HN ↗
odo1242 · · focus · HN ↗
(Specifically, there are two O(n^2) steps in an attention layer and KV caching makes the first one O(n) with caching - but the overall big-O is still O(n^2) because KV caching doesn't affect the second step.)
microtonal · · focus · HN ↗
Then
This gives us the normal quadratic complexity of prefills (of course, in an actual implementation an attention mask is used to ensure that tokens cannot attend to prior tokens and you may have things like ALiBi).Note that during decoding, the dimensions change. Since it is an autoregressive model, we do not need to recompute the values and keys of prior tokens, only of the token that we are currently decoding. Of course, the token that we are decoding still attends to all prior tokens and itself.
So attention during a decoding step is O(N), so when decoding N tokens, it is O(N^2) overall.The point-wise feed-forward layer does not matter, in decoding it only needs to be computed for the representation of the token that we are currently generating. We don't need the representations of the preceding tokens for the next layer, since we have already cached their keys and values for each layer.
Disclaimer: I was one of the developers of a widely-used inference engine and implemented several of these optimizations.
odo1242 · · focus · HN ↗
odo1242 · · focus · HN ↗