‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. velominati · · focus · HN ↗
    Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
    1. odo1242 · · focus · HN ↗
      Most likely: it does still have attention layers (the O(n^2) part), but it’s not autoregressive (which makes it O(n^3) because you have to run the whole model again for each predicted token)
      1. microtonal · · focus · HN ↗
        But nobody runs the whole model again for each predicted token in each decode step, this is what KV caching (and prefix caching) is for, to keep the model running in O(n^2) overall (and O(n) per decoding step).
        1. odo1242 · · focus · HN ↗
          Models are O(n^2) on input size even with KV caching - the KV cache only decreases the processing time by a constant factor.

          (Specifically, there are two O(n^2) steps in an attention layer and KV caching makes the first one O(n) with caching - but the overall big-O is still O(n^2) because KV caching doesn't affect the second step.)

          1. microtonal · · focus · HN ↗
            I think that I understand where you get the idea from, but you are mistaken. Check equation 1 of the attention paper:

                softmax(QK^T/sqrt(d_k))V
            
            So without KV-caching, if suppose you have N tokens, then you have

                Q: N x q_dim (leaving the batch size and n heads out for simplicity, they are constants for our purposes here)
                K^T: k_dim x N
            
            So

                QK^T: N x N <- this is the attention matrix
            
            The scaling and softmax are irrelevant here, since they are elementwise

            Then

                V: N x v_dim
                (QK^T)V: N x v_dim
            
            This gives us the normal quadratic complexity of prefills (of course, in an actual implementation an attention mask is used to ensure that tokens cannot attend to prior tokens and you may have things like ALiBi).

            Note that during decoding, the dimensions change. Since it is an autoregressive model, we do not need to recompute the values and keys of prior tokens, only of the token that we are currently decoding. Of course, the token that we are decoding still attends to all prior tokens and itself.

                Q: 1 x q_dim
                K^T: k_dim x (N+1)
                QK^T: 1 x (N+1)      <- Note that attention in this step is not quadratic anymore, since
                                        we only need to compute how the current token attends to prior
                                        tokens, the representations of prior tokens are frozen. K comes
                                        from the KV-cache.
                V: (N + 1) x v_dim   <- V comes from the KV-cache
                (QK^T)V: 1 x v_dim   <- Also not quadratic, the value is only computed for the current token.
            
            
            So attention during a decoding step is O(N), so when decoding N tokens, it is O(N^2) overall.

            The point-wise feed-forward layer does not matter, in decoding it only needs to be computed for the representation of the token that we are currently generating. We don't need the representations of the preceding tokens for the next layer, since we have already cached their keys and values for each layer.

            Disclaimer: I was one of the developers of a widely-used inference engine and implemented several of these optimizations.

            1. odo1242 · · focus · HN ↗
              Ah, I see
              1. odo1242 · · focus · HN ↗
                (To be specific, I was thinking about the point-wise feed-forward layer, but didn't realize you don't have to run the full layer for a decoding pass)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.