‹ BackHN Continuity

Thread

Transformers Explained Visually

663 points · 92 comments · aray07

  1. est · · focus · HN ↗
    It seem that everyone is getting into details of how transformers work, but I am more interested in why other setups didn't work.

    Or is it?

    1. ActorNightly · · focus · HN ↗
      >how transformers work,

      most people in ML have no idea what transformers actually are.

      Traditional networks, at every layer, used to be output = [weights matrix][input], where input is a vector, and weights matrix is the weights, where each row corresponds to the set of weights for each neuron.

      Transformers upscale the dimension of the data. Instead of the above, transformers do [output] = [input][weights_matrix]. When you multiply an input by a matrix, you get an output matrix back. Thats all that happens. Nothing fancy. You have weights matricies for K/Q/V, which when post multiplied with the input, give you the KQV vectors, and then you just simply multiply them together and apply a scaling factor.

      There is nothing magical about K/Q/V. There is nothing about any one doing any querying or any one representing some keys. The naming is just a carry over from how they that selection process is used in pre llm data science fields where you manually define the key and query matricies to define relationships between components.

      The reason of why it works is because is an extension of something called kernel tricks from pre LLM machine learning days - you map a lower dimensional space to an extra dimension based on some equation, and it lets you apply some classifier on the combination of existing values and new value. Thats what transformers are doing - they are mapping the individual token to the dk x n_heads latent space, which allows for a higher dimensional representation of the data, capturing complex relationships.

      You can do Transformers with 5 matricies instead of 3, you can do this with 4-dimentional tensors, and so on. The thing is, there really isn't any way to tell if any of that gives you more advantage - it certainly would give you more granularity, but as of right now, in terms of training to generate a specific token given previous ones before it, it seems that you don't need any more dimentions than dk x n_heads. Interestingly enough, you also can mathematically represent any such transformer including the starting one with a sequence of linear layers like in traditional networks, the only thing is that it becomes computationally inefficient due to having duplicates of data.

      The reason why RNNs and others and others didn't work is because RNN training is effectively trying to linearly regress on chaotic effects - i.e what set of starting conditions would evolve with a given process into what you want. This is an NP hard problem, and you can't really do it linearly.

      Transformer models on the other hand, use breadth instead of compute to capture interactions. In those learned weight matrices, you have a latent space of a bunch of "knowledge" compressed, and an algorithm to search on that "knowledge".

      But, its very possible that an RNN can be smarter than a frontier model while being much smaller in size - in the same way that its very possible that you can have the right set of prompts for an existing local inference smaller model that can basically be very close to AGI in terms of being able to solve any problem across any domain. Right now, the space is about exploring those prompts, which is the frameworks and harnesses, to get to there, as well as making the compute portion more efficient so you can explore that space faster.

      And the thing that comes after harnesses/efficiency in terms of progress should be obvious if you understand all of the above.

      1. octoberfranklin · · focus · HN ↗
        Hey I like your exposition of attention in terms of the kernel trick, but the big huge difference is that kernel methods use the inner product which is a commutative operation -- it's bidirectional (and both tokens are projected into the space by the same function prior to being dot-producted).

        This means that it can't capture unidirectional relationships, like "ball" is the object on which the verb "threw" acts in the sentence "I threw the ball". This relationship is true in only one direction; it isn't true to say "threw" is the object on which the verb "ball" acts.

        I do think it would be fair to say that transformers generalize the kernel trick to noncommutative relations by applying a different projection function (W^Q and W^K) to the two tokens being considered. This makes the overall operation (project then dot product) a noncommutative operation.

        And the thing that comes after harnesses/efficiency in terms of progress should be obvious if you understand all of the above.

        Nirvana? Singularity? Paperclips? Vernor Vinge rising from the dead? I'm curious; please share!

        1. ActorNightly · · focus · HN ↗
          I mean, given sentence construction, you don't really need to capture directionality, you just have a mapping of how sentences are constructed to the latent space of some representation.

          >Nirvana? Singularity? Paperclips? Vernor Vinge rising from the dead? I'm curious; please share!

          Simulated evolution. Thats how you "solve" highly nonlinear chaotic systems. And generally, if you think about it, you have to have some secondary system on top of the knowledge embedded in LLMs to drive them to select certain tokens, which then starts to eerily resemble what humans call emotions in themselves.

          1. octoberfranklin · · focus · HN ↗
            I mean, given sentence construction, you don't really need to capture directionality, you just have a mapping of how sentences are constructed to the latent space of some representation.

            Uh, no, you absolutely do need directionality.

            Without directionality you're simply grouping words into equivalence classes. That's nice, but graphs are strictly more powerful than equivalence classes.

            Simulated evolution. Thats how you "solve" highly nonlinear chaotic systems.

            Eh. I dunno. Evolution is a horribly inefficient way to do learning. Nature uses evolution because it's the only learning algorithm that you can implement using uncoordinated chemical reactions on DNA/RNA/proteins.

            As soon as evolution hits a point where it can build neurons all the effort switches over to gradient descent. Just look at humans.

            Even cellular/self-organizing systems can be driven more efficiently by gradient learning than evolution... if you want your mind blown, read the paper that this page summarizes:

            <a href="https:&#x2F;&#x2F;google-research.github.io&#x2F;self-organising-systems&#x2F;difflogic-ca&#x2F;" rel="nofollow">https:&#x2F;&#x2F;google-research.github.io&#x2F;self-organising-systems&#x2F;di...

            1. ActorNightly · · focus · HN ↗
              Im saying you don&#x27;t need to capture directionality if you capture all possible cases of sentence construction.

              As for evolution, you can still go gradients, the problem is that you can&#x27;t do gradients in a space with many false positives. You need some method of figuring out the true optimal point.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.