‹ BackHN Continuity

Thread

Transformers Explained Visually

663 points · 92 comments · aray07

  1. est · · focus · HN ↗
    It seem that everyone is getting into details of how transformers work, but I am more interested in why other setups didn't work.

    Or is it?

    1. octoberfranklin · · focus · HN ↗
      It isn't so much that "other setups didn't work". More like "the first thing we found that did work turned out to be the minimal thing that could work".

      Transformers are conventional feed-forward neural networks alternated with attention blocks. You can think of them as big huge "ordinary" neural networks augmented with this new kind of block.

      Attention blocks are basically just a differentiable hashmap. Think of it like a scratchpad memory.

      It turns out that a hashmap/scratchpad is pretty essential to being able to untangle language. I don't find this too hard to believe. Somewhere in there, you have to build the graph of which object is acting via which verb on which object.

      What is surprising is that this is all it takes! These simple little hashmap/scratchpad units (and massive scale) are really the only thing you need to tack on to a feed-forward neural network to get essentially general intelligence. This is totally surprising to me.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.