‹ BackHN Continuity

Thread

Transformers Explained Visually

663 points · 92 comments · aray07

  1. andblac · · focus · HN ↗
    Nicely done. For me the most fascinating thing about attention heads is the place where Attention matrix is already computed and is getting multiplied by Value vector. It behaves exactly like pushing Value vector through Dense layer of ordinary network where Attention matrix forms weights of that layer. So attention head is trained to construct this small single layer network dynamically during inference from Key and Query. And that's the point. That's rarely underlined in explanations of LLMs architecture and for me it's quite amazing that it works so well. This mechanism easy to observe in this particular visualization if you click through it.
    1. abirch · · focus · HN ↗
      I&#x27;m a huge 3Blue 1Brown fan and his series on this was great: <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=wjZofJX0v4M" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=wjZofJX0v4M
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.