‹ BackHN Continuity

Thread

Transformers Explained Visually

663 points · 92 comments · aray07

  1. mhl47 · · focus · HN ↗
    Love this! Not having spent too much time on understanding the architecture I always struggled to see how transformers get away with compressing all context in a flat vector (e.g. 768 numbers here) between attention and the MLP when computing the current token. But now I think I understand that since you often alternate between attention and MLPs it regularly mixes/queries the information of other tokens into the current computation. Probably common knowledge for everyone that learned about transformers but this made it so much quicker to see.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.