Love this! Not having spent too much time on understanding the architecture I always struggled to see how transformers get away with compressing all context in a flat vector (e.g. 768 numbers here) between attention and the MLP when computing the current token. But now I think I understand that since you often alternate between attention and MLPs it regularly mixes/queries the information of other tokens into the current computation. Probably common knowledge for everyone that learned about transformers but this made it so much quicker to see.
mhl47 · · focus · HN ↗