‹ BackHN Continuity

Thread

Transformers Explained Visually

663 points · 92 comments · aray07

  1. raluk · · focus · HN ↗
    One thing that was not intuitive for me is that attention is per token generator and in theory unbounded for context size. Per token approach does not generate attention matrix as presented here but attention vector. It could provide different view and I find that aproach easier to reason about.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.