‹ BackHN Continuity

Thread

Transformers Explained Visually

663 points · 92 comments · aray07

  1. E-Reverance · · focus · HN ↗
    I get that this is for explaining GPT-2, but I really hope laymen don't use it as an example of how modern models work (ex. absolute positional encoding is no longer used)

    edit: I know that it mentions its not modern, but these kinds of details have major implications in terms of the representations a model can learn, which is in many ways the most important part!

    1. ViscountPenguin · · focus · HN ↗
      Yeah it's a bit odd, especially since RoPe is a lot more conceptually simple imo.
      1. ianand · · focus · HN ↗
        Interestingly, I find absolute positional embeddings easier to explain to a layperson than rope, e.g. [1]. What am I missing?

        [1] <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=ZuiJjkbX0Og&amp;t=5712s" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=ZuiJjkbX0Og&amp;t=5712s

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.