‹ BackHN Continuity

Thread

Why Backprop Goes Backward (2018)

70 points · 10 comments · andsoitis

  1. akssri · · focus · HN ↗
    The intuition here is okay - but the math is hand-wavy with imprecise terms like "blow-up" etc.

    The statements however, if taken to mean optimality, are also incorrect. Reverse-mode AD (backprop) is generally quite efficient for scalar outputs (more generally, when n_inputs >> n_outputs), but it's not strictly optimal even for this particular scalar-output case.

    Consider for eg. a MLP, with 4-layers with dims (1, N, 1, N, 1) - reverse-mode here does ~3N multiplies, but the optimal is ~2N. The optimal ordering for gradient accumulation is in fact NP-hard on general DAGs, but such 'cross-mode' AD is apparently quite hard to implement and not often seen given the marginal gains.

    Griewank-Walther's excellent book is a excellent reference for this and much more,

    <a href="https:&#x2F;&#x2F;epubs.siam.org&#x2F;doi&#x2F;book&#x2F;10.1137&#x2F;1.9780898717761" rel="nofollow">https:&#x2F;&#x2F;epubs.siam.org&#x2F;doi&#x2F;book&#x2F;10.1137&#x2F;1.9780898717761

    They also had a library called ADOL-C that had mixed-mode.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.