Why Backprop Goes Backward (2018)
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Why Backprop Goes Backward (2018)
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
LoganDark · · focus · HN ↗
dkrylov · · focus · HN ↗
kazinator · · focus · HN ↗
TeMPOraL · · focus · HN ↗
The naive explanation I understand: because viewport is fixed, an it's much less rays to trace if you cast from viewport pixels in various directions, than if you cast from light sources in every direction and rejection-sample at the very end to those rays that happen to intersect the viewport. Plus, you get spatially uniform image this way, whereas realistic take would oversample the bright parts before giving you any samples for the dark parts.
akssri · · focus · HN ↗
The statements however, if taken to mean optimality, are also incorrect. Reverse-mode AD (backprop) is generally quite efficient for scalar outputs (more generally, when n_inputs >> n_outputs), but it's not strictly optimal even for this particular scalar-output case.
Consider for eg. a MLP, with 4-layers with dims (1, N, 1, N, 1) - reverse-mode here does ~3N multiplies, but the optimal is ~2N. The optimal ordering for gradient accumulation is in fact NP-hard on general DAGs, but such 'cross-mode' AD is apparently quite hard to implement and not often seen given the marginal gains.
Griewank-Walther's excellent book is a excellent reference for this and much more,
<a href="https://epubs.siam.org/doi/book/10.1137/1.9780898717761" rel="nofollow">https://epubs.siam.org/doi/book/10.1137/1.9780898717761
They also had a library called ADOL-C that had mixed-mode.
omnicognate · · focus · HN ↗
_0ffh · · focus · HN ↗
omnicognate · · focus · HN ↗
qwlk4 · · focus · HN ↗
eigenspace · · focus · HN ↗
Forwards mode AD (and finite differences) tell you how much wibble of the inputs corresponds to a given wobble in the outputs.
Reverse mode AD tells you how much wobble of the outputs corresponds to a given wibble in the inputs.
If you have more inputs than outputs (such as in optimization), it's cheaper to calculate the wibbles given a wobble, than the other way around.