‹ BackHN Continuity

Thread

How to win a beer with high-dimensional statistics

81 points · 11 comments · jamie-simon

  1. JHonaker · · focus · HN ↗
    After seeing the claim, I knew it would work because of a few intuitive things:

    1. Things that are the result of a very large number of small perturbations follow a Gaussian distribution.

    1a. Word embeddings are the result of millions/billions/trillions of small perturbations (i.e. gradient descent)

    2. As you increase the dimension, d, of a multivariate Gaussian distribution, the actual mass of the data is increasingly concentrated in the thin spherical shell. For MVN distributions with standardized marginal dimensions, this is on the sphere of radius \sqrt{d}.

    In fact when I searched "Gaussian curse of dimensionality" I got this at the top result: [1] which shows this phenomena perfectly. Or this one with pretty interactive pictures [2]

    [1]: <a href="https:&#x2F;&#x2F;www.miryusupov.com&#x2F;blog&#x2F;posts&#x2F;thin-shells&#x2F;index.html" rel="nofollow">https:&#x2F;&#x2F;www.miryusupov.com&#x2F;blog&#x2F;posts&#x2F;thin-shells&#x2F;index.html

    [2]: <a href="https:&#x2F;&#x2F;aseemrb.me&#x2F;blog&#x2F;high-dimensional-gaussians-on-a-sphere&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aseemrb.me&#x2F;blog&#x2F;high-dimensional-gaussians-on-a-sphe...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.