‹ BackHN Continuity

Thread

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

277 points · 81 comments · volotat

  1. jmatthews · · focus · HN ↗
    I built a smaller scale replication. Most of your claims prove out but the lever is essentially Biderman et als "learn less forget less". My first pass actually made this mistake initially and the findings didn't replicate. After controlling for step density I replicated your findings. Here is a write up:

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;dreddnafious&#x2F;mini-agi-replication" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;dreddnafious&#x2F;mini-agi-replicat...

    At the bottom of my write up I added some ideas to pin down the actual dominant parameters. I also added a pr to fix an issue in the codebase:

    PR: <a href="https:&#x2F;&#x2F;github.com&#x2F;volotat&#x2F;mini-AGI&#x2F;pull&#x2F;20" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;volotat&#x2F;mini-AGI&#x2F;pull&#x2F;20

    What it fixes: a crash in upstream&#x27;s GradSNR meter (issue #19), a diagnostic that tracks how much of the gradient is signal versus noise.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.