Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
Unofficial Hacker News client; not affiliated with Y Combinator.
abeppu · · focus · HN ↗
The "trunk learning rate" is set at 0.1x the learning rate for the experts, so learning on different subjects disproportionately happens in the experts, and the trunk portion is comparatively more stable. But the population of experts can grow and shrink:
> The pool grows when it is short of capacity and shrinks when parts of it stop being asked for.
So:
- doesn't the trunk then _eventually_ still undergo catastrophic forgetting, it just may take much longer?
- and before that point, catastrophic forgetting happens in stepwise chunks whenever the expert pool shrinks?
volotat · · focus · HN ↗
And here are the types of samples the model produces after about a week of training: