Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
Unofficial Hacker News client; not affiliated with Y Combinator.
ilaksh · · focus · HN ↗
That might explain why there are no benchmarks of any kind.
synctext · · focus · HN ↗
As a professor who published on continual learning I'm leaning towards agreement[1]. It lacks any substance. No relation to related work, no description of algorithm, no ablation study, just hand-waving that we're feeding some data and "Chess is not forgotten".
This "how-continual-learning-works" markdown text is not an algorithm [2].
[1] <a href="https://arxiv.org/abs/2301.12530" rel="nofollow">https://arxiv.org/abs/2301.12530
[2] <a href="https://github.com/volotat/mini-AGI/#how-continual-learning-works" rel="nofollow">https://github.com/volotat/mini-AGI/#how-continual-learning-...
ilaksh · · focus · HN ↗
synctext · · focus · HN ↗
"The model reads 524,000 characters of chess". This is 100KByte of training data in a toy model with rigid parameters and no global learning. Gap with real LLM and trillions of tokens.
This model really addresses the problem of preserving previously learned knowledge, but by restricting the LR of the trunk it stops acquiring new knowledge. Details: "Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural Perspective"