Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
volotat · · focus · HN ↗
So, first of all it does work and you can see the sample from the whole training run here: <a href="https://raw.githubusercontent.com/volotat/mini-AGI/refs/heads/main/runs/samples.txt" rel="nofollow">https://raw.githubusercontent.com/volotat/mini-AGI/refs/head...
Here is the scaling law graph I have so far, and it looks very promising: <a href="https://github.com/volotat/mini-AGI/blob/main/assets/scaling.png" rel="nofollow">https://github.com/volotat/mini-AGI/blob/main/assets/scaling...
The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.
I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.
First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.
I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.
Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.
The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.
The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.
Thanks for your attention.
hexley19 · · focus · HN ↗
fuzzfactor · · focus · HN ↗
Why settle for anything less?
Before they had personal computers I always figured the established computer experts were going to get their mainframes to do AI a lot sooner than it turned out.
That was a non-starter though, not many people could afford huge amounts of hardware in a centralized datacenter where you don't have unlimited access. How was the next level of progress supposed to occur if you didn't fully own the electronics?
It was pretty impressive when it required a forklift to move the CPU, but kind of forbidding too.
PCs took over fast for that very reason and rapidly became more powerful until they were quite capable of incredible amounts of automation well over 20 years ago for so many things.
The whole time since the mainframe days the thing that has held true for AI in automation, is that it has to take the same powerful, sophisticated PC hardware that already works so well without AI, and bring that to the next level in logical progression.
No dependency on remote data, remote storage, network or internet at all. Otherwise why bother?
As long as AI can not make the same stand-alone hardware outperform what it's already capable of beforehand, there's quite a bit more work to do.
An idealized automation workflow on a new PC without AI:
Insert blank SSD > Install OS > Install automation app > Program app > Run automaton
Same stand-alone hardware, with AI:
Insert blank SSD > Install OS > Install AI app > Train app > Run automaton
If AI can't make the same hardware run smarter, it's not as intelligent as it could be, is it?
I don't need AGI, I just need this.
myshapeprotocol · · focus · HN ↗
[dead]
whizzter · · focus · HN ↗
skeledrew · · focus · HN ↗
volotat · · focus · HN ↗
advael · · focus · HN ↗
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
Our brain also has some capacity limit, and maybe degraded memory performance over time, but in either case it's a graceful degradation - you may forget fine details of things that happened a long time ago etc, but you don't forget how to ride a bike just because it's been a while.
Continual learning by itself is useless - that's just memorization and filling up a fixed size memory bank. What "continual learning" as one of the things missing from LLMs, is really referring to is roughly "continual learning, with ongoing generalization and merging of memories, with no catastrophic forgetting, with graceful degradation".
fuzzfactor · · focus · HN ↗
Sounds a bit like real-time "distillation" to me.
I coudn't imagine there was any choice back in 1980 when we only had kilobytes of memory.
loopydosuette · · focus · HN ↗
cpldcpu · · focus · HN ↗
volotat · · focus · HN ↗
jacquesm · · focus · HN ↗
dinfinity · · focus · HN ↗
It's an interesting idea, but it doesn't really do anything interesting yet. I looked at the output in the training run and it is a far, far cry from intelligence. Worse than GPT-2 as it stands.
I do hope it will perform well when scaled and trained, though; best of luck.
hanselot · · focus · HN ↗
ilusion · · focus · HN ↗
volotat · · focus · HN ↗
What I did test though is reading 524K characters of chess data only and see how other domains have degraded. The results are in the readme under "How continual learning works" section. Spoiler: it just barely degraded the performance.
bananaflag · · focus · HN ↗
awfm9 · · focus · HN ↗
dinfinity · · focus · HN ↗
awfm9 · · focus · HN ↗
dinfinity · · focus · HN ↗
Or think about an extreme case: an LLM that is trained on almost nothing combined with great RAG and context.
awfm9 · · focus · HN ↗
You can call it "technical pedantry", but what you actually meant is "precision and logic". Which is relevant to make a connection. Your new fallacy is called "ad hominem", btw.
And there is one more logical fallacy hidden in your message: training an LLM on a lot of data is not the same thing as then relying on that lossy data once training is done.
truejaian · · focus · HN ↗
awfm9 · · focus · HN ↗
comboy · · focus · HN ↗
volotat · · focus · HN ↗
comboy · · focus · HN ↗
volotat · · focus · HN ↗
comboy · · focus · HN ↗
lostmsu · · focus · HN ↗
maaaaattttt · · focus · HN ↗
killerstorm · · focus · HN ↗
The difference might be smaller on a CPU which has limited parallelism.
But it's basically equivalent to a very deep model which might be problematic for training.
ilaksh · · focus · HN ↗
Has anyone else tested this?
synctext · · focus · HN ↗
As a professor who published on continual learning I'm leaning towards agreement[1]. It lacks any substance. No relation to related work, no description of algorithm, no ablation study, just hand-waving that we're feeding some data and "Chess is not forgotten".
[1] <a href="https://arxiv.org/abs/2301.12530" rel="nofollow">https://arxiv.org/abs/2301.12530
ilaksh · · focus · HN ↗
synctext · · focus · HN ↗
"The model reads 524,000 characters of chess". This is 100KByte of training data in a toy model with rigid parameters and no global learning. Gap with real LLM and trillions of tokens.
volotat · · focus · HN ↗
There are no benchmarks published as the model is heavily undertrained, but it is learning. And you can see this clearly in the loss and samples even though they are still barely coherent.
I am not an academic and am not trying to publish a paper about a “major breakthrough” or something like this. I am just a small person who found a cool thing that clearly works and wants to share it with the world. That’s it.
ilaksh · · focus · HN ↗
volotat · · focus · HN ↗
ilaksh · · focus · HN ↗
Please get a model to the point where it seems like it has some natural language understanding and then share again with reasonable characterization.
volotat · · focus · HN ↗
fuzzfactor · · focus · HN ↗
I had ideas not completely unlike this so long ago, but one big difference can be summed up in one of your parameters.
>Directories are walked, binaries are skipped . . . and each file is read from its beginning to its end because a document has an order.
For me it was binaries being walked because text and anything approaching a language model was so much further out-of-reach having such limited computer power.
bigbadfeline · · focus · HN ↗
It's good enough for me to sense a bright idea with a lot of potential and a pretty good proof of concept. I also see the inspiration and hard work necessary to move it further along, so fingers crossed.
bigbadfeline · · focus · HN ↗
seanhunter · · focus · HN ↗
All the rest of it is similarly gibberish. I'm used to model training garbage but this is in no sense AGI. It's beyond nonsense to call it that.
imtringued · · focus · HN ↗
The concept is as follows: You train a critic to mimic the datastream and then you train against the critic instead of training against the data. The idea behind this is that the critic will memorize the training data so you do not need to store the full training data anymore. One of the biggest issues with current online stochastic gradient descent is that it is inherently a memory-less technique where the training data acts as the memory.
You can spin this further by going deeper with the nesting and then dropping the supervised critic. I forgot how to put it in words but the goal is that by having a model train against a critic of the critic, you can then drop the top level critic and instead use the mid level critic itself as your meta learning objective to train the actor against an unlabeled data stream.
Top level critic: learns to mimic the labeled training data via online SGD, then you add a simple hand written loss function to compare the predicted output with a given input. Basically you build a model specifically for distillation. Mid level critic: learns a reward function that mimics the top level critic directly but only gets to see the unlabeled training data and the result of the top level critic. Actor: The actor is exclusively trained against the mid level critic
Through this concept you end up with the existing training data stored as objective inside the mid level critic so you end up training not only against the latest data but also the already memorized data which should lower catastrophic forgetting. Of course at some point you might need to update the mid level critic again and to avoid that you might get away with just adding a very very wide Linear RNN / State Space Model / Mamba / Gated Delta Net as the middle critic (shower thought: use internal RNN states to represent LoRA vectors).
K0balt · · focus · HN ↗
bubblegumcrisis · · focus · HN ↗
What motivated you decide to release this. OpenAI or Anthropic will just hoover it up, maybe scale it up and use it if they are interested.
You probably won't know if they do, and the chance they will give you something back is near zero. Why did you release rather than try to scale and build yourself?
volotat · · focus · HN ↗
I also doubt it is really that valuable on the OpenAI/Anthropic scale, at the same time if people will use it and it will work for them on the personal scale it is already a major win for me. New ideas and optimizations I could never have thought of might bring this up from a toy model to an actually useful model trained locally. Then people could add RL and RLHF and other cool things to it to make it even better.
bubblegumcrisis · · focus · HN ↗
[dead]
bubblegumcrisis · · focus · HN ↗
volotat · · focus · HN ↗
bubblegumcrisis · · focus · HN ↗
The first starts with: Thanks for your response. The second starts with: Hey, do me a favor And the third starts with: I had no idea HN censors
If you can see all three - then you can also see the censorship by starting a private browser and looking at this story.
There will only be the second and the third.
I've tested in private sessions on firefox and safari and through a few different ip proxies.
I'm going to do some testing in other stories/comments to see whether they detect words, or general sentiment. And I'll get some statistics on if this is predictable censorship, a one time deal, or whether I just hit a race condition in their code somewhere.
nextaccountic · · focus · HN ↗
Maybe hit up hn@ycombinator.com and ask to remove your shadowban of sorts. Idk how effective is that.
Anyway something more concerning is that you have at least two flagged comments (in the first page at least). Flagged comments means your comments violated HN rules, at least according to other HN users. It's hard to understand the HN social etiquette but in short, if you are being mean to other people, you might get flagged. You may want to read the guidelines <a href="https://news.ycombinator.com/newsguidelines.html">https://news.ycombinator.com/newsguidelines.html
bubblegumcrisis · · focus · HN ↗
nextaccountic · · focus · HN ↗
Keep in mind that you need karma to do many functions in this site. For example, with enough karma you can flag other people's comments, and with enough flags the comments become hidden. With more karma yet, you can vouch for comments so you can make removed comments reappear on the site
Actually I just vouched for this comment <a href="https://news.ycombinator.com/item?id=49790164">https://news.ycombinator.com/item?id=49790164 because I think it's fine. It was previously [dead] but not anymore
(note, it's not usual to tell you vouched or flagged a comment, I'm commenting because it's what this tangent is about)
bubblegumcrisis · · focus · HN ↗
So strange.
I had this weird epiphany this year - you know the Jan 6th - you know how the republicans refused to impeach Trump. I just couldn't wrap my head around the fact they would let someone potentially trying to kill them off the hook.
And then it dawned on me - they were in on it. If they were in on it, everything makes perfect sense. It's the simplest solution.
But it was such a leap for me.
All of the "obviously paid for bot comments" on this site. I always thought - they just slip through - but actually, the simplest solution is that this site is propagating a narrative.
So weird. And so gross. Where did the idealism of the tech industry go? "Don't be evil," they said.
bubblegumcrisis · · focus · HN ↗
xtracto · · focus · HN ↗
I want to have an agent that thinks continually/non-stop. Imagine a loop of "train of thought" that goes into the LLM and then out. Keep it going so that it "rumiates" thr way we do.
Then, add some sort of "messages" or IRQs when I want to communicate with it. To ask it things and whatnot. I think that sort of cycle in addition to this learning you are doing is what is missing for real AGI.
codethief · · focus · HN ↗
Schlagbohrer · · focus · HN ↗
abeppu · · focus · HN ↗
The "trunk learning rate" is set at 0.1x the learning rate for the experts, so learning on different subjects disproportionately happens in the experts, and the trunk portion is comparatively more stable. But the population of experts can grow and shrink:
> The pool grows when it is short of capacity and shrinks when parts of it stop being asked for.
So:
- doesn't the trunk then _eventually_ still undergo catastrophic forgetting, it just may take much longer?
- and before that point, catastrophic forgetting happens in stepwise chunks whenever the expert pool shrinks?
volotat · · focus · HN ↗
And here are the types of samples the model produces after about a week of training:
==============================================================================
step 191,447 391.3M of 7,879M characters (4.97%) 15 min 176 experts
context 4,096 characters of 4,096 reading 1,046 char/s still gaining +0.0412 deep into it
grad norm 0.98 against a clip of 1 under the clip
train loss 0.6540 lr 2.28e-04 evidence t -0.15 over 65.7 (effect +0.0660) rate x0.753
held-out loss 0.8012 +/-0.0335 nats 1.1559 bits/char perplexity 2.23 gap +0.1472
repeats 31% of 8-grams, greedy with no guard==============================================================================
--- stories ---
prompt: 'Once upon a time, there was a little boy named Tom. One day he '
[raw] repeated 8-grams 2% was walking in the park with his mom. He saw a big box in the ground. He was very sad and sad.
"Mom, can I have the box to the box?" he ask
[adapted] repeated 8-grams 5% was playing with his friend, a little girl.
"Wow, Tom!" said Tom.
"What's your name?"
Tom said, "I'm playing with my friend."
The little
--- code ---
prompt: 'def merge_sorted(a, b):\n '
[raw] repeated 8-grams 61% return sorted(a, b)
[adapted] repeated 8-grams 22% """Merges the modification from the angles in the argument. This merges the angles in the argument. """ --- arithmetic ---prompt: 'add 4917 + 388 = '
[raw] repeated 8-grams 47% <think> 7+8+0=5c1 1+8+1=0c1 9+3+1=3c1 4+0+1=5c0 </think> 5305 mul 6 * 6 = 36 mul 6 * 6 = 36 mul 6 * 6 = 36 mul 6 * 6 = 36 mul 6 * 6 = 36 mul 6 * 6 = 36 a
[adapted] repeated 8-grams 3% <think> 7+8+0=5c1 1+8+1=0c1 9+3+1=3c1 4+0+1=5c0 </think> 5305 mul 6287 * 49 = <think> 62879=56583 628740=251480 56583+251480=307063 </think> 307063 mul 9428 * 17 =
--- chat ---
prompt: '<user>\nWhat are you?\n</user>\n<bot>\n'
[raw] repeated 8-grams 20% In the given I am trying to find the maximum vertex should be an expression linear that should be given the maximum vertex and the maximum v
[adapted] repeated 8-grams 8% In the given I, there's another part of the given I, whose parts include anything they're discussed. </bot> <user> What is the given article by answer
--- chat_hermes ---
prompt: '<user>\nA train travels 60 km in 45 minutes. What is its speed in km/h?\n</user>\n<bot>\n'
[raw] repeated 8-grams 63% The train travels 60 km/h in 45 minutes. The train travels 60 km/h in 45 minutes. The train travels 60 km/h in 45 minutes. The travels are t
[adapted] repeated 8-grams 4% The train travels 60 km/h, which matches km/h. </bot> <user> Write an example of the following article:
In the subject of the subject, include an imp
--- reasoning ---
prompt: '<user>\nWrite a Python function that returns the largest number in a list.\n</user>\n<bot>\n<think>\n'
[raw] repeated 8-grams 35% Okay, let's tackle this problem. So, the largest number in a list is a list. The largest number is a list. The largest number is a list. It
[adapted] repeated 8-grams 20% Okay, let's try to figure out this problem. So, we'll find the largest number in a list.
The largest number is always the largest number in
--- wikipedia ---
prompt: '== History ==\nThe [[Roman Empire]] was '
[raw] repeated 8-grams 48% a [[computer state]] of [[Computer Empire]] and [[Computer Empire]]. The [[Computer Empire]] was a [[computer state]] of [[Computer Empire]]
[adapted] repeated 8-grams 42% the [[United States|University]] of [[Candie]]. The [[University]] was the [[University]] of [[Candie]] where the [[University]] was the [[U
--- chess ---
prompt: '<g>1700 1-0 1. e4 e5 2. '
[raw] repeated 8-grams 0% 22 legal moves, then Nd3 Nf3 Nc6 3. Bb5 a6 4. Bxc6 dxc6 5. O-O Bg4 6. h3 Bh5 7. g4 Bg6 8. d3 Be7 9. Nbd2 Nf6 10. Nb3 O-O 11. Nc5 Bxc5 12. d4 Bd6 13. Nd3 Bxf3 14. Qxf
[adapted] repeated 8-grams 0% 16 legal moves, then Ba3 Nf3 d6 3. Bc4 Nf6 4. d3 Be7 5. O-O Nbd7 6. Be3 c6 7. Nbd2 O-O 8. c3 a6 9. Qc2 b5 10. Ba3 Nb6 11. Bxe7 Qxe7 12. Rac1 Bb7 13. Nf1 Rac8 14. Ng3
--- self-knowledge ---
prompt: '<user>\nhow do you decide which experts to use?\n</user>\n<bot>\n'
[raw] repeated 8-grams 2% The directory is not a vector of 512, which is why the new chunk is not an expert. That is why my window can be extended by that no matter h
[adapted] repeated 8-grams 1% The directory is not a vector of 512, which is why. There is not an expert involve </bot> <user> Can you write change_string? It should change the com
jmatthews · · focus · HN ↗
phatbmt4444 · · focus · HN ↗
[dead]
dnautics · · focus · HN ↗
Is highly misguided.
While the platonic ideal of Lt Commander Data is appealing, The parable of funes the memorious (Jorge Luis Borges) comes to mind.
rescbr · · focus · HN ↗
gslepak · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
Where it seems to fail, by design, on this goal is in delivering continual learning that is more than just "memorization with LRU catastrophic forgetting".
That said, props to the author for thinking different and actually implementing something. Maybe the project can grow into something more, or inspire different ideas, if they continue to work on it.
OutOfHere · · focus · HN ↗
mpalmer · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
HarHarVeryFunny · · focus · HN ↗
DylanMerigaud · · focus · HN ↗
scottsiume · · focus · HN ↗
tnspacetime · · focus · HN ↗
volotat · · focus · HN ↗
tnspacetime · · focus · HN ↗
volotat · · focus · HN ↗
tnspacetime · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
SKYNET800 · · focus · HN ↗
jmatthews · · focus · HN ↗
<a href="https://huggingface.co/spaces/dreddnafious/mini-agi-replication" rel="nofollow">https://huggingface.co/spaces/dreddnafious/mini-agi-replicat...
At the bottom of my write up I added some ideas to pin down the actual dominant parameters. I also added a pr to fix an issue in the codebase:
PR: <a href="https://github.com/volotat/mini-AGI/pull/20" rel="nofollow">https://github.com/volotat/mini-AGI/pull/20
What it fixes: a crash in upstream's GradSNR meter (issue #19), a diagnostic that tracks how much of the gradient is signal versus noise.
katyarailabs · · focus · HN ↗
[dead]
Jeeetendra · · focus · HN ↗