Mini-AGI trains a continual learning model from scratch on an 8GB laptop GPU
Mini-AGI – dynamic continual learning model trained from scratch on 8GB VRAM

Mini-AGI is a byte-level language model that assembles its own architecture, trains on a single 8GB VRAM GPU, and keeps learning without catastrophic forgetting. It pages expert weights from disk, grows and prunes capacity dynamically, and uses a 0.1x trunk learning rate to retain 99.84% of progress against chance. It's a toy-level experiment showing personal continual learning is possible on modest hardware.
The trunk learning rate is the mechanism. Running it at 0.1x the experts' rate takes forgetting from +2.2300 to +0.0067 nats, which is 99.84% of progress retained against chance.
- K0balt
Interesting. I wonder how much could be gained from using tokenization, which makes the model work at a semantic level rather than a syntactic level? I think it’s a force multiplier, but idk if it works here.
- ilaksh
If you actually scroll through the transcript he links to, you will see that something that looks like it could be training is happening, but no coherent responses are coming out at any point. At least not that I saw skimming through.
That might explain why there are no benchmarks of any kind.
- whizzter
Nobody will throw rocks, I think most people are curious/suspicious about the big players and wants more hands-on since we suspect that this all will come down in cost soon enough.
- bananaflag
This is the first thing I see in my life that really looks like proto-AGI, it deserves its name.
- advael
Seems interesting, I've been messing with a lot of continuous learning approaches lately and it's cool to see something that's built from the ground up for avoiding catastrophic forgetting. Worth a clone for sure