A new open-source project, mini-AGI, demonstrates that a byte-level language model can achieve continual learning—learning from a continuous stream of data without catastrophic forgetting—on consumer-grade hardware with as little as 8GB of VRAM.
mini-AGI: A Continual Learning Byte-Level Model Trained on 8GB VRAM
The model utilizes a dynamic Mixture of Experts (MoE) architecture that assembles itself one character at a time. Unlike traditional Large Language Models (LLMs) that are trained on static datasets and then frozen, mini-AGI grows and prunes its own capacity during training. It uses a recurrent block that applies up to 26 block-applications per character, selecting experts from a shared pool based on the immediate input. To manage memory, the model's weights are stored on disk and paged into VRAM as needed, allowing the total parameter count to be limited by disk space rather than video memory.
To prevent catastrophic forgetting—the tendency of neural networks to lose previously learned information when trained on new data—the design employs a "trunk" learning rate. By running the trunk at one-tenth the rate of the individual experts, the model was able to retain 99.84% of its progress while learning entirely new subjects from a single data stream.
The model operates at a byte level, using a vocabulary of 265 tokens (256 bytes plus structural markers), which allows it to process any data without a fixed tokenizer. It is designed for PCs or laptops with at least 8GB of VRAM and is built to be truly personal, as the training and learning process is continuous and tied to the user's own hardware and data.
Sources
- Mini-AGI – dynamic continual learning model trained from scratch on 8GB VRAM (Hacker News Frontpage, 2026-09-21)