One person trained a 3.8B LLM to beat GPT-2 for $998
Training a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes
Hugo Vergnes built little-lm, a config-driven framework, and trained a 3.8B-parameter model on 65B tokens in 43 hours for $998 using 8× B200s. The model scores 0.384 on CORE, surpassing GPT-2 and nanochat d32. Key wins: Muon optimizer, trapezoidal LR schedule, ClimbMix data, FP8 with vocab padding, and fused cross-entropy. He shares what worked, what didn't, and the surprising impact of context length.
As the frontier moves, $1,000 takes you further and further.
- brainless
More and more such experiments. I felt sad for a couple months when I realized that writing code will not be the same since. Now I am on the other side.
LLMs are interesting in their own ways but as an engineer, this is a way to unlock a new way of building software.
I recently build a Claude-assisted Excel/CSV parser for a US based property management system (tax compliance). Uses Haiku and has a lot of deterministic code to extract column/row combinations to check known formats and finally handing out the headers to Haiku to give us a translation plan to our support columns.
These would eventually become part of the software, in a tiny LLM. The gap between training (such tiny LLMs) and inference will shrink. We can consult Claude for edge cases, create sample dataset and train a the tiny LLM on demand so we go to Claude less.
The tooling that a project needs is really important. Something I have been feeling as well. Not just in LLM building projects, but regular software projects that are LLM generated.
- rao-v
This is really neat! I'd be tempted to try this again targetting ~1B params and the entire cookbook of "current" small model ideas: gated delta nets (or is ~2K context too short to benefit?), per layer or n-gram embeddings, gated residuals etc.
- johnnylambada
I don’t blame you for using an LLM to write an article about an LLM that you built. I’ve been reading so much LLM output that now I see it everywhere. I wonder if humans will start writing more like LLMs?