How to Build a Diffusion Language Model

Diffusion models have revolutionized image generation, but applying them to text was long an open problem. This guide from the Kuleshov Group explains how masked diffusion works—training a bidirectional transformer to fill in randomly masked tokens, then generating text by iterative unmasking. It covers the probabilistic foundations, block diffusion for variable-length output, architectural choices, and techniques like distillation and post-training. By 2026, diffusion LLMs like Mercury 2, Gemma Diffusion, and Nemotron Diffusion are competitive with autoregressive models, offering faster generation, error correction, and bidirectional context.
Rather than producing text one token at a time, they generate the whole sequence at once, starting from an initial guess and iteratively refining it over a number of steps.
- radarsat1
Something I've wondered, maybe I should just do it if I can find some time, but... given DeepSeek's nice results on using rendered text as input, I'm wondering if anyone has given serious research efforts towards image-based diffusion methods for text.
As in, instead of all the complexities induced by discrete token generation, just generate the image of the text using standard image diffusion methods, then convert it to text.
If you used a single, monospace font, I bet this would be even pretty efficient, because the OCR problem becomes basically just direct template matching.
But I guess probably there is already a paper out there, I haven't searched. I'd be curious to know if it compares on par with token-based methods.
- quirino
I've been studying these a bunch for a project in university. Last week I went over the derivation of the ELBO for a couple hours and it was a very fun and elucidating exercise.
Once you give names to the larger mathematical structures and understand them a bit better it becomes quite simple. I wish some of the blogs/papers I'd read had named "Importance Sampling".
The probability notation can be pretty confusing too. Sometimes it's hard to understand the "types" of some variables. But I'm inexperienced.
ChatGPT was surprisingly helpful. If you put in the work to truly understand the where the gaps are in your mental model (which parts aren't completely intuitive), it can do an amazing job filling in the gaps.
- rottc0dd
A good video from welch labs on image generation with diffusion models:
https://www.youtube.com/watch?v=iv-5mZ_9CPY&pp=ygUVZGlmZnVza...