No Priors · Friday, September 18, 2026
Stefano Ermon highlighted a 2024 breakthrough where diffusion models matched the quality of autoregressive models like GPT-2 for text generation. Crucially, these diffusion models were significantly faster, achieving text generation speeds up to 10 times quicker due to their parallel processing capabilities.
“Uh, we had a breakthrough in 2024, we published a paper, uh, basically showing that for the first time, it was possible to match the quality of an autoregressive model at the GPT-2 scale, so less than a billion parameters, still fairly academic, but we were able to train basically still a transformer model as a diffusion model on the same data. We were able to match the quality like the same perplexity, uh, you know, you were fitting the data just as well as an autoregressive model with the same number of parameters, but the diffusion model was significantly faster.”
“Because it's diffusion, because you're outputting many tokens at the same time, we were able to generate text like 10X faster compared to the autoregressive model.”