No Priors · Friday, September 18, 2026
Stefano Ermon argued that diffusion models are fundamentally better than autoregressive models for inference time efficiency. He drew an analogy to the shift from RNNs to transformers for training parallelism, stating that diffusion models enable similar parallelism for inference, making them more suitable for GPUs.
“The the the workload that we have at inference time in a diffusion model, it's basically very similar to the workload that you have for training, where you are processing many tokens at the same time in parallel. And so it's built to, uh, have an inference workload that maps really well to GPUs.”
“But if you think about inference, inference generation. Uh, autoregressive models are still sequential. The computation is, uh, one left to right, one token at a time. You cannot generate the 10th token until you've generated everything that comes before it. That kind of workload is, um, does not map well to GPUs.”