Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon
In a Nutshell
Diffusion models will win AI inference because they generate tokens in parallel rather than sequentially, mapping far better to GPU workloads than autoregressive models. This yields 10x faster inference at equivalent quality, which becomes critical for scaling test-time compute and RL post-training. Inception has already deployed production diffusion-based LLMs that match speed-optimized frontier models like Haiku while running on standard Nvidia GPUs.
These notes were generated by AI and may contain inaccuracies.
Stefano Ermon is a longtime Stanford professor and co-founder and CEO of Inception. He has an extraordinarily broad body of work around generative modeling, and is especially well known as one of the fathers of diffusion. The company is challenging the large labs, with speed and efficiency being the name of the game in AI over the next few years.
Ermon has been doing research in generative models for his entire career. He started at Stanford in 2014 as an assistant professor working on building generative models. Back then the research area was not particularly hot. The models were not working well. Researchers were building little generative models over MNIST, and it was a big success if you could generate grainy images of digits. It was even hard to publish papers on that topic, and you had to justify training a generative model as a way to learn features from unlabeled data that could help with supervised learning.
Things took off and Ermon was at the right place at the right time working on the right thing. He has been doing research in that space since the beginning.
Ermon always felt that building a generative model was the right way to make sure you understand the structure in the data. He was thinking from a world models perspective, working a lot on images and thinking about having a world model where you can imagine what's going to happen if you stand up and walk out the door. Having this kind of model of the world requires some generative capabilities.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.