Back to No Priors: AI, Machine Learning, Tech, & Startups

Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon

Topics44
Introduction to Stefano Ermon and Inception0:05Research Background and Path to Company0:31Early Motivation and World Models Perspective1:30Evolution of Generative Models Research3:01Development of Diffusion Models3:32Extending Diffusion to Text and Code4:302024 Breakthrough and Quality Matching5:01Current State of Diffusion Models6:00Two Paradigms: Autoregressive vs Diffusion7:02Inference Time Scaling Advantages7:31Inference Workload Problems with Autoregressive Models9:00Why Diffusion Maps Better to GPUs9:32Inference Advantages for Post-Training10:32The Bitter Lesson and Parallelism11:00Challenges with Discrete vs Continuous Modalities11:30Current Performance Claims12:31Company Status and Team13:30Research Areas and Engineering14:00Focus Areas and Differentiation15:00Efficiency and the AI Factory15:30Where Speed Wins17:00Customer Example: OpenCall17:30Hardware and Software Complementarity19:00Competing as David Against Large Players20:01Working with Real Customers21:32Structure in Data and Compression22:00Handling Noisy Data24:01Pattern Recognition and Inductive Bias24:25Voice Pipeline and Alignment24:55Backward Compatibility and API Design25:00Performance and Service Continuity25:30Control and Steering Capabilities26:00Coarse-to-Fine Generation Process26:31New Product Experiences27:00Emergent Capabilities Beyond Performance27:33Data Efficiency Through Denoising28:00Workload Distribution Projections29:01Technical Challenges and Infrastructure30:01Post-Training and IP Considerations31:00Team Size and Hiring Strategy32:01Recursive Self-Improvement and Human Ingenuity33:00Organizational Structure and Resource Allocation34:01Academic Research Impact and Contrarian Bets35:30Academic Advantages in AI Research37:00
In a Nutshell

Diffusion models will win AI inference because they generate tokens in parallel rather than sequentially, mapping far better to GPU workloads than autoregressive models. This yields 10x faster inference at equivalent quality, which becomes critical for scaling test-time compute and RL post-training. Inception has already deployed production diffusion-based LLMs that match speed-optimized frontier models like Haiku while running on standard Nvidia GPUs.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Stefano Ermon is a longtime Stanford professor and co-founder and CEO of Inception. He has an extraordinarily broad body of work around generative modeling, and is especially well known as one of the fathers of diffusion. The company is challenging the large labs, with speed and efficiency being the name of the game in AI over the next few years.

Ermon has been doing research in generative models for his entire career. He started at Stanford in 2014 as an assistant professor working on building generative models. Back then the research area was not particularly hot. The models were not working well. Researchers were building little generative models over MNIST, and it was a big success if you could generate grainy images of digits. It was even hard to publish papers on that topic, and you had to justify training a generative model as a way to learn features from unlabeled data that could help with supervised learning.

Things took off and Ermon was at the right place at the right time working on the right thing. He has been doing research in that space since the beginning.

Ermon always felt that building a generative model was the right way to make sure you understand the structure in the data. He was thinking from a world models perspective, working a lot on images and thinking about having a world model where you can imagine what's going to happen if you stand up and walk out the door. Having this kind of model of the world requires some generative capabilities.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.