Back to Y Combinator

Inference, Diffusion, World Models, and More | YC Paper Club

Y CombinatorMay 28, 20261h 7m
Topics50
YC Paper Club Launch0:07First Paper: Speculative Speculative Decoding3:30Fast Inference Demo6:31Vanilla Speculative Decoding7:32Speculation as Currency Exchange10:33Speculative Speculative Decoding (SSD)11:32Implementation Details and Trade-offs15:00Performance Results16:32Second Paper Introduction17:32Diffusion Model Predictive Control18:33Core MPC Components19:32Diffusion MPC Motivation20:01Taxonomy of Related Approaches20:33Diffusion-Based Agents Landscape22:31Algorithm Overview24:30Diffusion-Based Multi-Step Planning in DMPC25:17Comparison to Prior Work26:32Performance and Adaptability Results27:02Component Ablations29:02Transition to World Models Presentation29:50World Models: Definition and Motivation30:30Observations, Actions, and Prediction in Real Systems32:31Model-Free vs. Model-Based Policies33:30Toy Environment Demonstration35:31Challenges in Training World Models36:02JEPA and the Lay World Model Approach37:31Capabilities: Open-Loop Prediction and Model Predictive Control40:00Surprise Quantification and Uncertainty42:01Discussion Points42:31Generalization in Deep Learning43:31Explaining Overparameterization44:31Flat Minima and Generalization Bounds47:38Benign Overfitting48:00Flexible vs. Biased Hypothesis Spaces49:00Deep Learning Mysteries Explained by Existing Theory49:30Intelligence per Watt and Intelligence per Sample50:30Pre-training Progress and Scaling Laws51:30The Data-Constrained Regime53:01Core Contribution: Scaling Laws for Data-Constrained Pre-training54:30Experimental Setup55:00Standard Recipe and Overfitting56:01Aggressive Regularization Baseline56:30Ensembling Results58:01Joint Scaling Recipe59:00Infinite Compute Performance1:00:01Data Scaling Laws1:00:30Distillation for Practical Inference1:02:31Self-Distillation Results1:03:30Downstream Benchmarks and Other Settings1:04:31Key Takeaways1:05:30
In a Nutshell

Three papers showed concrete algorithmic wins for scaling intelligence under real constraints. Speculative speculative decoding parallelizes drafting and verification on separate hardware to reach 300 tokens/sec on Llama-70B, turning extra FLOPs into lower latency and higher throughput. Diffusion MPC factorizes action proposals from dynamics models, enabling runtime adaptation to new rewards and dynamics where joint models fail, while a JEPA-style world model delivers fast latent-space planning and built-in uncertainty estimates using only 15M parameters. In the data-constrained regime, aggressive regularization plus ensembling yields a 5× data-efficiency gain over standard pre-training; distillation recovers most of the gain at inference time.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

YC Paper Club is the first of its kind, created as a community for founders and researchers. Over 1,000 people applied and roughly 100 were selected. The event is held at Pioneer, a location chosen to bring together AI talent from the Bay Area, including people at Anthropic, OpenAI, Cursor, Google DeepMind, Tesla, xAI, and Thinking Machines who may not regularly travel to San Francisco for YC events.

The speaker references their Winter 2016 YC batch experience, noting that 140 companies went through that batch and 10-15 became unicorns, including WPY, Astronis, and Deep Graham. During that period, early OpenAI discussions occurred at the same location.

Tanishk, a Stanford grad student, presents work done with Triau and Aar May. The talk focuses on inference, arguing that inference should be viewed as a capability rather than merely a cost or convenience factor. The core claim is that if a method's performance scales with thinking time, then tokens-per-second directly determines peak intelligence deliverable.

A live comparison shows three approaches on a code prompt from VLM: standard autoregressive decoding, speculative decoding, and a custom inference engine implementing a new algorithm. The custom engine achieves significantly higher speed, demonstrating that algorithmic improvements matter more than systems optimizations alone.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.