Back to Dwarkesh Patel

AI researchers debate how close we are to recursive self-improvement

Dwarkesh PatelSeptember 11, 20261h 37m
Topics71
Introduction to the Discussion0:00Why 2036 Might Not Feature Billions of Superintelligences0:30Human Advantages and Persistent Model Limitations1:30The Path to Fast Takeoff3:05The Risk of Asymptotic Curves4:02Deep Learning's Potential to Dominate R&D5:02Different Types of Research and Objective Specification7:03The Success of Next Token Prediction9:02AI Labor in Research Optimization10:32The 10x Speed-up Potential and Its Limits13:04Moravec's Paradox and Autonomous Research13:30The Last Human Role in AI R&D14:30Defining Objectives as the Final Human Job16:03Post-training Team Requirements17:05The Verification Bottleneck18:03Distillation as a Counterforce to Centralization19:02The Importance of Prompt Distribution20:02Automating Prompt Distribution21:31Distillation Advantages for Non-Frontier Labs23:01The Student-Teacher Gap Problem24:32Difficulty vs Realism in Environment Creation24:51Distillation and Distribution Matching26:01Post-Training Challenges27:37Training Models Capable of Automating AI R&D28:02Lineage Rollback and Self-Play29:33Environments That Exceed Human Capability31:01Types of Training Tasks33:30The Scaling RLVR Bet34:00Domain-by-Domain Training Progression35:30Transfer and Data Scarcity38:31Sim-to-Real Limitations39:03Inference Compute and Deployment Learning41:30Online Reinforcement Learning Challenges44:00Long-Horizon Task Simulation45:34Cumulative vs Non-Stationary Tasks47:02Current Model Weaknesses and Sample Efficiency48:54Thought Experiment on Context Windows and Taste50:34Meta-Learning Taste from Short Episodes51:31Hive Mind Learning and Economic Incentives52:33Learning from Data in Real Time54:00Stages of Learning Implementation54:33Continual Learning Limitations55:03Capacity Versus Technique Issues56:31The Technique Bottleneck57:00Scale and Noise Washing58:32Data and AI Progress1:00:34RL Environment Ladder1:01:03Real-World Information Asymmetries1:02:01Fine-Tuning Examples and RL Signal Problems1:03:01Signal Extraction from the Real World1:04:33Pre-Training Signal Versus Post-Training Signal1:05:30Data Versus Architecture Compute Efficiency Gains1:06:00Scale Dependence of Architecture and Data1:08:36Scale Dependence of Mid-Training and Post-Training Data1:09:32Parameter Scaling in RL-Heavy Regime1:10:37Model Size and Architecture Trade-offs1:11:47Data Efficiency and Parameter Scaling1:13:33Inference Efficiency and Hardware Constraints1:15:31Chinchilla Scaling Laws and Compute Constraints1:16:31Historical Scaling Law Challenges1:17:30RL Effectiveness and Learning Dynamics1:18:00RL Signal-to-Noise Advantages1:19:31RL Impact on Reasoning Traces1:20:31Qualitative Model Improvements1:21:34RL Generalization and Task Coverage1:24:01Creativity and Entropy Concerns1:25:35Remote Worker AI Predictions1:28:30Productivity and Research Acceleration1:31:04AI Research Acceleration Details1:32:33ASI Timeline Predictions1:34:04Closing Discussion1:36:56
In a Nutshell

AI researchers are debating whether current transformer-plus-RL paradigms can reach recursive self-improvement, with the key bottlenecks being generalization to open-ended objectives, verification of agent outputs, and continual learning without catastrophic forgetting. They predict that scaling RL across millions of diverse environments could produce drop-in remote workers within 1-3 years and ASI-level capabilities in 3-10 years, though the field may still need paradigm shifts beyond next-token prediction. Distillation and automated environment creation are accelerating progress by allowing non-frontier labs to match capabilities while reducing human involvement in defining objectives and collecting data.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Beren Millidge is CTO of Zyphra, developing open source models. John Schulman is chief scientist at Thinking Machines and previously co-founded OpenAI, leading the RLHF work that led to ChatGPT. Charlie O'Neill is head of model training at Baseten.

If 2036 lacks billions of superintelligences transforming the world, the most likely technical reason would be the absence of true generalization. AI might become extremely good at benchmark tasks without achieving the "spark of generalization." A persistent sim-to-real gap could block broader impact. This scenario seems unlikely because RL already shows generalization in practice, but if meta-learning proves impossible and continual learning remains unsolved, this would be the default outcome.

Humans maintain advantages over models in areas where models show weaker judgment and cannot check themselves well enough. A recurring cycle emerges where new models initially seem revolutionary, but after a month of use they feel limited again. This pattern could repeat more times than expected. Research and engineering remain bottlenecked despite AI writing more code than humans, preventing 100X productivity gains.

The question is how far the current paradigm of transformer plus RL is from the global optimum of a learner on a chip. An agent only 0.1% better than all humans at AI research could trigger fast takeoff through parallel execution of hundreds of thousands or millions of instances at increasing speeds. Moore's law required many discrete innovations to maintain its trajectory, and LLMs similarly moved from pre-training scaling laws to RL to overcome diminishing returns.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.