AI researchers debate how close we are to recursive self-improvement
In a Nutshell
AI researchers are debating whether current transformer-plus-RL paradigms can reach recursive self-improvement, with the key bottlenecks being generalization to open-ended objectives, verification of agent outputs, and continual learning without catastrophic forgetting. They predict that scaling RL across millions of diverse environments could produce drop-in remote workers within 1-3 years and ASI-level capabilities in 3-10 years, though the field may still need paradigm shifts beyond next-token prediction. Distillation and automated environment creation are accelerating progress by allowing non-frontier labs to match capabilities while reducing human involvement in defining objectives and collecting data.
These notes were generated by AI and may contain inaccuracies.
Beren Millidge is CTO of Zyphra, developing open source models. John Schulman is chief scientist at Thinking Machines and previously co-founded OpenAI, leading the RLHF work that led to ChatGPT. Charlie O'Neill is head of model training at Baseten.
If 2036 lacks billions of superintelligences transforming the world, the most likely technical reason would be the absence of true generalization. AI might become extremely good at benchmark tasks without achieving the "spark of generalization." A persistent sim-to-real gap could block broader impact. This scenario seems unlikely because RL already shows generalization in practice, but if meta-learning proves impossible and continual learning remains unsolved, this would be the default outcome.
Humans maintain advantages over models in areas where models show weaker judgment and cannot check themselves well enough. A recurring cycle emerges where new models initially seem revolutionary, but after a month of use they feel limited again. This pattern could repeat more times than expected. Research and engineering remain bottlenecked despite AI writing more code than humans, preventing 100X productivity gains.
The question is how far the current paradigm of transformer plus RL is from the global optimum of a learner on a chip. An agent only 0.1% better than all humans at AI research could trigger fast takeoff through parallel execution of hundreds of thousands or millions of instances at increasing speeds. Moore's law required many discrete innovations to maintain its trajectory, and LLMs similarly moved from pre-training scaling laws to RL to overcome diminishing returns.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.