8 Predictions for the Era of Continual Learning
In a Nutshell
Continual learning makes AI safety, regulation, and business models obsolete overnight: instead of static weights checked before deployment, models will improve daily from millions of real work sessions, requiring ongoing risk inspections and new techniques to preserve safety across weight updates. Leading labs gain accelerating moats through switching costs and economies of scale, as personalized weights become too expensive for small users and enterprises face lock-in or fall behind. The result is faster deployment pressure, greater AI diversity across domains, and a shift from pre-deployment alignment research to maintaining values during continuous learning.
These notes were generated by AI and may contain inaccuracies.
The speaker explains why actual continual learning is needed for AIs to perform whole jobs competently. The current approach of AIs writing markdown files from session to session is insufficient. An illustrative example compares this to students learning saxophone: each student enters the music hall, fails on their first attempt, writes notes about what went wrong, and leaves these notes for the next student waiting outside. With an infinite number of students outside, no sequence of text could allow subsequent students to nail the saxophone from the first try. The relevant experience must accumulate in the brain. The same principle applies to AIs needing to accumulate experience from different workplaces.
Current AI regulation proposals assume a model is trained and then deployed, allowing pre-deployment checks to prevent issues like aiding cyber attacks. This assumption won't hold when models improve daily based on millions of work sessions. Locking in safety regulatory regimes now risks creating archaic, counterproductive approaches. Government safety evaluations should involve monthly or quarterly risk inspections rather than focusing on a special moment after training but before deployment, as this distinction won't exist in the future.
Current technical alignment research focuses on ensuring frozen weights behave well during deployment. There is little research on maintaining safety with constant weight updates—preventing jailbreaks or shifts to deceptive personas. When AIs consolidate learnings between users, preventing malicious backdoors or inclinations injected into the base model becomes critical. This mirrors the human alignment problem where children learn new things but ideally retain common sense and basic values to avoid adopting weird beliefs or misanthropic ideas.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.