Back to Dwarkesh Patel

Dario Amodei — The highest-stakes financial model in history

Dwarkesh PatelFebruary 13, 20262h 22m
AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Broadly speaking, the exponential of the underlying technology has gone about as expected, plus or minus a year or two. Didn't predict the specific direction of code, but the march of models from smart high school student to smart college student to beginning PhD and professional stuff, and in code reaching beyond that, matches expectations. The frontier is uneven but roughly as expected.

Most surprising: lack of public recognition of how close we are to the end of the exponential. Wild that people inside and outside the bubble talk about tired political issues when near the end of the exponential.

Scaling hypothesis remains the same as in 2017, from doc 'The Big Blob of Compute Hypothesis' written when GPT-1 came out. Not specific to language models; covered robotics, reasoning separate from LMs, RL scaling like AlphaGo, Dota at OpenAI, StarCraft/AlphaStar at DeepMind. Aligns with Rich Sutton's 'The Bitter Lesson'.

Hypothesis: all clever techniques don't matter much. Key factors:

  • Raw compute
  • Quantity of data
  • Quality and distribution of data (broad distribution)
  • Training duration
  • Scalable objective function (pre-training, RL with goals; objective rewards in math/coding, subjective in RLHF)
  • Normalization/conditioning for numerical stability so compute flows laminarly

Pre-training scaling laws continue; gains persist. Now seeing same for RL: pre-training phase then RL phase. Other companies report log-linear performance on math contests like AIME with RL training duration. Seen across wide variety of RL tasks.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.