Back to Sequoia Capital

Robotics' End Game: Nvidia's Jim Fan

Sequoia CapitalApril 30, 202620m
In a Nutshell

NVIDIA's Jim Fan declares robotics entering its "endgame," mirroring LLMs' progression: pre-train world models on vast video data to predict next physical states (WAMs like DreamZero), align via action fine-tuning, and scale with RL—outperforming VLAs by prioritizing physics/verbs over language. Data shifts from costly teleop to scalable sensorized human egocentric videos (EgoScale: 21k hours pre-training yields dexterous policies) plus real-to-sim synthesis (DreamDojo) and wearables like DexOOI, unlocking neuroscaling laws for dexterity up to 10M hours/year. Predictions: physical Turing Test in 2-3 years, API fleets for factories/labs soon after, full agentic auto-research by 2040.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Andre introduces Jim Fan, who leads the embodied autonomous research group at NVIDIA, known as NVIDIA Robotics. Robots are thrilling; cars are big robots, but excitement is for robots that can lift things. Jim was a standout at last year's AIN.

Jim recounts a summer day in 2016 in this office: a guy in shiny leather jacket with big biceps delivers a large metal tray inscribed, "To Elon and the OpenAI team, to the future of computing and humanity, I present you the world's first DGX1." First time meeting Jensen. As a good intern, Jim signs his name on it (visible in photo), along with Andre. They joke about going to the computer history museum, feeling like dinosaurs. Back then, no clue what they were signing up for.

If you believe in deep learning, deep learning will believe in you.

Deep learning believed in them big time. Three-step functions in six years: (1) GPT-3 pre-training on next token prediction learns grammar and shape of language, simulating how thoughts, code, and strings unfold. (2) 2022 InstructGPT supervised fine-tuning aligns simulation to useful work or reasoning. (3) Auto research accelerates the loop beyond humanly possible.

All labs are at the final boss fight for LLMs, in the end game. OM folks speedrunning AGI on mythical creatures called mythos.

Why can't robotics have fun? Jim copies homework and calls it the great parallel: simulate next physical world state, align through action fine-tuning to a thin slice that matters for real robots, let reinforcement learning carry the last mile. If you can't beat them, join them. New episode: Robotics, the endgame.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.