Back to Dwarkesh Patel

Building AlphaGo from scratch – Eric Jang

Dwarkesh PatelMay 15, 20262h 37m
In a Nutshell

Eric Jang rebuilt AlphaGo from scratch using modern compute, achieving strong Go bots for ~$10K via neural-enhanced MCTS—combining policy/value networks with PUCT-guided search for efficient policy improvement through self-play and distillation, far cheaper than DeepMind's original millions. MCTS elegantly sidesteps RL's high variance by iteratively relabeling states with search-improved soft targets, enabling stable supervised learning that hill-climbs win rates without exploration failures or long-horizon credit assignment issues plaguing naive policy gradients in LLMs. Key takeaways: initialize with expert/small-board data, prioritize bug-free baselines before scaling laws, and leverage LLMs for hyperparameter tuning/experiments while humans supply research taste for track selection.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Building AlphaGo from Scratch – Eric Jang

Eric Jang was most recently vice president of AI at 1X Technologies, and before that, senior research scientist at what is now Google DeepMind Robotics. He's been on sabbatical for the last few months, rebuilding, improving, and hacking on AlphaGo. The discussion covers building AlphaGo from scratch and what it tells us about the future of AI research and development.

Why AlphaGo is interesting: AlphaGo got Eric into the field. Early breakthroughs in 2014, 2015, 2016 showed how smart AI systems could become and the computational complexity class they could tackle with deep learning. Go was long understood to be intractable for search, yet solved through deep learning, which was mysterious.

Eric's training is in deep neural nets for robotics, where decisions are intuitive. But AlphaGo involves very deep search, mysterious how a ten-layer network amortizes simulation of something deep in the game tree.

Plot of compute for strong Go bots: In 2020, open-source KataGo by David Wu from Jane Street achieved 40x reduction in compute to train a strong Go bot tabula rasa. Not certain if stronger than AlphaGo Zero, AlphaZero, or MuZero, but very strong; most Go practitioners train against it today.

Thanks to LLM coding, what took DeepMind team, millions of dollars, research and compute, now done for few thousand dollars of rented compute.

Go is simple game, easy to implement on computer. Objective: put down black and white stones to occupy as much territory as possible.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.