The whole AI safety drama, explained.
In a Nutshell
AI companies have created unstable systems the world now depends on, so their "safety" talk is just bug-fixing to avoid a Chernobyl-level PR disaster. Ex-employees like Jacob Coxon and Evan Hubinger warn that the race to self-improving superintelligence could kill everyone, while critics dismiss this as hype over normal software risk. The real issue is dependence on non-deterministic code, not sentient AI.
These notes were generated by AI and may contain inaccuracies.
All this AI safety stuff is not something the general public needs to worry about. The speaker argues that AI safety discussions are private matters between software companies trying to fix bugs in their products. It has nothing to do with the larger world or regular people's business. The speaker questions why Barack Obama is tweeting about this topic and notes that it creates content opportunities for everyone.
Jacob Coxon, who worked at both OpenAI and Anthropic, tweeted that he resigned from Anthropic after spending three years doing pre-training research. He stated that neither company is acting responsibly, claiming they are racing straight to self-improving superintelligence and gambling with people's lives.
Evan Hubinger, former alignment science lead at Anthropic, responded to Coxon's tweet saying Jacob is correct and that they earnestly believe AI could kill all humans. Hubinger stated he personally thinks there is greater than 10% chance of this happening within the next decade, then noted he needs to get back to work building AI systems.
The speaker criticizes the AI safety community for trying to spook people into believing software can kill all humans, suggesting they have a death fetish. Paul Cristiano joined the OpenAI board and stated that building superintelligence without robust alignment would result in permanently losing control of it, which could lead to most people dying.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.