Top AI News: Sonnet 4.6, Grok 4.2, Gemini 3 Deep Think, and OpenClaw | EP #231
In a Nutshell
AI model race accelerates with Anthropic's Sonnet 4.6 leading benchmarks like GPQA via performance gains at fixed prices, xAI's multi-agent Grok 4.2 beta, and Google's Gemini 3 Deep Think achieving 48.4% on Humanity's Last Exam plus Olympiad golds—all slashing costs 400-1400x and cooking knowledge work, math, physics, and coding. OpenClaw enables addictive 24/7 local AI agents (now OpenAI-backed), powering agent economies, Coinbase wallets, and Chinese integrations like Kimi Claw, amid exploding energy/data center demands (80GW US need) and job shifts (IBM triples AI-wrangling hires as juniors vanish). India surges as OpenAI's 100M-user land grab bellwether for AI-driven growth, while warnings flag OpenClaw security risks, benchmark saturation, and path to ASI via recursive self-improvement.
These notes were generated by AI and may contain inaccuracies.
AV is hard. They're struggling to get AV working in Germany at midnight. Sem has taken over. They're figuring out screen sharing with audio. Testing outro music. Dave was not AV certified in school. They preview the deck backwards, test video playback. Time loop glitch from Nick in the room. They're live. Welcome to raw backstage chaos at Moonshots. Good morning, good afternoon, good evening to another episode of WTF just happened in tech with DB2, Selen, Ismael, AWG, PhD in Germany in Stuttgart to get your future ready. Episode covers multbots, race between hyperscalers, dive into energy data centers. Jump into supersonic tsunami. Singularity is now. PhD in Stuttgart for longevity treatments. Did a pilgrimage to Stuttgart to visit Porsche Museum.
Race and leapfrogging continues between Sonnet 4.6, Grok, in living color. Sonnet 4.6 interesting release. Anthropic pioneering scaling phase space: keep prices of model tiers same but increase capabilities. Sonnet 4.6 same price per token as Sonnet 4.5 but increased capabilities. OpenAI reducing cost per token while keeping capabilities constant through distillation and other processes.
Progress on benchmarks astonishing. On GPQA benchmark (gross domestic product eval that OpenAI launched), Anthropic leading with Sonnet 4.6 (not even Opus 4.6) at state-of-the-art on GPQA and another eval for knowledge work. Knowledge work is cooked, charbroiled thanks to Sonnet 4.6. Computer use becoming killer app; Sonnet 4.6 state-of-the-art on handful of computer use benchmarks. Anthropic's thesis of focusing on software engineering and code generation as critical path to recursive self-improvement working versus distractions like images and video generation. Can accomplish borderline magical tasks with Opus 4.6.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.