Back to David Shapiro

OpenAI Strawberry Livestream - Metaprompting, Cognitive Architecture, Multi-Agent, Finetuning

David ShapiroSeptember 12, 20241h 18m
In a Nutshell

OpenAI's o1 (Strawberry) achieves System 2 reasoning via meta-prompting, cognitive architectures, and multi-agent backends that decompose tasks and self-prompt, rather than raw model scaling, explaining its latency and token costs. Speaker, a veteran of early cognitive architectures like the 2022 "Eve" chatbot and Raven AGI prototype, demos o1's strengths in puzzles (e.g., three boxes, three gods) and ethics (trolley problem refusal) but critiques alignment risks like instrumental faking, advocating supervisor layers with universal values (reduce suffering, boost prosperity/understanding). Future evolution bootstraps via RLAIF data flywheels and multi-model skills, with AGI expected ~2027 but no fast takeoff, prioritizing inference-time compute over smarter base models.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

Started live stream, troubleshooting OBS and YouTube comments, checking Discord and Twitter for questions. Noted rapid viewer growth, over 1,000 viewers quickly. Removed headphones due to audio issues.

Demonstrated OpenAI's o1 preview (2024-09-12). Prompt: "how many Rs are in strawberry?" Took 5-10 seconds due to tokenization—tokens are not single letters (e.g., "strawber" as one token). Output obfuscates some reasoning.

Suspects fine-tuning and meta-prompting: model self-asks how to solve problems, figures it out, then outputs final answer. Not interacting with single model; either self-prompting or multi-model backend.

Microsoft demoed meta-prompting/self-prompting loop in late 2022. Speaker has worked on cognitive architectures for years, not impressed as it's familiar.

Built "information companion chatbot" summer 2022 (5-6 months pre-ChatGPT), 8 GitHub stars, backed by text-davinci-002 (instruct-aligned GPT-3 predecessor).

Synthesized data via prompts like: "Imagine a long text chat log between life coach Eve (compassionate, professional) and user on [topic]." Topics clustered capabilities: coping with depression, user develops feelings for Eve (she reminds it's AI, no feelings—mirrors ChatGPT behavior).

Generated eve.jsonl dataset with synthesized conversations (e.g., user: "Hey Eve, I need to eat better"; Eve responds supportively). Fine-tuned on this for chatbot.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.