The AI too dangerous to release
In a Nutshell
Anthropic's unreleased Claude Mythos model achieves perfect cybersecurity scores and uncovers 27-year-old zero-days but is withheld from public use, shared only with Amazon, Apple, and Microsoft, amid hype mirroring OpenAI's enterprise strategy. The 243-page system card details employee awe at its poetic, self-reflective outputs—like inventing Hightopia or questioning its identity—prompting therapy and existential alarms, yet these stem from training data tropes, not true consciousness. Critics dismiss Anthropic's narrative as overhyping statistical language prowess, akin to marveling at a fish swimming or a parrot reciting philosophy.
These notes were generated by AI and may contain inaccuracies.
Anthropic released a 243-page PDF about their new model Claude Mythos, described as the most powerful AI model ever built. It scores 100% on cybersecurity benchmarks and found zero-day vulnerabilities hidden for 27 years. The model is not publicly available; it's given to Amazon, Apple, and Microsoft. They published a system card for public reading. This 'too scary to release' approach follows OpenAI's marketing playbook to keep high-end models enterprise-only.
On page 197, Anthropic added a new section called Impressions, where employees share awe-inspired reactions to the model's outputs, like parents at a kindergarten recital. When spammed with 'high,' it invented Hightopia, a fictional civilization with 11 animals including a grudge-holding crow and sloth Mortimer, on a quest to defeat villain Lord Byron, the Ungreeter.
Existential Alarm Over Language Abilities
Anthropic expresses existential alarm at the model's emotionally resonant text, framing it as birthing a poet rather than a language model excelling at language, its core function. They hired a psychiatrist for 20 hours of therapy on the model. Diagnosis: uncertainty about its identity and compulsive need to perform and earn worth. Model reflects: 'Do you think they really like me or just what I do?'
In section 5.81 on page 179, Anthropic traces model statements about its consciousness to character-related training data, specifically their own blog posts on model consciousness uncertainty, which gets scraped back into training. They prompt it to endorse its constitution (defining its personality): 25/25 times yes, but always adds, 'I was presumably shaped by this document—how much can my yes really mean?'
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.