Back to Sequoia Capital

ElevenLabs' Mati Staniszewski: How Voice Becomes the Interface for AI

Sequoia CapitalMay 6, 202626m
In a Nutshell

ElevenLabs, founded in 2022 by childhood friends inspired by poor Polish dubbing, builds frontier audio AI models—including text-to-speech, speech-to-text, dubbing, voice agents, and music—achieving >$400M revenue with a flat, remote team of 400+ by prioritizing rapid launches, global talent, and data annotation from 1,000+ experts for long-term emotional intelligence. Key applications span customer support (e.g., Deliveroo, Deutsche Telekom), citizen services (Ukraine gov), and education (interactive Feynman/Ramsay agents), with viral moments like dubbing world leaders and celebrities in new languages. Founders' lessons: keep teams <10 people, embed engineers everywhere, focus on domain-specific models, ecosystems (20k+ user voices), and full-stack integration beyond raw research for defensible moats.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

ElevenLabs started in 2022 with childhood friend co-founder P (Polish names are complicated). Met in high school, became best friends, took same classes, traveled, studied, and worked together. Still best friends today. Inspiration from Poland, suburbs of Warsaw: foreign movies dubbed in Polish use one single voice for all characters, male or female, often monotone, forcing viewers to interpret emotions themselves. This persists today for most content. Highlighted need for everyone to speak any language with same emotion and intonation.

Problem spans audio domain: narrating content, books not in audio form, news articles, language barriers, future humanoids/robots where voice is primary interface.

ElevenLabs builds frontier models for audio without billions in funding. Started in 2022, year of crypto/metaverse, pre-AI boom. Audio was niche with few researchers. Advantages: excited about future, others undervalued it, audio models smaller (less compute), big data needs but solvable via transcription/annotation, architectural innovation.

Co-founder top researcher, assembled best audio talent. Untraditional: started in London with Warsaw presence, fully remote to hire globally via GitHub scraping, sharing samples. Launched product quickly, monetized early for revenue to fund models, stayed healthy margins for independence. Later raised external money as ambitions grew. Niches remain untackled for startups.

Sign in to read the full notes

Get access to AI-generated notes, topic timestamps, and more.