Back to Lenny's Podcast

How to build products on a moving frontier | Dan Shipper (Every)

Lenny's PodcastSeptember 24, 202618m
In a Nutshell

Build on a moving frontier by creating a tiny labs team (1-2 people) that runs many parallel experiments and discards ~90% of them, while the product team focuses on scaling the 10% that prove valuable. Use "pirates and architects" in these small teams, dogfood every experiment, and move ideas through a clear pipeline—internal adoption, early-customer testing, then roadmap integration—only when they show repeated use, 10× improvement, and scalable cost. This structure turns new model releases into opportunities instead of threats.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

The world just changed again with the launches of Fable 5.1 and Astra 6. These models enable historically accurate 3D reproductions of events like the Battle of Waterloo from a single prompt after four hours of processing. A team member created a simulation with a thousand agents by feeding a scientific paper into the system, and Astra is now capable of performing video editing in Premiere. All animations in the presentation deck were generated by Astra without manual animation work.

Product leaders face the dilemma of whether to maintain focus on execution, ask customers what they want, or attempt dramatic product transformations. Most customers are unfamiliar with frontier capabilities and expect guidance from product teams rather than providing direction themselves.

Never make any major life decisions within 30 days of a meditation retreat, a psychedelic experience, or your first encounter with a frontier model.

Product teams must simultaneously execute roadmaps at a high level while staying at the frontier, which represents fundamentally opposed ways of working. Exploration is divergent and involves trying many different approaches with the expectation of discarding most experiments. Execution is convergent and requires focus, saying no to options, and delivering against planned roadmaps.

The solution is to separate concerns by creating a labs team. Product team members focus on improving and scaling what already works, while labs team members focus exclusively on exploring the frontier and running experiments with new models. This structure allows organizations to harness early adopters without distracting those responsible for roadmap delivery.

AI makes labs teams accessible even for small organizations because team members now have superpowers that reduce the resource requirements. A labs team of one person can effectively explore new model capabilities when they launch.

On a labs team, the expectation is to dispose of approximately 90% of experiments. The focus is on exploring new capabilities, particularly when new models are released. On a product team, the goal is to adopt about 10% of what the labs team produces, with emphasis on improving and scaling existing products for current customers.

Anthropic Labs produced Claude Code, one of the most successful productivity products, along with MCPs, Skills, and Claude Design. These emerged from a small experimental group, with Anthropic then investing more heavily in successful experiments. Thousands of experiments never reach public visibility.

Use very small teams of one or two people maximum, referred to as two-slice teams rather than the traditional two-pizza teams of eight to ten people. Larger teams create coordination overhead and conflicting visions that hinder progress in the AI context.

Team composition should include pirates and architects. Pirates are obsessed with finding value and generate many messy experiments. Architects take promising but messy systems and shape them into valuable, beautiful, and extensible products. Pairing these personality types creates powerful results.

The most important practice is dogfooding to create the tightest possible feedback loop between creation and validation. Building for personal use provides the fastest iteration cycles. When this is not possible, early adopter customers should be engaged for testing. All experiments should serve actual work requirements to distinguish between what is genuinely useful versus merely novel.

Run many experiments in parallel, including competing approaches to the same problem. When the frontier is unknown, having different team members attempt the same goal from various perspectives helps map what is valuable and what is not, even if the approach appears incoherent.

Even the 90% of experiments that do not enter the product should generate value. Convert experiments into external content showing what was tried, what worked, and what did not. This attracts customers to company content and products. Use experiments to feed early adopter programs, providing value to customers who want closer relationships. Labs teams should share capability learnings with product teams so they can incorporate frontier knowledge without needing to explore independently.

Create a research pipeline where ideas progress from lab-only experiments through stages of validation. Begin with internal adoption by other team members. If internal usage occurs, advance to early customer testing and product team evaluation for roadmap integration. Not everything reaches the final stage, but successful experiments eventually scale and release to customers.

For three years, experiments have attempted to automate copy editing for Kate, the editor-in-chief who possesses exceptional taste in edits. As the team grew to approximately 30 people, her capacity became insufficient. Historical copy edits spanning three years were downloaded and used to train Fable on her editing patterns.

The system, called Kate Bench, reached the point where capabilities became sufficient for actual use. An Every agent now performs Kate passes on drafts, filing suggested changes based on her historical editing patterns. Internal usage has begun, with measurable improvement of 12% reduction in Kate's remaining work on edits compared to the previous month.

When experiments reach internal use stage, architects refine messy but functional systems into robust, measurable platforms. A dashboard now tracks suggestion acceptance rates and remaining work per document. The system is approaching readiness for early customer release.

Review the pipeline regularly through weekly all-hands meetings where team members discuss pipeline status. This keeps product teams informed about emerging capabilities without requiring them to conduct their own experiments.

Define clear decision criteria for advancing ideas through the pipeline. Primary criteria include whether people are using the tool and returning to it, whether it represents a 10x improvement over existing solutions after sustained use, and whether it is affordable to serve at scale. Making advancement decisions into formal events maintains rigor in the process.

A small team built Codex outside OpenAI's main application while other teams simultaneously explored future coding interfaces. After launching a desktop app in February 2026, Codex grew rapidly enough to merge into ChatGPT, becoming its foundation. The Codex team ultimately took responsibility for an 800-million daily active user application through successful experimentation.

This methodology enables building the next version of products while continuing to scale existing ones. Success is indicated when new model releases become welcome events rather than sources of dread.

Keep Lenny's Podcast in your library

Save the videos and channels worth coming back to, and find them again in one place.