Where AI products go next: voice, agents, and self-driving software | Tara Sesha and Nan Yu (OpenAI)
In a Nutshell
OpenAI's product leaders argue that shipping imperfect AI products quickly is essential, as urgency and real user feedback outweigh polished perfection. Voice interfaces, autonomous "self-driving" agents, and direct user relationships are emerging as critical form factors and practices, while 2-3 month planning horizons and deep user empathy are now the core skills required for AI product development.
These notes were generated by AI and may contain inaccuracies.
Every AI product has two faces represented by a left toggle and right toggle. Tara Sesha explained that shipping imperfect products is essential in the current era because urgency matters and nothing compares to empirical evidence from users actually trying and using the product. She contrasted her background at Stripe, where everything was deeply polished and considered, with the current need to shift mindset from getting things perfect a priori to shipping and iterating as quickly as possible.
The toggle feature for ChatGPT is an imperfect solution but was necessary to get agentic harness capabilities to over a billion users without disrupting developer workflows. The decision acknowledges that products will evolve, and the next step might be removing the toggle entirely, with users being more accepting of change when they can follow a coherent story and have transparency about what's happening behind the scenes.
When building products at high speed, several constraints help maintain quality. The product must be additive for users, unlocking additional value that is genuinely useful. Internal testing is essential - shipping products internally first and checking for sufficient uptake, retention, delight, surprising greatness, or novel use cases. The product should aim for where models will be in 2-3 months rather than being overly anchored in the present or being overly futuristic and unusable.
Nan Yu emphasized that the biggest constraint is people's understanding and ability to absorb what is being given to them. There is significant capability overhang where models can do much more than people are taking advantage of, and this absorption gap is the real bounding process.
The statement that B2B customers cannot absorb change is partially true in that the current pace feels overwhelming to some enterprises. However, if companies aren't shipping the latest frontier capabilities, enterprises will get leapfrogged. Most enterprises were using ChatGPT merely for questions and answers while the agent revolution was happening, requiring agents to be shipped as fast as possible to unlock more value, even if it broke existing processes and conceptions about how enterprises should receive updates.
The principle is to do what users need rather than what users say they need, which applies more than ever before.
The discussion centered on whether to centralize on a single generic agent identity or create artisanal micro-agents for specific jobs. Nan Yu argued that successful efforts map onto human nature, noting that managing 40 agents is challenging and that people naturally create a chief of staff agent to manage other agents, effectively reducing the cognitive load.
Tara Sesha emphasized that practical product considerations matter significantly, including data access and permissions, how memory should behave, whether agents should use service accounts or user accounts, and how agents should manifest credentials to users. The use case determines these decisions more than philosophical positions.
Two fundamental skill sets are required: user empathy to understand mental models and how products should interface with user problems, and systems thinking to manifest products realistically through platform components and edge cases. Tara added a third element - relentlessness to try things out and keep iterating despite pain, combined with the ability to learn from feedback loops at much higher clock speed.
The platform strategy involves determining what functionality goes natively into the platform versus what should be built by first-party or third-party developers as plugins or additional interfaces. The goal is to expose hooks that allow third-party apps to be most effective, while providing layers of capabilities. If pre-built tools or ecosystem tools don't work, computer use serves as a fallback that will always work, even if slower or more expensive in tokens.
The ideal experience is that users feel like guests in a home where their needs have been anticipated, with the model increasingly providing this anticipation. Getting 99% of the way and then failing is worse than a non-starter because it creates an unfulfilled promise.
Working with research requires a wholly different approach than working with engineering. Product leaders should come with very specific use cases, clear understanding of user goals, and sample readings showing exactly what happened in sessions and why models didn't do what was needed. Writing good evaluations is the most important skill, allowing product teams to demonstrate that with specific prompts and skills, outcomes were achieved, which can then be post-trained into models from the beginning.
The ideal planning horizon is 2-3 months out. Planning years ahead leads to incorrect predictions, while building exactly for today means getting left behind. The appropriate planning cadence depends on understanding the pace and dynamics of the specific market. Stripe's business has relatively stable market dynamics compared to OpenAI's, making longer-term planning more feasible in some contexts than others.
Onboarding matters significantly more now than before. The initial experience and helping people understand the offering is critical given capability overhang and absorption challenges. Privacy considerations around user data and how it's used are also more important than ever.
As agents begin doing things semi-autonomously, users need to be able to predict what will happen with any new feature or primitive at any given point. Previously, unpredictable behavior could be managed through settings menus or popups, but this approach no longer works for autonomous systems. The importance of testing things empirically and moving faster has increased significantly. Product managers must maintain both high-level and low-level understanding simultaneously, really understand the design and product, but do all of this faster while testing with users has become much more important.
At OpenAI, the product and engineering team is highly DMable, highly accessible, and highly visible, receiving direct feedback IDs from users on a regular basis. Direct user relationships are becoming table stakes for product leaders, with the concept of the "DMable PM" being particularly visible at OpenAI. This level of accessibility, which has been standard in developer and enterprise contexts where people could always message to ask for features, now extends to large consumer products used by everyone.
The value of this direct contact lies in knowing specifically what users want and their use cases. The more accessible a PM or engineer is, the better, as this direct feedback is crucial for adding value to the product.
Engineers are extremely picky and detail-oriented, and good feedback often requires second or third follow-up questions. When designing agentic experiences with natural language interfaces, the problems people encounter are often subtle and require additional context to understand fully - such as what exactly was said, the context provided, or the user's emotional state during the interaction. Having a more direct, deeper relationship with users bears significant fruit that would not have been accessible before.
When asked about what will be different about product craft and the role in a year, the discussion focused on where products are heading in terms of form factors and user experiences. The consensus highlighted several emerging trends that will become prominent.
Voice is identified as a major upcoming trend. Voice interfaces have fundamentally changed work patterns, with natural conversation being preferred over traditional interaction methods. Voice has dramatically reduced the amount of tech support needed for family members and has proven to be an intuitive interface for onboarding teams onto new products. The models are finally getting good at voice, with speed improvements making it increasingly viable. Voice represents a natural way humans interact, aligned with how people have evolved to communicate over millennia.
Self-driving-like capabilities are predicted to become much more important and prominent. A major problem identified is the "empty input box problem" - users having a product but not knowing what to do next. Since products are now intelligent, they can use themselves to help users and provide a gentle on-ramp to functionality. This approach will become normal, with users looking back and questioning why products didn't always work this way.
Hardware that replaces carrying around laptops is predicted to emerge, with concepts like portable devices that eliminate the need for traditional computer interaction being envisioned as the next evolution in personal computing devices.
Keep Lenny's Podcast in your library
Save the videos and channels worth coming back to, and find them again in one place.





