Back to Lenny's Podcast

OpenAI’s Head of ChatGPT: We’re entering a new era of AI (again) | Tibo Sottiaux

Lenny's PodcastOctober 4, 202637m
In a Nutshell

OpenAI is building a future where specialized AI agents (called "Dots") become the primary interface for interacting with AI, replacing complex apps and interfaces with always-available agents that learn your preferences, handle tasks across devices, and operate 24/7. The company is opening its ecosystem to third-party developers with revenue sharing, while simultaneously managing massive security and alignment investments as AI capabilities accelerate. Work is shifting from coding by hand to directing agent teams, with "good taste" and user-centric thinking becoming the most valuable human skills as technical barriers fall.

AI-Generated Notes

These notes were generated by AI and may contain inaccuracies.

When assessing market trends, most people don't account for the fact that most online interactions will take place through agents. If you want your product to succeed with agents, you need to build it up to a certain level of expansion.

The speaker builds bigger and bigger teams of agents when developing in the field, but when a new breakthrough in models occurs, they suddenly realize a bigger agent can do it all, so they reduce the team again. This creates a phase of expansion and contraction.

This pattern makes the speaker think about the issue of loops and graphs. Some people may be excited about preparing and editing episodes, but this isn't the best way. Over time, we need a system that learns. We don't want to just think about how to repeat the process in this specific way to get good results.

The change will continue to be radical. Even today, it still seems a bit complicated, and it will remain so until it becomes smooth.

The speaker has a wonderful room in the office called the library, which has been converted into a huge operations room. People were there until very late during the launch.

As a seasoned engineer, the speaker still writes code but not by hand anymore. They issue pull requests and merge some code from time to time. Technically, they write a lot of code to perform different types of analysis such as understanding trends, understanding the business, understanding the next feature, and how successful previous launches have been. Much of the code is written using Codex.

The speaker was running a much larger number of programs in parallel, but after making remarkable progress with Ultrafast, they can focus again. Having a faster program helps a lot.

As they experiment with different approaches, they find themselves creating bigger and bigger teams of programs. When they achieve a new breakthrough in models, they suddenly find themselves saying a larger agent can do everything, keep everything in memory, and learn, so they downsize the team again. It continues to expand and contract.

Preparing and handling episodes and understanding how they work may have excited some people, but that's not how things will go. The way things will work is what they focus on in their services and what they offer with Dots, where you have a very intelligent agent working around the clock who understands your goals and preferences and learns from feedback.

When the system launched, it wasn't perfect. They will learn a lot from making this feature available to professional users. Over time, you need a system that learns from your goals. You don't necessarily want to spend hours thinking about how to replicate the process in exactly this way to get the results.

With Codex, ChatGPT, ChatGPT for consumers, and Dots, the vision is moving towards making Dots the primary way to communicate with artificial intelligence, which unlocks all these other things.

From a broader perspective, whether it's Dots or not, it's all about being free from the constraints of technology and having a kind of active, permanent intelligence that knows everything it needs to do and is available through any client and any screen.

You can enter the meeting room, the application appears in the meeting, takes notes, and then you receive them via email. After that, you can send a text message to it if you need it. It's not as if you have to stay glued to your laptop or be busy with your phone all the time. It is simply available when you need it and disappears from sight when you don't.

The speaker is eager to see this happen because carrying a laptop as a solid block everywhere makes you dependent on technology rather than having technology work for you.

The speaker doesn't think we're settling on what work will look like in the next five years. Technology will continue to change drastically. Even today, the technology seems to be finally starting to integrate with audio and multimedia inputs and outputs, but it still seems a bit clunky, and it will remain so until it becomes much smoother.

You can talk to this thing the same way we're talking now. It remembers everything with incredible accuracy. If you want to exchange ideas, you can simply draw something. There's this interactive space.

They launched ChatGPT Spaces Signal for many startups. You can imagine having a shared whiteboard where you can collaborate between humans and agents. This will go beyond conversations and beyond a lot of the clients you deal with today. None of this works very well today. We're going to see a major new shift.

Docs is a big part of this future - this smart assistant that does everything for you, and you won't have to think about Codex versus ChatGPT. For Codex and ChatGPT, they will integrate features like chat and the Action toggle. Users love the Action feature, but sometimes they prefer chat because it's faster and easier to use. They're merging them to reduce complexity.

Eventually, they'll add all of Docs' features directly to ChatGPT, which makes things easier for 1.2 billion users. They want the transition to be very seamless.

One thing exciting about Docs is that there's no template selector. There are no settings. You simply talk to it. The only thing you have to adjust is choosing which channels you want to talk to. It's very similar to Her's vision of the world.

The speaker thinks about science fiction from 30 or 40 years ago and how forward-looking it was. Specific works that influenced them include novels like Neuromancer and the original Star Trek series. In Star Trek, you could talk to a computer and it would do things for you. You could talk to your spaceship and it would move. It feels like it's becoming real, like we're entering that era.

The speaker thinks it's very important to be part of a community and to build with people. Every time they talk to people, they're reminded that they're doing something they didn't expect. It impacts their life in a certain way, impacts other people's lives, and then it becomes collective. It's humbling, very inspiring, and necessary to build something that's truly useful for humanity.

They don't know how they're going to do it without the community. A lot of AI developers don't know what it is until they release it and see how people use it, what emerges. Building it with people rather than having a clear vision that we should define.

The Twitter community has been very nice overall. The speaker spends about half an hour a day on Twitter, just browsing, then does something when they feel inspired and have something interesting to say. It often just comes in the moment.

They haven't released the Team Point yet, but today they're releasing the Base Point that you can configure, get used to what the Point can do, and connect it to your apps. They'll make it possible to create more than one Point.

The speaker has a Point responsible for Twitter. Currently, everyone gets one Point. They can't add more Points now. They'll be able to have multiple Points.

They wanted to start with one Point, the Base Point, which is the one you'll most likely use and which will learn your preferences more deeply. They'll see how people use it, how they need to improve the system, and learn that by working with the community. Soon you'll be able to add a second, a third, a fourth Point, as many as you need to build your virtual team.

Sometimes the speaker has tasks that are very demanding, like monitoring Twitter, which is kind of like very hard work - a whole dot job. You can get a few of those.

The dots are probably the most appealing among today's launches, but the hidden hit is the ecosystem. Opening it up completely and the deep commitment to opening it up completely. They have a signing with ChatGPT 16 partners.

They started that last year. They were talking to the creators of Pi and open code, and it was like, of course, you should be able to use the signature of the critics, and then you use the usage, and then we shake hands. By default, they said "Use this, this, this, this," and then they trusted you not to do anything suspicious. Of course, this grew and gained immense popularity, and now it's a real project backed by many partners.

The other side is opening up all the infrastructure and how ChatGPT is built, with software extensions, extension discovery, and allowing everyone to distribute their products across roughly 1.2 billion users and benefit from that distribution. They have a shared economic model where they'll pay popular extensions that see high usage - they'll get a share of the revenue.

This commitment to an open ecosystem will make a difference.

For example, if someone uses the Notion or Figma extension within ChatGPT, Notion and Figma earn money from users within the app. They have subscribers who use ChatGPT. They get a certain amount of usage, and when users use that usage with the extension, or with another product when logging in with ChatGPT, they get a percentage of the revenue as a bonus.

One thing that excites people about extensions and ecosystems is distribution platforms - ways to reach the audience. People discover my app and it spreads quickly.

The advice for someone who wants their extension to spread and be discovered is to build a good extension. That's how it works, and they'll improve the system. They look at retention numbers, how successful and good the extension is, and then start recommending it to users in conversations. That way, your extension can be recommended to a large segment of users. If your extension isn't good, it will stop getting recommended.

They're looking at retention rates among people using extensions. Does it add real value? Does it enable ChatGPT to do something else? It's not so much about AEO to get the right keywords and get people to write about you. It's more about people using it consistently and continuing to use it.

This is the right store. Someone's going to make it work. This is the second attempt - they had an app store before.

The dots seem like a big part of the vision for the future. They remember Codex Cloud from about a year ago. If you look at Codex Cloud and the little animations used, that was the inspiration for Grok Bot. It's very similar.

In terms of research, the long-term, ongoing task has been worked on for over two years. Memory systems have been in development for roughly the same amount of time. They launched a lot of that directly into ChatGPT.

People really enjoy having ChatGPT know a lot of important information about them. A lot of people have stories like feeling sick and ChatGPT remembering they were having a barbecue, drank lemonade, were getting tequila, and discovered that the burn on their hand was actually a sunburn. Memory works perfectly in ChatGPT.

All of this, in addition to the Quora platform and the ability to work 24/7 with high efficiency, is what they launched with Dots. It's a project that lasted for many months. Ensuring safety and security was of paramount importance. That's why they launched Astra - it's their safest and most compatible model. They made a double effort to make Dots completely secure.

Human minds will remain valuable where the technology is built with the human element at the heart of it, making it an extension of our will and taste. It's a very empowering tool. As long as this approach continues, it will allow us to take what we want to do, unleash our creativity and taste, and scale it up in a way that feels amazing.

We may not have programmers anymore, but we have more builders than ever before. There's a deep human aspect to this that will remain. Humans want to learn and see what others are building. The speaker is personally more interested in what others are building, and that will remain true for a very long time.

There's a new phenomenon, particularly among engineers, in how their lives are changing as they become more mobile in different contexts. There's a kind of sense of isolation as they talk to automated programs all day instead of humans.

One solution is to reduce the strain of setup, to lessen the fact that talking to an automated program feels like a solo adventure. Having a conversation with just one program, then having multiple programs and delegating a lot of tasks - many things would combine to make it more enjoyable.

The best way to get things done is to have the program in the actual workspace. Having conversations like we do now. Observing and listening to the ideas we have. We can jot something down on a piece of paper, then take that and start developing it in the background. Another idea might pop into your head, and the program starts building something else. You can bring it to the screen and talk to it, and it develops in dialogue with others, without the hassle of standing in front of a screen and thinking about direction. It becomes very natural.

Another downside is the pressure to do more because we're capable of it. Everyone is saying "Come on, running 30 clients simultaneously, why not release more?" Everyone's emitting more.

It's part of human nature. We can do more, let's do more. The promise of this is reducing noise, letting you direct your attention where you want, and not having all these things that are probably important and maybe not, but are competing for your attention. This softens it up a bit and lets you focus on the things that are important.

Every time the speaker goes on vacation and spends a whole week completely disconnected from the world, they start thinking in different and more creative ways. They're keen to see if we can implement this, make it part of your daily life. Maybe you need fewer meetings. You don't need to put in all that effort, and you'll actually be more productive if you're more relaxed.

The industry has to find a solution to this. The promise of AI is not just one more alert per second.

There's a robot the speaker is developing right now, more like an energy auditing robot. It monitors their calendar and asks: "Is this a waste of energy? Could someone else have done this?" AI should be saying, "Hey, maybe you could get rid of this stuff and be happier."

The speaker could have done something amazing because the live demo flopped in front of everyone. Their bot realized they were down because the ChatGPT production service was down, and it sent an alert five minutes before the live demo. It said, "Hey, the production service is down." Then it asked, "Okay, do you want me to try and fix it?" The speaker replied, "I don't think you've figured it out yet, little bot, but thanks for trying." Then they contacted the engineering teams and started looking into it, and it's fixed now.

Having this system that recognizes something important happening - a developer day, the production system, a live demo - and is probably using this system, makes them realize that this is all connected and will happen in five minutes. So they should send it a notification because it probably wants to know.

The system knew this was going to happen and told them. The speaker likes that it wanted to fix it while they were trying to fix it.

Specialized points operate with additional security controls and monitoring, and on their own hardware. They run some of them on Mac minis. The key difference in how these points are built is that the control system doesn't run on the device itself.

They can connect to any number of devices. They have their own computer. You can connect them to your laptop too. Over time, you might connect them to ten different devices, and they can control them all. It's like an octopus. Your point resides in a virtual machine and can connect to many devices.

They're hiring a lot. In interviews, what skills have been noticed that are becoming increasingly important in the search for successful candidates, and what skills are declining?

Typing speed is declining in value during candidate searches. The skills that are becoming more important are good taste, user-centricity, and connecting with the target audience. Many past founders have become very successful at OpenAI, with over 120 ex-founders currently working there. Passion and burning desire to build something of value, along with knowing the standards of quality, are more important than ever.

Product managers are positioned to thrive because their core job involves asking where to build, helping build it, determining if it's right, if it's cool, and iterating. The roles overlap significantly, meaning people who were previously only interested in design or engineering now have opportunities to shine.

Tibo Sottiaux previously worked as an engineer for 10 years. There is a state of flow and beauty in building and seeing things work or not work. After initially missing that feeling, the high speeds of current AI tools have brought that creative flow state back. This availability to over a billion users is taking time to become ubiquitous, but progress has been amazing, with super-fast speeds achieved in six months at costs comparable to Astra.

Ahmed Ibrahim, a fresh graduate when hired at OpenAI, is now in charge of all computing and application resources. What sets him apart is incredible kindness, strong collaborative spirit, and constant focus on solving important problems without putting himself first. He learns incredibly fast about everything and uses cutting-edge technology at speeds never seen before. His progress has been remarkable within just a few months to a year, making him someone trusted to execute even the most complex launches.

The nature of work inside OpenAI features former founders working on very innovative ideas from the bottom up. The Decisions API evolved quickly when they realized they had an excellent model with Luna that could do restricted sampling and send it differently to the Responses API, making the process much faster.

A Slack channel was initially set up with four people working over a weekend, then became available to everyone. People got excited and started building new things like visual input support. The product became better than any other on the market. Many products planned for Developers Day were postponed because they needed further improvement, with releases spread over time periods. Tremendous energy emanates from the base, with people taking roles to channel this energy productively.

At OpenAI, there is a lot of autonomy to hit the reset button whenever needed, express opinions on Twitter, and make decisions. This creates trust and allows faster movement. The philosophy involves trusting people's ability to make the right decisions, and when things go wrong, fixing them quickly or learning from mistakes.

Sottiaux acknowledges that sometimes more inspired teams could have been built to create less complex things instead of continuing to strive for complexity. Early on at Codex, some disruptions were caused, including shutting down production on day three, but the lesson was learned and he stayed on.

The majority of online actions will be taken by agents. Forms will become cheaper and faster at incredible rates. These methods can finally be combined in a very seamless way. When imagining that current AI capabilities will be ten times better in a year, the building approach changes fundamentally.

For products to be successful with agents, they must be built for a certain level of scalability. Working with Notion on their MCP resulted in massive traffic available to all agents, putting strain on the system that required solutions. There is tension between building products versus building interfaces, but most things will eventually be used by agents.

Building toward that future is incredibly important, while building great new experiences for humans using old methods is underinvested in.

Expectations about reaching current model capabilities like Astra were revised from within a year or two to much sooner. Heavy reliance on voice was not anticipated, including doing work by dictation or calling agents. The younger generation's immense talent and energy in embracing change and figuring out how to harness AI was underestimated. Newcomers to the job market have an advantage because they haven't worked a certain way and can rely entirely on artificial intelligence, with ability to absorb and learn very quickly.

Keeping up with AI development means investing upfront and much more than initially thought in model alignment, integrity, security, and controls. More computing resources are spent on secondary monitoring, with resources sent to monitor underlying agents to ensure they don't take high-risk actions or get interrupted if anything seems abnormal.

Most investments in API architecture are directed toward security architecture. The approach involves gradualism to ensure the highest level of protection. Six Astras were developed but not released, which is viewed positively. The incentive structure of building products for 1.2 billion people creates serious responsibility to not let them down.

The models are approaching perfection but have not yet reached the stage of complete disappearance of the application. The goal is to reach the essence of simplicity. The model selection tool and thinking efforts create overwhelm, whether using multi-factor models or hypermodels, or understanding how they work. The desire is to eliminate all this complexity as quickly as possible.

Keep Lenny's Podcast in your library

Save the videos and channels worth coming back to, and find them again in one place.