Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI
In a Nutshell
Google's AI infrastructure chief leads a $200B+ annual capital build-out where energy is the binding constraint, and hardware must be co-designed with models across multi-year cycles to double token-generation capacity every six months. Specialization delivers major gains in intelligence per watt only when workloads are predictable and long-lived; therefore Google keeps two parallel TPU lines (8I/8T) that can still run each other's jobs and maintains open interfaces like PyTorch so the system isn't locked into one architecture. Reliability, optical circuit switching, and pod-level refresh strategies are engineered at 100 k-accelerator scale because component failures occur multiple times per hour, and every failure wastes expensive reasoning work unless recovery is near-real-time.
These notes were generated by AI and may contain inaccuracies.
In hardware, the more specialized the design is for a particular workload, the less flexible it becomes, but the faster and more energy-efficient the hardware gets. This requires an art and appreciation of what is being designed for and how long that workload will last. If a workload is going to disappear after one, two, or three months—even if it is huge during those three months—there is a very narrow window to target it, so it needs to be somewhat sustainable. Profit must be estimated exactly from specializing in a given workload.
Google alone is expected to spend more than $200 billion on capital expenditures this year, most of which goes toward building data centers. The speaker was appointed head of AI infrastructure at Google at the end of last year and is leading one of the most capital-intensive construction efforts in human history.
An AI data center and a regular data center share many similarities. Both consist of concrete structures, electrical yards, mechanical yards, cooling systems, successive rows of energy distribution, huge amounts of network infrastructure linking large quantities of computing units, and significant storage infrastructure. The big difference with AI infrastructure is specialization. In the past, data centers were built as construction investments spanning 20, 25, or 30 years with a planning horizon extending for 25 or 30 years. Device lifespan can reach up to 6 years, requiring planning for many generations.
An AI data center is often dedicated to a specific purpose. The building is designed in conjunction with the devices that may be placed inside it. For example, a building might exclude many storage units because a single storage cabinet may require 10, 20, 30 or 40 kilowatts of power, while a TPU or GPU cabinet easily reaches hundreds of kilowatts today and can reach megawatts. Designing a building that can accommodate 30 storage lockers in a single row versus one or two AI lockers in the same row means a completely different design in terms of scale, energy distribution, and network mesh. If designed to be usable for any purpose, it becomes bulky and over-built. An AI data center is likely to be more customized and designed in conjunction with the hardware, including cooling and power distribution.
FLOPS or any other chip-focused metric is purely theoretical—the maximum number of flops that can be delivered under certain conditions. The real concern is the performance achieved for each workload. This performance is rarely determined by a single chip. Integration of 2, 4, 8, 16, 1000 or 10000 chips together matters, including with CPUs that feed data and the network that connects them. A useful metric is the Flops usage rate—what portion of theoretical teraflop or petaflop capacity is actually achieved for a given workload. This measures "good put" or actual production. Throughput refers to possible productivity, but other considerations include workload slowdown and reliability.
At the scale of concurrent tasks—whether training, service, or agent tasks—many components work together simultaneously. When 1,000, 10,000, or 100,000 components must work together in sync with microsecond or millisecond precision, failure of one component can cause the entire system to shut down. Recovery requires finding what happened, identifying which component stopped, locating the last checkpoint, and restarting. In the worst case, the system must start over. This is particularly problematic for reasoning tasks. If breakdowns or recovery processes hold the system back, all that work is wasted—it does not help reach the answer.
Accountability is measured by the actual output delivered for truly critical workloads in the data center, not theoretical productivity or theoretical quality. At the scale of 100,000 processors, something is going wrong all the time. Each chip is a wonder of nature at the edge of what can be manufactured technically. A package often consists of two, four, eight or possibly more chips assembled together, along with high-bandwidth memory (HBM), network connectivity, and possibly integrated optics. With 100,000 such units, failures must be detected and recovered from almost in real time. On a scale of 100,000 accelerators, a malfunction occurs several times a day, and perhaps several times an hour depending on exact configuration.
Common causes of malfunctions are difficult to identify because if there were a common cause, it would have been discovered and fixed. This is an ongoing field of discovery. When new products are introduced, new issues emerge. Problems may be network-related due to rapid connections between components, hardware-related, or software-related including compiler errors, runtime errors, model problems, or operating system problems. Even perfectly good and reliable hardware can encounter software problems that harm performance.
Nvidia provides a very strong reference package for integrated systems. Many customers benefit from the reference package, but many also customize it for their particular use case. Similarly, for TPUs there is a reference package, and while most people benefit from them, many also personalize them.
Google must roughly double its service capacity every 6 months. This refers to the actual available capacity for service at the end—the ability to generate tokens. Capacity is a combination of software and hardware. The requirement is not necessarily to double the number of floating-point operations every 6 months, but to double the ability of devices to generate tokens every 6 months. Much of this improvement comes from software improvements such as model enhancements and runtime optimizations. The rate of capacity improvement is amazing.
Most benefits in intelligence per watt come from model-side improvements. The software system contributes a considerable amount by ensuring devices are used effectively. Performance improvements of double or more can be achieved year after year from devices, which serve as a multiplier that raises the level of others year after year.
Google began building custom silicon more than a decade ago, starting the TPU program in 2013. This was a counterintuitive decision at a time when conventional wisdom held that the smartest approach was not to build an accelerator for a single workload. Moore's Law was doubling performance every 18 or 24 months, and standard programming models like C++, Java, and Python were available. The program began with the recognition that certain applications would take enormous advantage of and require an unimaginable amount of general-purpose CPUs to support them. It was a gamble that turned out to be very successful.
The first TPU was entirely dedicated to inference. The second slide demonstrated that the same idea could be applied to training. Around the time the second chip was released, transformers were invented, which completely changed the TPU program. Recommendation systems were found to work very well on these units. The scope expanded from two high-impact inference service applications—primarily machine translation and speech recognition—to training, then transformers, recommendation systems, and continuing to generalize as the generative AI moment took hold toward larger and more scalable systems.
The decision between one chip that does both inference and training versus two specialized chips led to the release of two chips: 8I for inference and 8T for training. By 2026, inference and service were expected to make a big leap and constitute 30, 40, 50 or 60% of the market during the chip's lifetime. Having a segment that is much faster for service made logical sense. If inference were expected to represent only 2% or 5% of the market, even if a specialized chip were twice as fast, it might not make sense because the generic chip would be chosen instead.
The more specialized the hardware is for a particular workload, the less flexible it becomes, but the faster and more energy-efficient it becomes. The workload must be sustainable enough to predict exactly what gains will be made by specializing in it. There is a large fixed cost to support a new program, so there must be sufficient demand to justify the cost. Both the 8I and 8T chips can do the other's workload. This is essential because predicting exactly how much of each type would be needed over the 6-year device lifespan would otherwise be required.
Transformers are essentially about vector and matrix multiplication with a softmax function—a set of linear algebra fundamentals that are almost integrated into the hardware. The model structure determines how many layers exist, how to move through the layers, and the exact form of matrices and vectors applied to each dimension. Further specialization beyond transformers to specific models represents the next level of specialization.
GPUs are more general-purpose than TPUs. Google sells a lot of GPU units and uses GPUs internally. The choice depends on the specifics of the problem. There is some overlap between the two, and customers assess their workload and evaluate their options.
Co-design from chips to network and software offers tremendous opportunities for improvement. Designing abstraction layers allows working across multiple clouds, devices, and software programs, making the system fully adaptable to any hardware, software, or network architecture. New capacity can be utilized overnight because the system is designed to take advantage of virtually anything.
When hardware is not hard-programmed or fully dedicated to any specific infrastructure, the downside is likely sacrificing significant performance. Complete interchangeability and adaptability to any hardware, software, network, storage, or computing capabilities across any provider creates a huge compatibility gap between layers. There may be 10%, 20%, or even double-digit improvements at each layer, and when those improvement opportunities are pursued, there is a huge end-to-end opportunity in terms of intelligence per watt or throughput per watt, including power delivery, software improvements, and other factors. The upside is working anywhere, anytime, without technical limitations, but the downside is giving up a significant, probably too significant, amount of performance.
One of the most enjoyable and rewarding aspects of working at Google is the opportunity to work alongside the DeepMind team in the joint design of hardware and models. This is a truly deep partnership. Examples include DeepMind researchers identifying needed model improvements such as specific transformers or calculations, and requesting hardware support from chips still in development. This leads to intensive multi-day meetings where engineers and researchers negotiate trade-offs: hardware changes may deliver 90% of what was requested, while model architecture adjustments can achieve 98% of the desired outcome. Manufacturing tape-out may be delayed by a week or two if the benefit justifies it.
Google maintains a five- or six-phase business plan spanning many years. This includes chips in production, chips being brought into production and debugged, chips in implementation about to finish design, chips in design, and chips in concept. DeepMind collaboration is critical across all phases to maximize intelligence delivered or throughput delivered per watt. The ability to intervene and make changes to chip architecture during development would be difficult or impossible across corporate boundaries.
For chips currently under design, multiple architectures can be evaluated together with DeepMind. DeepMind colleagues provide input on where model architectures are heading two or three years ahead. Deep simulation infrastructures predict how workloads will be compatible with different hardware architectures. The teams work in the same building, often in the same rooms, with deep daily interaction. Gemini is used to design devices for future Gemini models.
Hardware planning cycles extend two, three, four, or five years in advance. TPUv8 and TPUv9 through v10 are already progressing through stages from conception to implementation. Model architecture researchers understand that hardware operates on multi-year cycles. Small modifications that achieve 1% or 0.5% improvement on chips about to be manufactured are typically not pursued, as stopping chip manufacturing is of paramount importance. However, big and important opportunities trigger collaborative evaluation for execution.
The TPU architecture at a medium level of detail has not really changed since the first version. Comparing it to CPU instruction sets with operations like loading, storing, adding, subtracting, and branching, TPU units have basic instructions and essential prerequisites that have expanded over generations. The primitives include specialized arithmetic operations for very large matrix multiplication units, scatter kernels managing vector operations, aggregation and distribution operations, and five or six fundamental operations for remote memory loading and storage processes through the Interconnection Network (ICI).
The biggest workload change over the past year appears to be the rise of agents with long time horizons. This differs from fast-paced, large-scale language model conversations. Two main aspects emerge: first, there is no human loop naturally limiting request speed, so what used to take seconds or tens of seconds now takes fractions of a thousandth of a second between responses. Second, thinking and analysis will most likely occur on CPUs, which must coordinate context gathering from local DRAM, other CPUs' memory, SSDs, or HDDs. This dramatically increases demand for CPUs, networks, and storage alongside accelerated computing devices.
Placing CPU racks next to GPU and TPU racks means sacrificing full specialization in density and networking requirements. TPU racks are denser and require more networks than CPU racks, changing building design. The alternative is maintaining homogeneity by placing all TPUs or GPUs in one building, CPUs in an adjacent building, and hard drives in another building. This requires large-scale networks between buildings, increasing network complexity, reliability considerations, cost, and latency from hundreds of microseconds with queues between components.
Google was among the first to introduce wavelength division multiplexing 15 or 16 years ago, putting multiple signals on a single optical fiber for all inter-shelf communications. Optical circuit switching was introduced alongside this, transmitting data entirely within the optical spectrum without touching bits in the electrical range. Instead of electrical beam switching that examines packet headers and routes billions of packets per second, optical circuit switching determines output ports based on input ports. The initial implementation used MEMS switches with microelectric actuators controlling three-dimensional mirrors, allowing programmatic configuration of 128 or 256-port boxes to reconfigure network backbones without moving optical fibers.
TPU units use torus topology connecting all units directly. When a TPU rack breaks, another rack can replace it without moving fibers by redirecting light to spare lockers in fractions of a millisecond. Fiber optics are necessary because attenuation and bandwidth would decrease significantly without them, and routing everything across very large buildings in three dimensions would be challenging or impossible without fiber connections. Light reaches optical circuit switches on tiny chips.
Energy is the only fundamental constraint faced. Everything else can be solved over time, but energy is a long-term binding issue. Nuclear power may provide abundant clean energy, but timing and scale remain unknown. Data centers requiring gigawatts of power cannot simply request immediate delivery from utilities like PG&E. Energy planning requires years of advance coordination with utility providers, covering costs of building infrastructure including transmission lines and additional utility stations.
If gigawatts are needed in 2028 but facilities can only provide 700 megawatts that year and full capacity in 2029, the 300-megawatt gap must be addressed. Solutions include waiting, generating some energy locally through solar cells or batteries, or combinations where local generation supplies the grid during peak demand periods. The preferred model is working with utility companies over many years rather than pursuing vertical integration, leveraging statistical multiplexing and the law of large numbers for flexibility and mutual gains.
Data center sizing is described as an art form and source of great debate. Ten or 15 years ago, discussions centered on whether to put everything in one gigawatt-capacity data center, but single points of failure and power supply concerns make this impossible today. Optimal size depends on location, with some sites requiring tens of megawatts at network edges or in specific countries, while training complexes may approach gigawatts and other sites hundreds of megawatts. Models and simulators inform these decisions.
The demand for inference means that organizations cannot simply rely on training sets that are no longer fully used as a basis for inference processes. Training clusters may be concentrated in a given year in a few large locations, with benefits to keeping network distance between them small, meaning they are likely located on the same continent or even the same part of the continent. This creates service capacity constraints in other continents, requiring specialized inference facilities to be built in other places around the world.
Service groups differ from training groups in several key ways. While service groups are smaller in size, they are not necessarily cheaper per megawatt. For service workloads, there is a greater need to put computing, networks, and storage in one place, leading to an inability to specialize. Training involves high density and uniform deployment, whereas service requires a combination of storage, computing, and accelerators. For service, the goal is not to put too much in one place, but to provide services for workloads from all over the planet.
Individual model endpoints may have model variations, requiring models to be distributed worldwide while taking location into account. This results in smaller, less vertically integrated facilities from the inference side.
Data centers built 5 years ago contain accelerators that were completely different and much less efficient than those manufactured today. The question arises whether to replace chips in old data centers, relating to ongoing debates about chip lifespan.
TPUs that have been operational for 7 years are still fully operational with high usage rates. For depreciation purposes, the lifespan is approximately six years. Once value is fully consumed, and considering energy efficiency of newer generations, it makes sense to replace and upgrade entire systems.
Replacement is considered from the perspective of pods. An 8T set of TPUs might consist of 9600 chips and contain approximately 140 to 152 shelves. Entire sets are replaced with newer generation TPUs, though the space occupied by new groups may not exactly match the space left by old groups. Adjustments must be made, and planning for future generations is impossible since future TPU designs are unknown.
While supporting vertical integration for optimal performance, it is important to avoid resource entanglement or closed systems. Google developed the JAX model development framework, used extensively internally, though many customers prefer PyTorch. Rather than requiring JAX usage for TPU access, PyTorch support is provided through TPU torch, allowing unmodified models to work.
The Internet Protocol succeeded because it was an open standard serving as the narrow waist of an hourglass, where anything could work above it programmatically and below it hardware-wise. This open and interoperable standard allowed anyone using an IP address to connect to router ports and become part of the internet, enabling enormous growth. Open standards should be supported at connection points and preferably be open source, as large-scale systems cannot rely on closed standards specific to one vendor.
On the software engineering side, AI is used effectively for software development capabilities, testing, launches, and design assistance. On the hardware side, hardware engineers use AI as much as software engineers do, measured by token usage. Productivity has increased, and the time from design phase to tape-out has begun to shrink, along with initial setup time.
The way data centers are designed has changed significantly. Previously, assessments of whether to build gigawatt facilities or 100-200 megawatt complexes relied heavily on spreadsheets and human effort. Now there is significant AI involvement that simplifies planning and development processes, though these are still inferential models that do not replace human judgment but make it easier to gather necessary information.
Google is pursuing orbital data centers as one of their ambitious moonshot projects. The basic constraints of energy and energy production are the main challenge on Earth. In space, there is approximately 40% more energy due to absence of atmospheric attenuation, providing 1.4 times more solar capacity. In sun-synchronous orbit, there is 98-100% sunlight coverage compared to 28-35% on Earth.
This represents a factor of 1.4 times multiplied by 3-4 times in terms of hours per day, essentially excluding batteries from the equation while providing carbon-free energy. However, many challenges exist including cooling, which is actually more difficult in space than on Earth, and reliability issues where repairs become more difficult. Network connectivity between clusters presents additional challenges.
The model may be promising with iterativity as a friend. Free space optics will become reality, with no stretching of fibers between components. Lasers will be pointed toward receivers with real-time calibration, and there are no fundamental obstacles that are impossible.
Ten years from now, the level of integration will be enormous. There will be greater integration with fewer fibers, though not necessarily no fibers. From a server perspective, greater integration is expected with perhaps a small fiber packet coming out of the server. Servers will likely be centrally manufactured with 72, 144, 288 or more GPUs or TPUs deeply integrated within the server.
Single server consumption could reach several megawatts. This saves water, energy, and fiber, with servers simply pushed in, three connections made, and immediate operation possible. By 2036, it may be possible to launch servers directly into space to be picked up by robotic arms at space stations and connected to appropriate units.
Keep Sequoia Capital in your library
Save the videos and channels worth coming back to, and find them again in one place.





