The largest capital expenditure buildout in human history is now underway. From the "goodput" metric that measures real system performance, to optical circuit switching networks, orbital data centers, and the shape of compute a decade from now, Google's AI infrastructure lead Amin Vahdat laid out the technical logic and strategic trade-offs behind this buildout in a recent in-depth conversation.
Google's capital expenditure this year is expected to exceed $200 billion, with most of it going toward data center construction. Speaking with Sequoia Capital partner Sonia Huang, Vahdat said the rise of long-horizon agents is fundamentally changing data center design logic, not only driving up demand for accelerators but also causing demand for CPUs, networking, and storage to surge in tandem. He also revealed that Google must roughly double its serving-side token generation capacity every six months, and that the primary source of this growth is not hardware itself but the continuous compounding of software and model optimizations.
For investors, Vahdat's judgments carry direct reference value: power is the most fundamental bottleneck to AI infrastructure expansion, not chips or manufacturing capacity; software optimization likely contributes no less to capacity gains than hardware; and on the TPU roadmap, inference and training chips were split into separate product lines for the first time this year, reflecting the rapid expansion of the inference market.
Not FLOPS but Goodput: Accountability Under Real Failure Conditions
Vahdat characterized FLOPS as a vanity metric, arguing it only reflects the theoretical ceiling of a single chip, while actual workload performance depends on how thousands or even tens of thousands of accelerators, CPUs, networks, and storage work together.
"At the scale of 100,000 accelerators, something is always breaking, always," he said. Once any component in a synchronous workload fails, the entire system may need to roll back to the last checkpoint and recompute, and this wasted effort is exactly the gap between goodput and throughput. He likened it to solving a problem on paper: every time you return to step one because of an error, the process consumes compute resources but does not count as effective delivery.
Failure causes follow a classic long-tail distribution. Vahdat admitted, "If there were a single most common failure cause, we would have found and fixed it long ago." Problems range from network interconnects and hardware itself to compiler bugs, runtime errors, and even operating system issues, with each new product generation introducing new failure modes.
The Agent Era Reshapes Data Center Architecture: CPU and Storage Demand Surge Together
Vahdat sees the rise of long-horizon agents as the biggest structural variable affecting data center design over the past year.
In the traditional human-machine interaction model, users read the model's response, think, and input again, with a "human cadence" of seconds or even tens of seconds in between, which naturally limits request frequency. In agent scenarios, the model does not need to wait for a human, response latency can compress from seconds to milliseconds, and request density rises by orders of magnitude.
At the same time, agents rely heavily on CPUs for inference and orchestration: parsing the previous response, deciding the next action, and fetching context from local DRAM, remote memory, SSDs, or HDDs. These operations run on CPUs rather than accelerators. "Accelerated computing demand is rising, but demand for traditional data center components like CPUs, networking, and storage is also surging," Vahdat said.
This directly creates a data center layout dilemma: placing large numbers of CPU racks next to GPU/TPU racks would break the specialized design optimized for accelerators; placing the two types of racks in separate buildings would introduce hundreds of microseconds of queuing latency through inter-building networking and significantly raise reliability and cost pressures.
TPU Splits Inference and Training for the First Time: Where Are the Limits of Chip Specialization
Google's eighth-generation TPU released this year was split into two separate chips for the first time: 8i for inference and 8t for training. Vahdat framed this decision as a direct result of forecasting the scale of the inference market.
Under the previous single-chip strategy, one chip had to handle both inference and training, and neither workload could be optimized to the extreme. As inference's share of total compute demand is expected to rise to 50% or higher, the marginal gains from specialization began to outweigh the flexibility lost through dedicated design.
However, Vahdat emphasized that this split included a key safety valve: both 8i and 8t retain the ability to execute the other's workload, with only some performance loss in the non-specialized direction. "If 8i could not do training at all, we would have to precisely predict the demand for each of the two workload types across a six-year lifecycle in advance," he explained, which would create unbearable forecasting risk.
On the broader philosophy of specialization, Vahdat described a continuum from general-purpose GPUs to hardening Transformer primitives, and even "burning a specific model architecture into silicon," noting that some companies have already begun exploring the final step. Google's current judgment is that linear algebra primitives such as matrix multiplication and softmax that Transformers rely on are already largely hardened in hardware, and further specialization down to the specific model level is a "very, very interesting direction" but has not yet been implemented.
Optical Circuit Switching: From MEMS Mirrors to Millisecond Failover
About 15 years ago, Google was among the first to bring wavelength division multiplexing (WDM) into data centers, and later layered on optical circuit switching, which routes data entirely in the optical domain, bypassing the electrical-domain processing of traditional electrical packet switching that looks up forwarding tables packet by packet.
Vahdat explained that Google initially used MEMS microelectromechanical switches, precisely controlling the angles of tiny mirrors in three-dimensional space to direct optical signals from any input fiber port to a designated output port, enabling programmable optical-domain routing. This technology initially served two goals: creating optical-domain shortcuts between compute clusters and storage clusters that communicate frequently, and dynamically reconfiguring network topology through a software controller for scale-up and scale-out without physically moving any fiber.
In TPU clusters, optical circuit switching also provides critical fault tolerance: when a TPU rack fails, the system can redirect the optical path to a spare rack within milliseconds, without any physical operation, directly reducing goodput loss.
Vahdat revealed that in orbital data center scenarios, communication between components will genuinely shift to free-space laser communication, which is precisely why he avoids that approach in terrestrial settings, where attenuation loss and the difficulty of large-scale three-dimensional alignment are too great, but in the space environment those limitations are greatly weakened.
Power Is the Fundamental Bottleneck: Multi-Year Planning With Utilities, Gigawatt-Scale Demand Reshapes Power Supply Models
Asked about the biggest constraint on AI infrastructure expansion, Vahdat's answer was direct and clear: "Power is the most fundamental constraint we face. Everything else we seem to know how to solve over time; power is the long-term binding problem."
Google's preferred model is grid interconnection with utilities rather than building independent power sources. The reason is the statistical multiplexing effect: full self-supply at reliability above 99.99% would mean building twice the installed capacity, whereas long-term coordination with utilities smooths demand fluctuations over a much larger base, allowing both sides to reduce redundancy costs. Google commits to bearing the cost of new transmission lines and substations required for its interconnection, avoiding passing costs on to other users.
But supply-demand mismatches are real. Vahdat gave an example: if 1 gigawatt is needed in a given year but the utility can only provide 700 megawatts that year, Google might temporarily build solar plus storage to fill the gap and feed local generation back to the grid during the hottest months.
On the optimal size of a single data center, he acknowledged this is "a source of enormous debate" and highly site-dependent: training clusters favor gigawatt-scale concentrated deployment to minimize network latency, while serving clusters need to be distributed globally close to users, are smaller individually, and because they require a mix of storage and compute, their degree of specialization is correspondingly lower.
DeepMind Co-Design: Intercepting Architecture Mid-Flight
Vahdat described the collaboration with DeepMind as a "deep partnership" and used the phrase "making changes mid-flight to chip architecture" to capture its closeness.
Google is simultaneously advancing five to six generations of TPUs: in mass production, in bring-up, about to tape out, in design, and in concept, forming a parallel pipeline spanning several years. Different stages correspond to different depths of collaboration windows: for chips already in production, the focus is jointly maximizing goodput; for chips about to tape out, if DeepMind proposes a model optimization with major benefits, both sides can urgently assess whether it is worth delaying tape-out to incorporate hardware changes. "Stopping a tape-out is a big deal, but if the opportunity is big enough, we will absolutely work together on it." For chips in the design phase, the two sides jointly forecast model architecture trends over a three-to-four-year horizon, relying on simulation infrastructure to evaluate how different hardware architectures fit future workloads.
Vahdat also disclosed that Google is already using Gemini to design hardware for future versions of Gemini, forming a self-reinforcing loop in which models assist chip design. He and DeepMind's Koray and Demis "talk multiple times a week."
The 2036 Supercomputer: Multi-Megawatt Racks, or Launched Directly Into Orbit
For a prediction of what compute will look like a decade from now, Vahdat offered two parallel paths.
On the terrestrial path, he expects integration to rise dramatically, with racks evolving from the externally visible tangle of fiber bundles into highly integrated units where "only a small bundle of fiber comes out of the rack." A single rack could reach multi-megawatt power levels, and on-site installation would require only three connections, power, water, and fiber, to operate, with manufacturing and deployment trending toward modular factory prefabrication.
On the orbital path, Google has made orbital data centers a formal moonshot project. The core logic lies in energy advantages: the absence of atmospheric attenuation in space naturally yields about 40% more usable power; a sun-synchronous orbit can achieve 98% to 100% sunlight coverage, compared with typically only 28% to 35% for terrestrial facilities, compounding to an energy density advantage of three to four times, with no carbon emissions.
The main challenges are cooling, which is actually harder in space, reliability and repair, since fixing broken hardware is far harder than on the ground, and inter-component communication, which will be forced to use free-space lasers rather than fiber. Vahdat said these issues have "no fundamental showstoppers," and joked that by 2036 it might really be possible to "launch a rack into space, have a space station robotic arm catch it, and plug it into the right module."
The Full Interview Transcript
Host (Sonia Huang, Sequoia Capital): Welcome Amin Vahdat to the show. Thank you very much for joining us today. I am really looking forward to this conversation. Today's topic excites me because we are at the center of a historic large-scale capital expenditure buildout.
Amin Vahdat (Google AI Infrastructure Lead / CTO): I am excited to be here too. Today's topic is very exciting.
Host: You are at the center of the largest capital expenditure buildout in human history. Google alone is expected to spend more than $200 billion this year, most of it on data centers. And you are the central figure in all of this: at the end of last year you were appointed head of Google AI Infrastructure, leading one of the most capital-intensive buildouts in human history. So I am really looking forward to digging into this with you today.
Amin: This has indeed been a very significant year at the industry level and for Google. Honestly, I think we have never seen anything like it before; certainly not at Google, and I even think not in human history, whether in terms of the scale of the buildout or the speed of transformation. It is incredible.
Host: Before we go deeper, let me quickly bring the audience up to speed: What is an AI data center? How is it different from a non-AI data center?
What Makes a Data Center an AI Data Center
Amin: That's a great question. AI data centers and non-AI data centers actually have a lot in common; they are not entirely different. They are both made of concrete, with an enclosed space; they both have electrical yards, mechanical yards, cooling systems; rows of power distribution; and a lot of networking infrastructure, meaning connecting large amounts of computing equipment to each other; data centers also have quite a lot of storage infrastructure.
I think the biggest difference for AI infrastructure is specialization. In the past, building a data center was essentially making a 20-, 25-, 30-year building investment. We would think about how the building would evolve over 20, 25, 30 years. It might house servers, networking, storage, and maybe some accelerators such as GPUs, TPUs, or other chips; but the planning horizon was 25 to 30 years. Hardware life might be six years, so we had to plan for many generations of hardware.
AI data centers, by contrast, are often more purpose-built. In other words, we often co-design the building and the hardware that will go into it. For example, we might decide not to put too much storage in a given building. Why? Because a storage rack might draw 10, 20, 30, 40 kilowatts; putting it next to a TPU rack or GPU rack, where racks today easily reach hundreds of kilowatts and may go even higher in the next few years, is a very different design. Designing a building to hold 30 storage racks in a row versus only one or two AI racks is very different. You can understand that difference from the standpoint of size, power, and how power is distributed within the building.
The same is true for networking. Storage racks need very little network bandwidth, especially mechanical hard drive storage, compared with AI racks. If you want a data center to be fully interchangeable and fully general-purpose over a 30-year cycle, you will likely build it too large and overdesigned. AI data centers are likely more specialized, co-designed with hardware, even down to cooling and power distribution. That is, there will be more joint optimization.
Host: You delivered a large Vera Rubin cluster to one of my portfolio companies, Ineffable Intelligence, and I saw the photos when it shipped. It is truly a masterpiece. It is almost a monument showing what kind of engineering humans can produce.
Amin: Indeed. It was only a few racks, but the cabling between racks and the fiber distribution are really beautiful. Beauty is in the eye of the beholder, but for me, for you, and for many listeners, it really is engineering beauty. We put that photo on social media, and people really liked what Ineffable Intelligence is doing; that is an outstanding team and it got a lot of attention. But actually, what people liked most was the fiber photo, especially the fractal-like structure of the fiber. That post became one of our most popular posts ever. So we were very excited.
Host: I got goosebumps when I saw the photo.
Not FLOPS but Goodput: How to Be Accountable When the System Fails Every Hour
Host: In this massive buildout, how do you measure yourselves and hold yourselves accountable? We talked before the show that FLOPS is more of a vanity metric and that you prefer another metric. Can you expand on that?
Amin: Whether it is FLOPS or another chip-centric metric you prefer, it is fundamentally theoretical. That is, under certain conditions, for a given chip, this is the maximum FLOPS it can deliver. But what we ultimately care about is the performance actually delivered by a workload. Workload performance is rarely determined by a single chip. It may depend on FLOPS, HBM, SRAM capacity, and so on, and those factors are very important; but it may also depend on how 2, 4, 8, 16, 1,000, 10,000 or more chips are put together. That includes not only accelerators, whether TPUs or GPUs, but also the CPUs feeding them data and the network connecting everything.
So the question is: what workload are you running? What is the performance of that workload? A meaningful measure is FLOPS utilization: for a specific workload, if you theoretically have some teraflop or petaflop capability, what fraction of it did you actually deliver?
That is one measure of goodput. What is goodput? Everyone is familiar with throughput. That is the theoretically possible throughput. But now you have to consider other factors. One is the inherent slowdown of the workload itself; another very critical factor is reliability.
In terms of accountability, if one chip fails in a synchronous workload, and these workloads, whether training, serving, or agentic workloads, are often synchronous, meaning many components work together at the same time. Suppose 1,000, 10,000, 100,000 components are working simultaneously and need to coordinate at microsecond or millisecond granularity. If just one fails, it can stop the entire system, because each component depends on the others completing their part of the work in order to jointly answer a very hard question.
One component stops, and now we have to figure out what happened, which one stopped. What checkpoint did we save at some earlier point? How do we restore that checkpoint? How do we restart? In the worst case, we may have to start over, which is very bad. In some cases that can indeed happen, especially more commonly on the inference side. The point is: if after a failure you have to go back and redo a large amount of computation, if you have to pause and wait to troubleshoot, none of that work truly helps you get the answer.
It is like solving a problem on paper: step one, step two, step three, step four. If you have to go back to step one because of a mistake, of course you are still doing work, that is throughput; but what really matters is goodput, which is the total time it takes you to deliver the answer. If there are failures, failure recovery, or anything that interrupts the work, those all count in the problem.
So the way we hold ourselves accountable is by the goodput delivered, not the theoretical benchmark throughput, and not theoretical goodness. For a real workload, under real failure conditions, what actually happened? Unfortunately, at the scale of 100,000 accelerators, I have to say: at that scale, something is always breaking. Always. And each of those chips is a miracle of nature, at the very frontier of manufacturing capability. These chips are not just one chip now, but a package of two, four, eight or more chiplets, plus HBM alongside, network connections, and perhaps co-packaged optics. There is nothing to blame, but there are indeed many things that can fail. And once you have 100,000 such devices, something will always fail. You have to be prepared for it: detect it in near real time and recover in near real time. The telemetry problem is enormous, like continuously finding needles in haystacks; of course it is continuously online at the second and minute level, and for some tasks, at the hour, day, and week level as well.
The standard by which we hold ourselves accountable is how much goodput the workloads that truly matter in the data center deliver.
Host: Is goodput a Google internal term, or an industry term?
Amin: It is a Google term, but I see more and more people in the industry adopting it.
Host: Help me calibrate: at the scale of 100,000 accelerators, how often do failures happen? Once a day? Once an hour?
Amin: At that scale, definitely multiple times a day, depending on the configuration, and possibly multiple times an hour. Something is always failing.
Host: What is the most common cause of failure?
Amin: That is the thing. It really is a great question. If there were a single most common cause of failure, we would have found and fixed it long ago. The reality is a long tail of constantly discovering new problems. Whenever a new product is introduced, new problems hit us. Frankly, in many cases precisely because it is at the frontier, failures may come from networking, from how we connect these devices at extremely high speeds, or from the hardware itself. We solve these one by one. But many issues can also be software problems. That is another part of how we hold ourselves accountable. A chip may have some FLOPS capability, but if there is a compiler bug, runtime bug, model issue, or operating system issue, that does not matter, it still hurts system performance. You may have completely reliable, perfect hardware and still be dragged down by software problems.
Doubling Token Capacity Every Six Months, and Where the Gains Actually Come From
Host: From accelerator companies, Nvidia or the TPU team, is there a standard reference stack such that if you build that optimal system around their accelerators, you are fine? Or do you still have to do a lot of data center design beyond what the semiconductor company provides?
Amin: There is indeed a reference stack. I would say Nvidia is an excellent full-system company; it is obviously a semiconductor company, but not only a semiconductor company. They provide a very strong reference stack. But what we observe is that most customers use that reference stack, while many customers also specialize. That is, they find that for their specific use case there is a naturally emerging optimization opportunity, or there are certain different requirements they must handle, and so they do it. It is similar on the TPU side. We also have a reference stack that most people use, but many also specialize further.
Host: Understood. On overall capacity again: you told your team that Google has to double serving capacity roughly every six months, right?
Amin: Let me clarify. What is meant here is effectively available capacity. Ultimately we look at token generation capability from the serving perspective. Capacity is a combination of software and hardware. Hardware may have some so-called intrinsic FLOPS level. I am not saying you have to double the number of FLOPS every six months; that is only one path. You do have to double the hardware's ability to generate tokens every six months. And the part that comes from software is likely no less than the part that comes from hardware. That is, the gains may come from model optimization or from runtime optimization. In fact, it is more likely dozens or hundreds of individual optimizations landing continuously, stacking again and again, that make all of this possible. But the pace of capacity improvement is truly incredible.
Host: Over the past few years, we have seen several years of data center construction, model progress, and software progress. Measuring capacity growth by intelligence per watt, empirically break it down: how much comes from silicon itself, how much from models, how much from other software, and what are the other big components?
Amin: Good question. I do not have an exact breakdown, but in our experience, most of the gains come from model-side improvements, that is, intelligence per watt at the model software system layer. By the way, intelligence per watt is a very good metric. We more accurately say goodput per watt. We can talk later about why the denominator has to be watts. Goodput can essentially be a measure of delivered intelligence per watt; goodput itself is a workload-specific metric. But I think most of the gains often come from model-side software systems, and then from other parts of the software system. Why? Because they ensure that the hardware you already have is used effectively.
Hardware itself is also amazing. In other words, we are living in an era where annualized 2x or even higher performance improvement is entirely possible. Hardware can indeed support that. If you want an analogy, it is like a free multiplier, not completely free, but everything above it can count on it: a multiplier that lifts everyone year after year.
The TPU Bet: From the Contrarian 2013 Judgment to the First Split Into 8i and 8t
Host: I want to talk about the TPU project and co-design. Google started building custom chips more than a decade ago, which was a contrarian judgment at the time. How has the TPU project evolved?
Amin: It has changed quite a lot. The project began in 2013, and it was indeed very contrarian at the time. It is hard to put yourself back into the context of 2013: the conventional wisdom then was that all the smartest, most experienced people would say you should not build a custom accelerator for a single workload. Why? Because Moore's Law was still alive, performance doubling every 18 or 24 months was still happening; you could use standard programming models, your C++, Java, Python, and so on all worked. It was a bit like the bitter lesson in chips: specialization will never win.
But in this particular case, we had one application, or a few applications, that would benefit enormously from specialization, and supporting them would require an unimaginable amount of general-purpose CPUs. So in 2013 it really was a bet, and even inside the company quite a few people were unsure it would succeed. It turned out to be a hugely successful bet.
The first-generation chip was entirely focused on inference. The second-generation chip was about the idea that we could actually extend the same approach into a training chip. From then on, more and more use cases began using TPUs. Around the time of about the second-generation chip, Transformer was invented, which was a major moment and completely changed the direction of the TPU project. We also found that recommendation systems ran very well on TPUs, so advertising and related use cases joined as well. So this change and evolution is essentially an expansion of scope and impact: it began with one or two highly impactful inference service use cases, mainly language translation and speech recognition; then expanded to training, then Transformer, then recommendation systems; and as the generative AI moment arrived, it continued to generalize into ever larger and more scalable systems.
Host: You mentioned the emergence of Transformer as an important moment. I am curious, TPU is a specific architecture, but it is not specifically for Transformer. How do you strike the balance on that fine line of how specialized a chip should be for a workload?
Amin: That is a great question, and the core is chip applicability to a specific workload. Every generation we think about whether to specialize further. For example, a little more than two years ago, the question we faced was: by 2026, should we make two chips or one? We could make one chip that does both inference and training reasonably well; or we could make two chips: one further specialized for inference and another further specialized for training. That analysis and work ultimately led to the two chips released this year: 8i for inference and 8t for training.
We ultimately realized that while that may not have been true a few years ago, by 2026 we saw inference and serving truly taking off. So a chip that is clearly faster at serving began to make a lot of sense. We had estimated that inference might account for 30%, 40%, 50%, 60% of the market over its lifecycle. By contrast, if we expected inference to be only 2% or 5% of the market, even a 2x specialized chip might not make sense; you would instead favor a general-purpose chip that is not fully optimized but has advantages such as uniformity.
So this really is a calculation: do you specialize further? How big is this workload? How big is it expected to be in the next two, three, or four years? Can this workload support continued growth at the hardware level? As I said earlier, the opportunity is that the more you specialize for a workload, the less flexibility you have, but the faster and more power-efficient the hardware becomes. So this is both an art and a prediction of design targets: how durable is this workload? If it disappears after a month, two months, or three months, even if those three months are large, your interception window is very narrow. It has to be durable enough, and you have to be able to accurately predict how much benefit specialization will bring.
Host: The rough trade-off seems to be: supporting a new project has a very high fixed cost.
Amin: Yes. You have to believe there will be enough demand for that specific project to justify the cost. For example, for 8i and 8t, it is also critical that both can actually do the other workload. If 8i could only do inference and could not do training at all, meaning training performance of zero; conversely, if 8t were extremely strong at training but had zero inference performance, that would also be a critical limitation. Why? Because we would have to predict in advance how much of each we needed across the chip's roughly six-year lifecycle. As it stands, both chips are better at their specialized workload, but if there is spare capacity somewhere, both can also take on the other's work. Depending on the degree of specialization, you could make yourself too specialized and lose flexibility and interchangeability.
Host: Since one end of the rough trade-off is project scale, let me ask a provocative question: so much of the modern AI market is based on Transformer, so why not just burn the Transformer architecture into the chip?
Amin: Transformer is indeed foundational, but the next-layer question is: Transformer is essentially about vector and matrix multiplication operations, plus softmax. That is, a series of linear algebra primitives. We and other vendors have largely hardened these primitives into hardware. Everyone is of course based on Transformer, but there is another layer of model architecture: how many layers do you have? How do you advance within a layer? How do you advance across layers? What shapes of matrices and vectors do you apply in each dimension? You can specialize further, not just to Transformer, but to your model. That is the next layer of specialization. I know some companies are thinking in this direction, and I also think it is a very, very interesting direction.
Host: So now, do your customers view TPU and GPU as roughly interchangeable equivalents? Or are some problems better suited to one or the other?
Amin: There are definitely some problems better suited to one or the other. GPUs are more general-purpose than TPUs; that is clear. Google and Google Cloud have very strong offerings, we sell a lot of GPUs, and we also use GPUs internally. But what really matters is the specifics of your problem. Our customers evaluate their workloads and options in the overlapping range of the two. What we at Google like to do is give customers choice: provide the right solution for their needs, and provide as much as possible the option that best meets their needs.
The Case For and Against Co-Design
Host: From chips to networking to software, what is the case for co-design? And what is the case against?
Amin: Co-design offers enormous optimization opportunities. You can imagine that if you have an end-to-end stack and want to freely choose any component, for example if you run across multiple clouds, or across many hardware and many software environments, you can design abstraction layers and declare: I can run on any hardware, any software, any network topology; no matter how much network you give me, my system will fully adapt. Then you gain a strong capability: if new capacity appears overnight, you can run on it, because the system was designed this way and can immediately use any resource. You have no hard-coding and no specialization for any particular infrastructure.
The downside is that if you are completely general-purpose, you will likely leave a lot of performance on the table. If you are fully flexible across anyone's hardware, software, network, storage, compute, and so on, then large impedance mismatches will appear between every layer if you pursue complete generality. There may be a gap of 10%, 20%, or even 2x between each layer. You start multiplying these optimization opportunities level by level, and in the end a large end-to-end opportunity remains, whether you call it intelligence per watt or goodput per watt, that can be exploited, even all the way down to power distribution, power availability, software optimization, and so on.
The case for co-design is: you can run anywhere, anytime, with no lock-in. The case against is: you will leave a significant, and quite possibly very significant, amount of performance on the table.
Host: My understanding is that if you look at OpenAI and Anthropic: OpenAI is mainly built on a relatively homogeneous compute stack, while Anthropic is built on a relatively more heterogeneous compute stack. Does that partly explain co-design and become one of the reasons people say the two companies have very different model architectures?
Amin: I cannot and do not want to speculate on what OpenAI and Anthropic are doing. It is one possibility. But without knowing exactly what they do, I would imagine there could be many reasons they ended up with different architectures.
Fighting Side by Side With DeepMind: Intercepting Architecture Mid-Flight on the Chip
Host: What about Google? I am curious what your team's working relationship with DeepMind is like: at what stage of model development does everyone get into the same room? How do you jointly make decisions on co-design?
Amin: This is one of the most interesting and most rewarding parts of working at Google: the opportunity to truly work shoulder to shoulder with the DeepMind team in co-designing hardware and models. There is also a third element: we also extend this collaboration to consumer services and the cloud. Let's set that aside for now and come back to it later. As far as DeepMind is concerned, it really is a deep partnership.
Take a past example: they propose a model optimization, perhaps for Transformer or for a specific mathematical operation, and at that time we have a chip in progress that is not yet finished. They might say: "Wow, if the hardware supported this, our end-to-end training or serving could be significantly faster and more efficient. What would it take to change the hardware definition we are already executing?" That pushes engineers and researchers from both sides intensively into the same room for a few days, a week, two weeks, discussing: "Okay, we can do it." More often, "We cannot fully do what you want, but we can do another version." And then maybe you can also change the model architecture a bit in another direction, so you get 98% of what you wanted, 90% of what we targeted, and then we go back and adjust the hardware. Or: "You know what, we will delay tape-out by a week or two, because this level of benefit is completely worth it."
Similarly, when we plan the roadmap, there are many generations of chips advancing in parallel at any point in time: one generation is already in production; one generation is being pushed into production, has come back from the manufacturer, and we are debugging and making it work; one generation is in the implementation stage, about to tape out and be handed to the manufacturer; one generation is in the design stage; and one generation is in the concept stage. So from mass production to still in our heads, it really is a pipeline of five or six stages spanning many years.
Collaboration with DeepMind is very important for chips already in production, because together we can maximize delivered intelligence or delivered goodput per watt. We know exactly what is happening in the model, what is happening in the hardware, and everything in between. We can also work deeply on chips about to tape out. Why? Because we can intercept and literally make changes mid-flight to the chip architecture. That would be difficult, even impossible, across company boundaries; not impossible, but much harder. If there are only a few weeks or months left before the chip is finished, then saying "Oh my God, let's sit in the same room shoulder to shoulder and see whether we should disrupt the project" is possible, but harder. That is my judgment about crossing company boundaries.
For chips still in the design phase, we can evaluate many architectures together. We will ask DeepMind colleagues: where do you think model architectures are heading in two or three years? Here is one set of options we can do, and here is one set of directions in which model architecture may evolve. In fact, we have deep and substantial simulation infrastructure that can predict how workloads will map onto different hardware architectures. Deep and rapid iteration. The teams are not working separately; much of the time they are in the same building, the same room, interacting deeply every day. I talk with Koray, Demis, and others multiple times a week. So this is truly a very interesting side of the work.
Host: That's great. Ultimately their models will also help chip design.
Amin: We are also using Gemini to design hardware for future versions of Gemini.
Host: That is really cool. I want to come back later to the collaboration itself. I guess there is still an impedance mismatch here: the hardware cycle is obviously slower than the cycle DeepMind faces on the software side, and your planning horizon is certainly further out. So how much room for adjustment do you actually have?
Amin: Good question. There is no doubt that we plan hardware two, three, four, five years in advance. We just talked about the eighth generation of TPU that we have announced, but you can imagine the ninth, tenth, and maybe others already in concept, execution, or other stages. They may still be years away. And if you are doing model architecture work by default, you usually do not think several years ahead.
But that is exactly the great thing about growing up together as a company. Google Research, DeepMind, the invention of Transformer, all of that also happened in these overlapping rooms. That is, a whole generation of researchers has become used to being able to influence hardware, and also knows that hardware moves on multi-year cycles. They also know that if it is just a small change that would bring 1% or 0.5% benefit to a chip about to tape out, they probably will not come to us, because they understand well enough: this is not like software, where you can submit a change list and be in production two weeks later. In fact, stopping a tape-out is a big deal. But they also know that if there is truly a very big opportunity, we will absolutely work together to see whether we can get it in. So the sources of gain are multiplicative. Hardware lifts all boats, right?
A substantial part of the DeepMind team, and we are very grateful for this, it is an amazing team, thinks carefully about how to influence the roadmap, because it is a pipeline. The idea you had two years ago is now in production and helping all Google workloads run faster, and that is a good feeling.
Host: I completely agree. Still, predicting the most common workloads five years from now, and what algorithmic breakthroughs will appear, seems almost impossible.
Amin: It does sound impossible. I understand why you would think so. But what is striking is that we are actually writing about this in great detail, and it is fun to write, the TPU architecture has not really changed since TPU v1 at a medium level of detail. Not that it has not changed at the highest level, but at the medium level of detail. You can think of it as analogous to a CPU instruction set architecture: you have load, store, add, subtract, branch. What has software done on top of that over all these years? The same is true for TPU. We have some basic instructions and basic primitives. Of course we have extended it, not that the instruction set is completely unchanged; but those primitives, such as specialized numerical formats, very large matrix multiplication units, a sparse core that manages vector operations and scatter/gather operations, and so on, what really defines it is five or six things. There is also one that is effectively load remote and store remote: we can read and write remote memory through the ICI network. These fundamentals have remained and have extended across many generations of models and many generations of even deep neural network algorithms and model structures.
Long-Horizon Agents Change the Shape of Data Centers
Host: Speaking of workload changes, one of the biggest changes over the past year or so, and I think actually from the beginning of this calendar year, is the rise of long-horizon agents. I guess this workload is shaped very differently from the fast-turnaround LLM dialogue of the past few years. What does it mean for data center demand?
Amin: I think there are two huge aspects. First, it is no longer a human-to-accelerator interaction model. In other words, when you type a prompt into a web browser or on your phone, of course a lot of computation happens in response to your prompt; but after the response comes back, you have to read it, think about it, and then perhaps have a follow-up. That is usually a multi-second interaction time. Now, in long-horizon agents, there is no human in the loop, so there is naturally no limit on how fast the model receives requests. Interaction times that used to be measured in seconds, perhaps tens of seconds, can now drop into milliseconds: as soon as I get a response, I can parse it, do a little reasoning about it, and determine what the next request is.
The second big change is that all of this inference and parsing likely happens on CPUs. And that CPU then has to think: what other state do I need to gather in order to generate the next prompt? In other words, I learned something from this response, I want to take the next step; but I actually need to fetch some context, perhaps from local DRAM, perhaps from someone else's DRAM on another CPU, perhaps on SSD, or perhaps on HDD elsewhere. So now a lot of orchestration is also required.
So the design is changing quite significantly: accelerated computing demand is rising, but demand for traditional data center components such as CPUs, networking, and storage is also surging.
Host: Does that mean you have to put more CPUs next to GPU racks? Are they GPU/TPU racks?
Amin: This is the key question. Back to the question of optimization and specialization: if we start putting a lot of CPU racks next to GPUs and TPUs, and that is indeed a problem, then it means we cannot fully specialize according to the density and network requirements of TPU racks relative to CPU racks. TPU racks will be denser than CPU racks and may also need more networking than CPU racks. In other words, our building design will change.
The other option is to maintain uniformity: put all TPUs or GPUs in one building, CPUs in the building next door; and perhaps by the way, hard drives must be in yet another building, because hard drives have another set of requirements. Now you need a fairly substantial level of networking between these buildings. Once you leave a building, network complexity rises significantly from the standpoints of reliability and cost; latency may still be acceptable, but there may now be hundreds of microseconds, or even more queuing latency, between components. So the considerations change in very interesting ways.
Host: Your job is hard.
Amin: Hard, but also very fun.
Optical Circuit Switching and the State of Networking
Host: What is happening on the networking side? I hear Google has long been at the forefront of the latest networking technologies, including optical networking. Can you talk about the state of optical networking?
Amin: About 15 or 16 years ago, Google was one of the earliest vendors to bring wavelength division multiplexing, WDM, into the data center. We can put multiple signals into a single fiber and use it for all communication between racks. At the same time, we also introduced a technology used together with WDM called optical circuit switching. Essentially, optical circuit switching is the opposite of traditional electrical packet switching: it transmits and moves data entirely in the optical domain.
You can understand the underlying technology this way. In a packet switch, you get a packet with a header, look at it in the electrical domain, and based on the header decide where it is going; the header may have an IP address, so you look up a table: for this destination, which port should it be forwarded from? Basically billions or even trillions of packets per second pass through and are forwarded at extremely high speed. Optical circuit switching says: I do not touch these bits in the electrical domain. What I do is, for an input port, decide where to send the light at the output port. There are multiple ways to implement this. We initially used MEMS switches: microelectromechanical motors control mirrors in three-dimensional space. So we can programmatically control a box that may have 128 ports, 256 ports, or a similar number, mapping each input port entering each fiber to some output port, and then change the mirror angles so that the light really hits these mirrors and reflects to the correct output port.
We initially did this for two reasons. First, to create locality between groups of racks. Suppose I have a compute cluster and a storage cluster jointly supporting, say, search. We know these two clusters will communicate frequently, so we configure the mirrors to create a purely optical shortcut between them, connecting the two rack clusters. Second, we wanted to be able to expand and contract the network without actually moving any fiber. I will not go into the details here; I can draw it on a whiteboard. Optical circuit switching allows you to reconfigure the network spine to expand or contract it without any human action, just a controller managing it.
Fast forward to TPU. TPU has a Taurus topology, directly connecting all TPUs together. Earlier I talked about throughput and goodput. One thing we can do is: if a TPU rack fails, replace it with another TPU rack without moving any fiber. Here too I have to wave my hands or use a whiteboard; but essentially, we can say we always keep a spare rack, and when a rack fails, redirect the light to the new rack, which can be done in milliseconds.
Host: Then why still use fiber? Why not just use free-space optics entirely?
Amin: Good question. Attenuation loss and bandwidth drop sharply. And across a very large building footprint, aligning all the optical paths in three dimensions without fiber would be very difficult, perhaps even impossible. We have discussed it, and there have indeed been some very interesting discussions. But yes, mainly through fiber; once the light enters optical circuit switching, essentially the light shines directly onto these tiny chips. That is one big direction in networking. But there is much more to say, and frankly both the capabilities and the demands of networking in data centers are exploding.
Host: Very cool. We could absolutely do an entire separate episode just on this topic.
Amin: Yes, very cool. Back to that Ineffable photo, what caught everyone's eye was those cables. And ultimately some of those cables will connect into optical circuit switching in our data centers.
Power Is the Constraint: Utilities, Gigawatts, and Data Center Size
Host: Understood. I want to turn to power. You keep talking about goodput per watt and various other per-watt metrics, which makes me think that per watt means power is in some sense the constraint, the scarce condition, or the expensive condition.
Amin: I am often asked: what is the biggest constraint we face? The reality is that there is no single biggest constraint. All constraints are constraints, all are super hard, and they keep changing. But if I had to answer fundamentally, I would say power is the most fundamental constraint we face. Everything else seems to be something we know how to solve over time; power is the long-term binding problem. Of course, nuclear energy may bring abundant clean energy and solve many problems when it arrives; when and at what scale that happens is still unknown.
Host: How does this work in practice? You want to build a new data center and need 1 gigawatt of power. I imagine you cannot just call PG&E and say, "Hey, please send 1 gigawatt over." What does distribution actually look like? Do you have to vertically integrate all the way and even build your own turbines? How do you solve the power bottleneck?
Amin: This is another big and important question. Our preferred model at Google has always been to interconnect with utilities and the grid. So it is indeed like calling the favorite utility company at the data center site, very politely, of course, and with a lot of advance notice, of course. In other words, if you are talking about gigawatt scale, you cannot say, "I want 1 gigawatt tomorrow, when does billing start?" This is something we plan together over many years. For us, we also place great importance on covering infrastructure costs when working with these utilities. This topic could go on for a long time. Because of billing mechanisms, in theory a utility building capacity for us could cause rates to rise for others; what we ensure is that, for example, transmission lines that need to be built or upgraded, additional utility substations, and so on, are paid for by us.
So this is a long-term planning process. It is absolutely possible that, say, in a year I just make up, 2028, we need 1 gigawatt, while the utility can give us 1 gigawatt in 2029 and only, say, 700 megawatts in 2028. Then we may face the question of how to cover that 300 megawatts. One answer is to wait; another is to study how to generate some of that power ourselves, perhaps solar, perhaps with batteries as backup, perhaps from other sources. Then we again work with the utility. It could also be a combination: we keep some local generation capacity, and even if the utility later comes fully online, say at gigawatt scale, we can actually feed power back to the grid. That is, feed it back when they need it. For example, during the two hottest weeks of the year, residential demand is huge; if we have local generation, we can send power back to the grid.
So this really is a multi-year process of working with utilities. Why do you prefer this rather than vertically integrating yourself?
Host: Exactly.
Amin: The main reason is flexibility and uplift for both sides. You can think of it as statistical multiplexing, or the law of large numbers. If we need 1 gigawatt of power and want it at reliability above 99.99%, that may mean we have to build 2 gigawatts of power. At that level, 99.99 or 99.999, you have to have 1+1 redundancy, and that is expensive. And you also have to ensure it is ideally clean energy, ideally right next to the data center, which is also very challenging. Now if we work with the data center, maybe we bring some of our own power, can give power to the grid when they need less, and can draw their power, then through statistical multiplexing over a much larger base, everyone actually wins: we win, the grid wins, residents win. We prefer this approach. In a very small number of cases we also do behind-the-meter projects, but even then, under the assumption that we will work with the utility and interconnect, say, a year later. So this is uplift for us and uplift for the grid.
Host: That makes sense. How do you decide how big to build a data center?
Amin: This is an art, and also a source of enormous debate. I remember more than ten or fifteen years ago, there was even a big debate inside Google: should we put everything into one data center? At that time it would have been 1 gigawatt, which now seems like the laughter of simpler years.
One obvious concern is single points of failure. If you look over a 30-year horizon, that is a huge concern. On the other hand, if you are talking about today's training workloads, bigger is better; from a networking standpoint, you really want things as concentrated as possible over shorter distances. But then there are two problems: single points of failure, and power availability. A gigawatt was large 10 or 15 years ago, but still imaginable; now, putting all of Google's demand in one place is impossible anywhere in this country or the world. So what is the optimal size? We do have models, simulators, and so on. But it also depends on location: in some places we will be at the edge of the network, or in a country where there may be only tens of megawatts; we even work with ISPs, perhaps in Iraq. Okay, that is not a training cluster. Training clusters may be closer to gigawatt scale. Other sites may be hundreds of megawatts, and so on.
Training and Serving Clusters, Seven-Year-Old TPUs, and Open Standards
Host: Very interesting. I want to understand how you manage portfolio lifecycle. I guess the biggest newest clusters are used to train the latest frontier models, and then older equipment is recycled for inference. Is that framework right? Do you also build inference-specialized clusters? How does all of this work?
Amin: Your framework is very reasonable and does make sense. But I would say inference demand is so large that we cannot rely only on old training clusters that are no longer fully occupied by training as the basis for inference. Think again: in a given year we may concentrate training clusters into a few large sites; keeping network distances between them small has benefits. So wherever they are in the world, in a particular year they may be on the same continent, or even in the same region of the same continent. As a result, other continents may not have enough serving capacity. Then we have to build dedicated inference clusters distributed around the world. So your intuition is close, but not entirely. It really has to be: here are training clusters; yes, a few years later they will probably be used for serving; but we also have to build serving clusters at the same time.
Host: Are your serving clusters different from training clusters? Smaller? Cheaper per megawatt?
Amin: Not necessarily cheaper, because serving more needs to put compute, networking, and storage together, which brings us back to the problem of not being able to specialize fully. Training can be done at high density with uniform deployment and so on; serving needs a mix of storage, compute, and accelerators. In addition, there is a very interesting aspect of serving, which we can expand on later: you actually do not want to put too much serving load in the same place; you want to serve user workloads from all over the world. But now we have a single model endpoint, which may actually be different variants of a model, so we also have to distribute these models globally while considering locality. So serving clusters will be smaller, and on the infant side may also be less vertically integrated.
Host: Very interesting. If you built a data center five years ago, the most advanced accelerators then were completely different from today's and far less efficient. Do you really go back and replace the chips in old data centers? I know this relates to an ongoing debate: what is the actual useful life of a chip?
Amin: Yes. I have said this publicly before, and I was a bit surprised by how much response it got: our TPUs from seven or eight years ago still maintain 100% utilization. My intention was not to make some major statement, but apparently it became a meaningful statement. Our older TPUs, and GPUs, but our old TPUs really are still heavily used. Eventually we will replace them. The question is: once the depreciation life of about six years ends, and considering the energy efficiency of new generations and so on, replacement and upgrade really do make sense. It is not that I replace chips, but really replace systems. In other words, we think in pods. For example, a TPU 8 pod may be 9,600 chips, about 140-plus, 152 racks, and so on. So we say we need to pull this pod out and replace it with a TPU 12 or 13 or 14 pod. The new pod's footprint may not perfectly match the old pod's vacancy, so we have to think about that and study how to retrofit. We cannot plan this in advance because we do not know what so many generations of TPU will look like. So this too is an art, and it involves a lot of hard practical work figuring out how to break down old systems and then replace them with new TPUs as quickly as possible.
Host: You have written publicly about open standards. Can you talk about that?
Amin: We talked earlier about interoperability. For us, although we support vertical integration and allow you to extract as much performance as you like, a very important point is: do not force lock-in, do not force an end-to-end walled garden. For example, at Google we developed a model development framework called JAX. We love it and think it is very good, and we use it heavily internally; but many of our customers like PyTorch. One approach would be for us to say: "Hey, if you want to run on TPU, you have to use JAX, because it is best." I am half joking; maybe it really is best, but it is not the only option, and you cannot offer only that choice just because it is very good. The other approach is: if you like JAX, we like JAX, and you can use it; if you like PyTorch, we have Torch TPU, and your unmodified models and so on can also run.
We have seen many examples of this historically. I use the example of IP: why did the Internet Protocol win? In fact, in the 1970s and early 1980s, IP had many competing protocols. IP won because it was an open standard, the narrow waist in an hourglass structure: any software could run on top of it, and any hardware could run underneath it. Open and interoperable; any device that brought IP could plug into a router port and immediately become part of the internet. It was beautiful and allowed the internet to explosively grow and scale around the world. So we truly believe in these open standards and in plug-in access points. If you want to connect something highly specialized, you can; if you think there is something better that connects to our framework, you absolutely can. But we want to truly support open standards, ideally with open source around them. This buildout is too big and too important to become any single vendor's closed proprietary stack. We truly believe this; it must be a choice. This is also why, for example, we fully support and have TPUs, GPUs, other accelerators, and so on.
Host: How is AI changing your team's daily work? At what layer has it changed your function the most?
Amin: The easiest answer is on the software engineering side. There has already been a lot of external documentation on this. I think Google and my team are using it very effectively to improve software development capability, and frankly also for testing and releases, and even to help with design. But on the hardware side, perhaps with less external coverage, there have also been significant changes. In other words, my hardware engineers, and I looked at the data earlier today, are using AI as much as software engineers. You can look at token counts, which is not the best metric, but still a metric. Hardware engineers are already using AI at levels comparable to software engineers. Productivity has also risen. The time from design start to tape-out is shortening; bring-up time is also shortening. Another place where productivity has risen significantly, and perhaps more unexpectedly, is that the way data centers are designed has also changed markedly. That is, how we do data center design and how we evaluate it. You just mentioned whether we build a 1-gigawatt monolithic building or a 100-megawatt campus, a 200-megawatt campus. In the past that would have been a very painstaking, very spreadsheet-driven, human-driven process; now it is still that to some degree, but there is already a lot of AI involvement, genuinely streamlining planning and development processes.
Host: Interesting. Are they reasoning models?
Amin: It is not entirely accurate to call them reasoning models yet. It has not replaced human judgment, but it has indeed made it much easier to bring the necessary information together in one place, basically putting the right information in front of the decision maker.
Orbital Data Centers and the 2036 Supercomputer
Host: I will close with two somewhat freewheeling but fun questions. First: orbital data centers. I saw Google is taking this quite seriously. You want me to comment on it. Does the fact that people are seriously calculating orbital computing itself mean there are serious constraints on Earth? What do you think about orbital data centers?
Amin: This is an exciting direction. We really are pursuing it, and we call it a moonshot without any irony; it is one of the big projects we are very happy to invest in. Back to the point you correctly raised earlier in the conversation: from the standpoint of fundamental constraints, energy and energy production are the key challenge. Fundamentally, in space, because there is no atmospheric attenuation and so on, usable power is about 40% higher. That is, just the sunlight capacity is 1.4 times. That is one thing. But in a sun-synchronous orbit, your solar cells can get 98% to 100% sunlight coverage, whereas on land it may be only 28%, 30%, perhaps 35%. So energy is extremely abundant: obviously, you get 1.4 times, multiplied by three to four times the number of sunlight hours per day, and largely remove the battery from the equation, so now it becomes possible to get something truly substantial. Of course, it is carbon-free, with many benefits, and also many challenges.
There are indeed many challenges. For example, cooling: you might naively think cooling is easier in space, but it is actually harder. Reliability: we just talked about how these devices sometimes fail. Repair becomes harder in space, not impossible, but harder. Your earlier pre-question about free-space optics now becomes real: we probably will not run fiber between these components, so it really will be free-space optics, with lasers pointed at receivers and calibrated in real time. There is no fundamental showstopper here.
Host: One last question. I keep picturing the supercomputer you built for Ineffable. So the question is: what will the most advanced supercomputer look like ten years from now?
Amin: Oh boy. Ten years is exactly on that boundary. I will say very honestly: some people may already be thinking about what computers in 2036 will look like. We may also have a few people thinking that far out. But the cone of uncertainty there is simply too wide. Looking at trends, integration will be astonishing. We saw the beautiful fiber in the VR200 rack assembled for Ineffable. My guess is it will be more integrated, with less fiber. I will not say there will be no fiber in 2036, but from the rack perspective, I think racks will appear more tightly integrated, possibly with only a small bundle of fiber coming out of the rack. And from a modular manufacturing standpoint, these racks in 2036 will very likely be centrally manufactured. Whether inside there are 72, 144, 288, or, since we mentioned the Ineffable case, perhaps 576, 1,152, choose any multiple of GPUs; the same applies to TPUs; GPUs and TPUs will both be deeply integrated into the rack. Will a rack be multi-megawatt? We do not know exactly how, but by 2036 we can imagine: a single rack might reach several megawatts. At that point you bring in water, bring in power, bring in fiber, roll the rack in, plug in these three things, and off it goes.
Host: So by 2036, it will not be a big alien sphere floating in space?
Amin: I will not say it is impossible by 2036 to launch it directly into space and have a space station robotic arm catch it and plug it into the right module. Maybe by 2036 that is exactly what happens.
Host: It is fun just to think about. I really enjoyed this conversation. This is the most intense and largest capital expenditure buildout in history, but I also think it is a technological revolution, and a beautiful thing. You have a very deep grasp of the beauty of technology, the constraints, and how to balance all of this. Google is in good hands with you. Thank you for taking the time to share what you are doing.
Amin: Very happy to. Thank you very much. This was a great conversation. We have the chance to live through all of this, and the chance to define it. Thank you.