OpenAI’s Compute Chief: We Can’t Build Fast Enough | Sachin Katti

Quick Overview

OpenAI cannot build compute infrastructure fast enough because demand for AI performance is growing much faster than the physical capacity to supply it, necessitating an active, multi-pronged strategy to secure power, chips, and data centers. While the company is on a clear path to its committed $500 billion, 10-gigawatt goal, the current reality involves constant, intensive efforts to manage the physical, logistical, and financial constraints of building the world's largest supercomputers.

Key Points: OpenAI plans to spend $50 billion on computing power this year alone, a figure that continues to grow rapidly. Demand for AI compute currently far outstrips supply, forcing the company to take an active role in building and funding necessary infrastructure. The company's strategy involves a portfolio approach using multiple cloud partners, chip platforms, and direct investment in grid and power generation. Data centers are evolving into massive, liquid-cooled supercomputers, with everything from cables to power transformers requiring specialized cooling and design. OpenAI and its partners are investing in new U.S. data center capacity, including a 4.5-gigawatt agreement with Oracle and significant partnerships with Microsoft, Amazon, and CoreWeave. The company is co-designing custom chips, such as Jalapeño, to optimize for its specific LLM inference and training workloads.

Context: This video features an interview with Sachin Katti, the Head of Compute Infrastructure at OpenAI, discussing the massive, unprecedented scale of infrastructure required to power the current AI boom. The discussion centers on the physical, financial, and logistical challenges of building the world's largest data centers, covering topics like liquid cooling, power grid constraints, custom chip design, and the company's multi-partner investment strategy.

Detailed Analysis

The AI boom is driving an unprecedented need for compute power, forcing OpenAI to move beyond simply renting cloud capacity to actively designing and funding the world's largest infrastructure projects. Because physical supply chains for data center components like power transformers and gas turbines are slow to move, the company faces a constant, insatiable demand that necessitates a diverse, multi-partner approach. This includes partnerships with major cloud providers, direct investment in grid infrastructure, and the development of custom silicon like Jalapeño to maximize tokens per watt. The company rejects the notion of a simple 'training vs. inference' divide, treating both as fundamental, compute-intensive tasks. They are also prioritizing reliability and efficiency by co-designing custom hardware, utilizing liquid cooling at both the data center and chip level, and implementing advanced networking protocols like MRC to ensure that training runs on massive GPU clusters remain resilient against component failures.

Raw markdown version of this recap