Part 3 began with the weight of an RTX 4090: two kilograms of aluminium fins, copper pipes, and fans wrapped around a square of silicon the size of a postage stamp. It said the metal was there to carry away heat, and moved on. Watts have turned up in nearly every part since. 450 W for that card, 700 W for an H100, 1,000 W and more for a Blackwell, 120 kW for the rack in Part 13, an estimate of 600 kW for the rack in Part 20. This part stops and asks the question those numbers have been begging: where does the electricity actually go?
The answer runs from the smallest scale in the series to the largest. It starts inside a single transistor, where flipping a switch costs a definite amount of energy, and it ends at sites that draw as much power as a city, because a modern AI data centre is, in the end, a machine for turning electricity into w × x + b and heat. Following the watts is also the last way of seeing the memory wall, since it turns out that moving a number costs far more energy than multiplying it, and that fact, measured in joules, explains the design of every chip in Parts 17 to 20 as well as anything in Part 10 did.
Where a watt goes inside a chip
Go back to the MOSFET of Part 3. To turn it on, its gate has to be charged up to about a volt, and to turn it off, that charge has to be drained away. The gate is a tiny capacitor, and charging a capacitor costs energy, a fixed amount set by its size and the voltage squared. It is a preposterously small amount, a few hundred attojoules for a modern transistor, a billionth of a billionth of a joule. Then you multiply. A chip has billions of gates, a fair fraction of them flip on every tick, and there are a few billion ticks a second. Small times billions times billions is tens or hundreds of watts, and every one of those joules ends up as heat in the silicon. This is dynamic power, the cost of doing work, and it rises with three things: how many switches there are, how fast the clock runs, and, most steeply of all, the voltage, because the energy per flip goes with the voltage squared.
There is a second bill that has to be paid even when nothing flips. A transistor that is "off" is not a perfect insulator. A trickle of current leaks through the channel and through the insulating layer under the gate, which is only a few atoms thick, and the smaller the transistor the worse it leaks. This is static power, or leakage, and on a modern chip it can be a quarter or more of the total. It is the reason a GPU that is doing nothing at all still draws tens of watts, and it is the physical reason behind the end of Dennard scaling that Part 3 described. For thirty years, shrinking a transistor let you lower its voltage, which cut the energy per flip, which let you raise the clock for the same heat. Around 2005 the voltage could not go much lower without leakage swamping everything, so the clock stopped rising, and the only way left to do more work per second was to do more of it at once. A GPU is what that looks like: a lower clock than a CPU, and thousands of times more switches flipping in parallel. It does not use less power. It gets more work out of each joule.
The joule cost of a byte
Now put a price on the running calculation. In 2014 Mark Horowitz of Stanford gave a talk at the ISSCC circuits conference, Computing's energy problem, that has been quoted in every hardware paper since, because it puts the whole of this series into one table. The figures are for a 45 nm process, so a modern chip does each of these for several times less, but the ratios are what matter and they have barely moved.
| Operation | Energy | Compared to a 32-bit integer add |
|---|---|---|
| 8-bit integer add | 0.03 pJ | a third |
| 32-bit integer add | 0.1 pJ | 1 |
| 16-bit floating-point multiply | 1.1 pJ | 11 |
| 32-bit floating-point multiply | 3.7 pJ | 37 |
| Read 32 bits from an 8 KB SRAM cache | 5 pJ | 50 |
| Read 32 bits from DRAM | 640 pJ | 6,400 |
A picojoule is a millionth of a millionth of a joule, so the absolute numbers are meaningless on their own. Read the last column instead. The multiply at the heart of w × x + b costs a few picojoules. Fetching one of its operands from main memory costs about two hundred times as much as the multiply itself. Part 7 measured the memory wall in time, hundreds of ticks per trip against four per multiply-add. This is the same wall measured in energy, and it is steeper. Compute is cheap and bytes are expensive, in joules as much as in nanoseconds.
Two more rows carry the rest of the story. A 16-bit multiply costs about a third of a 32-bit one, and an 8-bit add a third of that again, because the multiplier of Part 4 shrinks with the square of its width and a smaller circuit charges fewer gates. That is the precision lever of Part 11 as an energy saving: FP8 and FP4 are not only more arithmetic per square millimetre, they are less energy per operation and, because each number is fewer bytes, less energy per operand moved. And a read from a small on-chip SRAM costs a hundredth of a read from DRAM, which is the register file, the shared memory, and the L2 cache of Part 10, and also HBM on the package rather than GDDR across a circuit board. Every design move since Part 10 was a way of moving fewer bytes, and each one was saving joules as much as time. The 4090's 450 W is not mostly arithmetic. A large share of it is the cost of getting the numbers to the arithmetic.
What the number on the box means
When a spec sheet says 700 W, it is quoting the thermal design power (TDP), which Nvidia calls total graphics power (TGP) on gaming cards and total board power on others. It is not what the chip draws at every moment. It is the amount of heat the card is designed to shed continuously, which is the same thing as the most power it is allowed to draw for more than an instant. A GPU has a power limit set in its firmware, and when a heavy workload pushes it there, the chip lowers its clock until it is back under the line, a process called throttling. The boost clock on a spec sheet is the speed the chip will run at if the power and the temperature allow, and under a sustained matrix multiply on every SM it often does not. Two identical chips with different power limits are, in practice, two different products.
That is the difference between the two H100s you will see on price lists. The H100 PCIe is a card that plugs into an ordinary slot, is cooled by the server's airflow, and is limited to 350 W. The H100 SXM is the same silicon on a bare module, bolted flat onto an HGX board with a heatsink the size of a hand, wired to NVLink, and allowed 700 W. The SXM version has more of its SMs enabled and runs them at a higher clock, so it is a good deal faster, and its FP8 figure of 1,979 TFLOPS is the one every table in this series has used. SXM is nothing more than the name of that module form factor. Every "SXM" in Parts 18 to 20 means "the version with the bigger power budget", and the gap between 350 and 700 W is most of the gap in performance.
The card as a power system
A graphics card gets its electricity at 12 volts from the computer's power supply, and it arrives by two routes. The slot itself can supply 75 W, which is enough for a modest card and nothing more. The rest comes down cables from the power supply into connectors on the card's edge: the older 8-pin connector carries 150 W, and the 12V-2x6 connector introduced with the RTX 40 series (as 12VHPWR, then revised after a run of melted plugs) carries up to 600 W through one small socket. The RTX 4090 draws up to 450 W and the RTX 5090 575 W, which is why both need the new connector and why the makers recommend an 850 to 1,000 W power supply for a PC built around one.
The chip cannot use 12 volts. Its transistors run at around one volt, so between the connector and the die sits a block of voltage regulators, the row of inductors and capacitors you can see along the edge of any card. Their job is to turn 450 W at 12 volts into 450 W at about one volt, and that arithmetic is worth doing: 450 watts at one volt is 450 amps. Recall the car starter motor from Part 1, which draws one to two hundred amps for a few seconds through a cable as thick as your finger. A GPU under load draws two or three times that, continuously, into a piece of silicon two and a half centimetres across, through the layers of the package and hundreds of tiny solder bumps. Delivering the current is as hard a problem as delivering the data, and the regulators are a good share of what makes a high-end card long, heavy, and expensive.
Then all of it becomes heat. Every joule that goes into the chip comes back out as warmth, with no exceptions, because the switches do no mechanical work and store nothing. A 450 W card is a 450 W heater, indistinguishable from a small electric fire, and it has to get rid of that heat from a surface of about six square centimetres. That is the two kilograms of metal: a copper plate pressed against the die, heat pipes carrying the warmth out to a stack of aluminium fins, and fans blowing room air through them. It is the same problem ENIAC's engineers faced in Part 2, with a ventilation system for 150 kW of tube heaters, and the answer has not changed in principle: move air, or move water.
The ladder: chip, node, rack, hall
Now climb. The unit of "a GPU" grew through Part 13 from a die to a rack, and the power grows with it, faster than the count, because everything around the GPUs draws power too.
| Unit | GPUs | Power | For scale |
|---|---|---|---|
| RTX 4090 | 1 | 450 W | a small electric heater |
| H100 PCIe | 1 | 350 W | |
| H100 SXM | 1 | 700 W | |
| B200 | 1 | 1,000 to 1,200 W | a kettle, on all day |
| DGX H100 (one 8-GPU node) | 8 | 10.2 kW maximum | eight homes' average draw |
| DGX B200 (one 8-GPU node) | 8 | 14.3 kW maximum | |
| An ordinary rack of web servers | 0 | 5 to 10 kW | |
| GB200 NVL72 (one rack) | 72 | about 120 kW, up to 132 kW | a hundred homes |
| GB300 NVL72 (one rack) | 72 | up to 142 kW | |
| Kyber rack (Rubin Ultra, 2027) | 144 | estimates of 600 kW, since revised | five hundred homes |
| A 100 MW AI data centre | tens of thousands | 100,000 kW | 100,000 homes, by the IEA's figure |
| The largest AI site as of mid-2026 | hundreds of thousands | about 950 MW | a city |
The node figures are from Nvidia's own DGX H100 and DGX B200 user guides, and they show the overhead: eight H100s at 700 W is 5.6 kW, but the node draws up to 10.2 kW, because the two CPUs, two terabytes of memory, the NVSwitches, eight network cards, and the fans and power supplies all count. The rack figures come from the server makers, since Nvidia does not publish them, and the "homes" column uses the US average of about 10,800 kWh a year, which is 1.2 kW around the clock.
Something changes in the middle of that table. An ordinary data centre was designed around racks drawing five to ten kilowatts, fed by power strips and cooled by cold air blown up through the floor. A single NVL72 rack draws more than ten of those. Power is no longer delivered by cables to each server: the rack has its own busbar, a copper bar down the back carrying around 50 volts DC from a bank of power shelves in the same rack, and every tray clips onto it. And air can no longer carry the heat away, which is the next section. The "AI factory" that Nvidia talks about is largely this: a building designed from the start for racks that draw a hundred kilowatts or more, which almost no data centre built before 2023 can accommodate without rebuilding its power and cooling.
Try it: the power ladder
Pick a unit and see what it draws, what it costs to run, and how many homes it stands for. The two sliders are the two numbers a data centre operator argues about most: the price of electricity, and how much extra the building itself uses on top of the computers, which the next section explains.
Two things to try. Press RTX 4090 and then 1 GW campus: the same arithmetic spans nine orders of magnitude, and the bill goes from a few hundred dollars a year to something over a billion. Then take any unit and drag the PUE slider from 1.1 to 1.6: that is the difference between the best hyperscale buildings and the industry average, and at the campus scale it is hundreds of millions of dollars a year spent on nothing but fans, pumps, and chillers. The last three rows are the water, and the section after next explains them.
Air gives out, and liquid takes over
Air is a poor carrier of heat. A rack cooled by air needs a large volume of it moving through every server, and there is a limit to how much you can push through a box without the fans themselves becoming a major consumer of power and the noise becoming dangerous. In practice air stops making sense somewhere between twenty and forty kilowatts per rack, which is why the racks of Part 13 needed something else.
The something else is water, or more precisely a water-and-glycol mix, piped straight to the chips. In direct-to-chip liquid cooling, a small metal plate with channels inside it, the cold plate, is clamped onto each GPU and CPU in place of a heatsink, coolant flows through it, and the warm coolant is piped out of the rack to a coolant distribution unit (CDU), which passes the heat into the building's water loop and sends the coolant back cool. Water carries heat far better than air, so the loop can be run warm, forty-five degrees on the supply side in some designs, and still keep the silicon within limits. Every tray of a GB200 NVL72 has cold plates and quick-release couplings, the rack has a manifold running down the back beside the busbar, and a row of them shares a CDU the size of a large fridge. Some operators go further and submerge whole servers in a tank of non-conducting fluid, immersion cooling, which does away with fans entirely.
Liquid cooling is not only about keeping up. It is cheaper to run. A warm-water loop can often shed its heat to the outside air through a dry cooler without a compressor-driven chiller, and pumping water takes far less power than blowing the equivalent air. That shows up in the number the next section is about.
PUE, and what the building costs
A data centre draws more power than its computers. Cooling, pumps, lighting, the losses in every transformer and power supply between the grid and the chip, and the batteries and generators kept ready for a grid failure all take their share. The ratio of the whole building's draw to the computers' draw is called power usage effectiveness (PUE), and it is the number by which data centres are judged. A PUE of 2 means that for every watt reaching a server, another watt is spent around it. A PUE of 1 would mean a building with no overhead at all.
Uptime Institute's annual survey puts the industry average at about 1.56, where it has sat since 2020, having fallen from 2.5 in 2007. Buildings less than five years old average around 1.45. The hyperscale operators do much better: Google reports a fleet-wide annual PUE of 1.09, which means nine per cent overhead on top of the computers. The demo above lets you see what that difference is worth. At a 100 MW site, moving from 1.56 to 1.09 saves 47 MW, the whole draw of a mid-sized town, for identical computing.
PUE is a ratio, though, and the AI buildout is pushing the other number, the computers' draw, up faster than any efficiency gain pulls the ratio down. That is the last rung.
Where the water goes
Every watt in the table above becomes heat, and the heat has to leave the building. There are two ways out. It can go into the air, through a radiator and fans, which costs electricity. Or it can go into water that evaporates, which costs water. Evaporation is a remarkably good way to move heat: turning a litre of water into vapour carries away about 0.65 kWh, what a kettle uses in about twenty minutes, with no compressor involved. That is what a cooling tower is, a box in which the building's warm water trickles down through a draught of air and some of it evaporates, taking the heat of the whole hall with it into the sky. A 120 kW rack cooled that way evaporates about 180 litres an hour, 1.6 million litres a year, and the number scales with the ladder like everything else in this part.
The word to watch is evaporates. Water that runs through a pipe and back to the river is borrowed. Water that leaves a cooling tower as vapour is gone from the local supply, and the industry's own accounting separates the two as withdrawal and consumption. With clean water, roughly 80 per cent of what a tower withdraws is consumed, and the rest is periodically dumped as the minerals in it concentrate and replaced with fresh. So the honest measure of a data centre's water is consumption, and the metric for it is a sibling of PUE: water usage effectiveness (WUE), litres of water consumed on site per kilowatt-hour the computers use.
This is where the liquid cooling of the previous section causes confusion, because "liquid-cooled" sounds like "uses more water" and usually means the reverse. The loop through the cold plates is closed. The same few hundred litres go round and round and are never consumed, exactly like the coolant in a car. What consumes water is the last step, where the building's heat is finally dumped outdoors, and that step can be a cooling tower or a dry cooler, a radiator with fans that uses no water at all but more electricity, and which only works well when the loop is warm relative to the outside air. The forty-five degree supply water of a GB200 rack is warm enough for a dry cooler to shed its heat in most climates for most of the year. That is the trade at the heart of the water question: a site can spend water to save electricity or spend electricity to save water, and PUE and WUE move in opposite directions when it does. Microsoft's zero-water design, adopted for every new site from August 2024, makes exactly that choice: chip-level cooling in a closed loop with no evaporative step, which the company says avoids more than 125 million litres a year per data centre and raises its energy use by a "nominal" amount. Its fleet WUE was 0.30 litres per kWh in its last reported year, down from 0.49 in 2021.
Geography matters more than design. A 2023 study by Pengfei Li and colleagues, Making AI Less Thirsty, worked from Microsoft's own site-level disclosures and found the on-site figure ranging from 0.01 litres per kWh in Denmark, where the outside air does the cooling most of the year, to 1.63 in Arizona, where it does not. The same rack consumes a hundred and sixty times more water in one place than the other, and the places with cheap land and cheap power are often the dry ones.
There is a second bill that most company figures leave out. Making electricity uses water too. A coal, gas, or nuclear plant boils water to spin its turbine and evaporates more to condense it again, and a hydroelectric reservoir evaporates from its surface, so a kilowatt-hour from the average American grid consumes about 3.1 litres of water before it reaches the data centre, against nearly none for wind and solar. The same study calls the on-site water scope 1 and the power plant's water scope 2, and for a site with a dry cooler and a thermal grid, scope 2 is nearly all of it. When you see a water figure, the first question is which of the two it counts.
The totals are large in one sense and small in another. Lawrence Berkeley National Laboratory's 2024 report for the US Department of Energy put the direct on-site water consumption of American data centres at 17.4 billion gallons, 66 billion litres, in 2023, and projected 38 to 73 billion gallons by 2028 as the AI sites come online, with the hyperscale sites about half of it. Against the country, that is under half a day of the public water supply, spread over a year. Against a town, it is a different matter: a single large campus can consume a few million gallons a day, which is the water of a city of thirty thousand people, and it is drawn from one aquifer or one river in one county, quite often in the American southwest. The national number is a rounding error and the local number can be a crisis, and both are true at once.
The site
The International Energy Agency's 2025 report Energy and AI put the world's data centres at 415 terawatt-hours in 2024, about 1.5 per cent of global electricity, and projected that figure to more than double to around 945 TWh by 2030, with AI the largest driver. Its comparison is the one to keep: "a typical AI-focused data centre consumes as much electricity as 100,000 households, but the largest ones under construction today will consume 20 times as much." A 100 MW site is a normal one. Epoch AI, which tracks the largest sites, recorded the record for a single AI campus passing 900 MW in the first half of 2026, and notes that the record has doubled roughly every ten months since 2024. A gigawatt is the output of a large power station, and the sites now being built are planned in multiples of it.
At that scale the constraint stops being chips. Getting a gigawatt from the grid means transmission lines and substations that take years to build, so the largest sites are being placed next to power plants, or building their own gas turbines on site, and the electricity supply, not the HBM supply that Part 20 worried about, is what sets the timetable. The unit of "a GPU" that Part 13 watched grow from a die to a rack has one more level: the campus, whose spec sheet is written in megawatts and whose lead time is set by the power company.
What one answer costs
All of that is the supply side. The question people actually ask is the demand side: what does it cost, in electricity, when I send a chatbot a message?
For a long time the only answers were estimates from outside, and they varied by a factor of a hundred. In August 2025 Google published measured figures for its own service: the median text prompt to the Gemini app used 0.24 watt-hours of electricity, which it compared to watching television for under nine seconds. That figure counts the whole stack, the TPUs doing the arithmetic, the CPU and memory beside them, the idle capacity kept ready for peaks, and the building's overhead. Counting only the accelerator, the way most outside estimates had, gives 0.10 Wh, so the machinery around the chip roughly doubles the bill, which is the PUE and the DGX overhead of this part seen from the other end. Google also reported that the figure had fallen 33 times in a year, through software and batching, which brings the roofline of Part 10 back one last time.
Recall that generating text is memory-bound: for each token the GPU reads every weight once and does one multiply-add with it, so the tensor cores sit mostly idle. Idle silicon still draws power, the leakage from the first section plus everything on the board that does not switch off. A GPU serving one user at a time therefore burns most of its watts waiting for bytes. Serve sixty-four users in one batch and the same bytes feed sixty-four times the arithmetic, the lanes do real work, and the energy per token drops by nearly that factor. Batching, quantisation, and the split into prefill and decode machines from Part 20 are as much energy engineering as speed engineering, and they are why the cost of an answer keeps falling even as the sites keep growing.
The same question gets asked about water, and the answers in circulation differ by a factor of a thousand, so it is worth showing where each comes from. Google's August 2025 measurement included water: the median Gemini text prompt consumed 0.26 millilitres, about five drops, counting the cooling water at its own sites. OpenAI's Sam Altman gave a similar figure two months earlier, 0.000085 gallons, which is 0.32 millilitres, or roughly a fifteenth of a teaspoon. Both are scope 1 numbers, on-site water only, at the operators' fleet-wide WUE. Add the water behind the electricity, 0.24 Wh at the American grid's 3.1 litres per kWh, and the median prompt comes to about a millilitre all in. Then there is the figure everyone has heard, that a ChatGPT query uses a 500 ml bottle of water. It comes from the 2023 study above, which estimated that GPT-3 consumed a bottle's worth for every 10 to 50 medium-length responses, on-site and off-site water together, using power figures from 2020. Somewhere between the paper and the headline the "10 to 50" fell off. Read correctly it says 10 to 50 millilitres per response for a five-year-old model on older serving software, and the operators' measured numbers for today's models sit a further ten to fifty times below that, which is in line with the efficiency gains since, Google alone reporting a 33-fold drop in energy per prompt in a single year. A long conversation, a large reasoning model, or a site in Arizona push the number back up by ten or more. So the defensible statement is: a few drops to a few spoonfuls per exchange, depending mostly on where the data centre is and what its grid burns, and a ten-minute shower is somewhere between twenty thousand and three hundred thousand prompts. What makes the total matter is not the per-query figure but the multiplication: a billion prompts a day at one millilitre is a million litres a day, which is the site-level number from the previous section seen from the other end.
The chips themselves get more frugal too, though not as fast as the headlines suggest. Divide each GPU's dense FP8 throughput by its power and the RTX 4090 manages about 1.5 trillion operations per second per watt, the H100 about 2.8, and the B200 about 4.5. Three times better in two generations, mostly from HBM, bigger tiles, and fewer bits per number rather than from faster transistors, which is exactly the list this series has been keeping since Part 10.
Where this leaves us
Every joule that enters a GPU leaves as heat, and the joules go three places. Flipping switches is the arithmetic. Holding switches still is leakage. And, above all, moving numbers to the switches costs a hundred times more per byte than the multiply that uses it. That last fact is the memory wall in its final form, and it is why the last four generations spent their transistors on memory, interconnect, and precision rather than on lanes. Climb the ladder and the same watts add up: 700 for a chip, 10,000 for a node, 120,000 for a rack that has to be plumbed rather than fanned, and a hundred million for a site that stands for a city's worth of homes and is now large enough that the power company, not the chip company, decides when it opens.
The series has one more thing to look at before the picture is put together. Every chip in it so far has been an Nvidia GPU, and the phone in your pocket, the laptop on your desk, and the racks at Google run their neural networks on something else.
Next: NPUs, TPUs and the Chips in Your Phone: what a neural processing unit is, why a five-watt chip advertises 50 TOPS, how a systolic array does w × x + b without a warp or a cache, and where the GPU stops being the right tool.