Part 3 showed how one layer of a chip is printed, why chipmakers chase density rather than size, and why no die can be larger than the 858 mm² a scanner can expose in one shot. It stopped at the wafer. This part follows the wafer the rest of the way: where it came from, what the machine that prints it actually is, what happens after the last layer is printed, which is the test, the saw, and the sorting into products, and the packaging that decides how much memory a GPU has and how fast it can reach it. Along the way it puts numbers on two things Part 3 described in words, how many dies fit on a wafer and how many of them work, because those two numbers are most of the economics of the chips in the next few parts. The words to collect are wafer, die, yield, binning, package, and interposer.

It is also the answer to three questions that have been left hanging. Why does an RTX 4090 have 128 working SMs when the chip has 144 on them, and why does the H100 have 132? How do the two dies of Blackwell in Part 19 actually get joined, and the eight stacks of memory beside them? And why does the memory of a data centre GPU sit on the same package as the processor, while a gaming card's memory sits in separate chips around it? All three are manufacturing questions, and none of them makes sense until you have seen a wafer.

From sand to a wafer

The raw material really is sand, or rather quartz, which is silicon dioxide. It is heated with carbon to strip the oxygen, then purified through a chain of chemical steps until fewer than one atom in a billion is anything but silicon. Nine nines, the industry calls it, 99.9999999 per cent. Then the purified silicon is melted in a crucible, a small seed crystal is dipped into the surface, and it is pulled upward very slowly, rotating, over a day or two. The molten silicon freezes onto the seed in the same crystal orientation, and what comes out of the furnace is a single flawless crystal, a cylinder 300 mm across and a metre or two long, called an ingot. The method dates from 1916 and is named after Jan Czochralski, who discovered it by accidentally dipping his pen into molten tin instead of an inkwell.

The ingot is sawn into discs about three quarters of a millimetre thick, ground and polished until the surface is flat to within a few atoms, and those discs are wafers. A wafer is the unit of the whole industry. A factory's capacity is quoted in wafers per month, a foundry's prices are quoted per wafer, and every chip that exists began as a small square on one of them. The size has grown over the decades, from 100 mm to 150 to 200, and since around 2002 the leading edge has been 300 mm, a dinner plate. A move to 450 mm was planned and abandoned, so 300 mm is the size to keep in your head.

The machine that prints it

Part 3 walked through one layer: coat the wafer with resist, expose it through a mask, develop, etch, strip, and repeat for the sixty to a hundred layers that make a chip. What it did not describe is the machine, and the machine is where the money and the limits are.

The machine that does the exposing is called a scanner, and it does not expose the whole wafer at once. Its optics can only project a sharp image over one rectangle 26 mm by 33 mm, the exposure field of Part 3, so it exposes one rectangle, steps the wafer over, exposes the next, and so on across the whole disc, a hundred or so times per layer, each field aligned to the layers beneath it to within a few nanometres. Each layer of each chip on the wafer is the same image, printed from the same mask, which is why the cost of a mask set, well over ten million dollars for an advanced process, is spread over every wafer that ever runs through it.

The light matters, because optics cannot draw a line much narrower than the wavelength they use. For twenty years the industry used deep ultraviolet light at 193 nm and stretched it with a chain of tricks: shining it through a layer of water to shorten the effective wavelength, and splitting one dense layer into two or three separate exposures, which is called multi-patterning. Since 2019 the finest layers have used extreme ultraviolet at 13.5 nm, made by hitting droplets of molten tin with a laser fifty thousand times a second and collecting the glow with mirrors, because no lens is transparent at that wavelength. The machine is built by one company, ASML in the Netherlands, weighs about 180 tonnes, ships in forty freight containers, runs in vacuum, and costs on the order of €180 million. Its successor, with wider mirrors, costs over €350 million and can only print half the field, 26 mm by 16.5 mm, which is going to make "how big can a die be" an even sharper question later this decade. A 5 nm-class process uses EUV on a dozen or more of its layers and the older light for the rest, and the whole run through the fab takes about three months.

Wafer, 300 mm the same pattern printed ~90 times Die, ~25 mm square one copy, cut from the wafer HBM HBM interposer (blue) on substrate (green) Package die + memory stacks, wired together Card or module package on a board, with power and cooling
Four things that all get called "the chip". The wafer carries many dies. One die is cut out and mounted on a package, with memory beside it on a data centre part. The package goes on a board. Spec sheets quote the die's size and transistor count and the package's memory.

What a die is

Look at the wafer figure above. The scanner printed the same pattern over and over, in a grid, until it ran out of disc. Each copy of the pattern is a die. Part 3 used the word for the silicon rectangle you would see under a chip's lid, and that is exactly what it is: one printed copy, still attached to its neighbours until the wafer is sawn along the narrow scribe lanes between them. It is the object everyone means by "the chip" when they quote a size in square millimetres or a transistor count. The word comes from the same root as the dice you roll, and the plural is dies (or, in older texts, dice).

How many fit is simple arithmetic with an awkward edge. A 300 mm wafer has about 70,700 mm² of surface. Divide by 608 mm², the AD102 of Part 17, and you get 116. But dies are square and the wafer is round, so the ones that would straddle the edge are lost, and a few millimetres round the rim are unusable anyway. In practice about 90 AD102-sized dies fit on a wafer, and only about 65 of the 814 mm² GH100. That is the first reason die size is on every spec sheet: it decides how many chips a wafer, which has a fixed price, can produce.

Yield, and why the 4090 has 128 SMs out of 144

The second reason is worse, and Part 3 stated it in words: flaws land on the wafer at random, and a big die is a big target. Here are the numbers. Not every die works. A speck of dust, a flaw in the resist, a wobble in one exposure, and a single track is broken or shorted somewhere among the billions. On a mature process the defects land at something like 0.1 per square centimetre, which sounds tiny until you multiply by the area of a GPU. A die of 1 cm² has a good chance of being clean. A die of 8 cm² is eight targets side by side, and the odds of all eight being clean fall fast. The rule of thumb is that the fraction of clean dies is e to the power of minus (area times defect density), so at 0.1 defects per cm² a 100 mm² phone chip yields about 90 per cent, a 608 mm² AD102 about 54 per cent, and an 814 mm² GH100 about 44 per cent. More than half the biggest dies come off the wafer damaged. That is the yield, and improving it is a large part of what a foundry does with a process in the years after it launches.

Now the trick that turns a disaster into a product line. A GPU is 144 nearly identical SMs plus a small amount of shared logic. If a defect lands in one SM, the other 143 are fine. So the chip is designed so that any SM can be switched off after manufacture, the die is tested, the damaged SMs are fused away, and the chip is sold as a product with fewer of them. That is why Part 9 could say the RTX 4090 has 128 of 144 SMs and the H100 has 132 of 144. The number 128 was chosen so that most dies with a handful of defects still qualify, and a die with more damage can drop a tier further: the H100 PCIe card uses 114 SMs of the same silicon. Sorting dies this way, by how much of them works and how fast they will run, is called binning, and it is the reason one die design turns into a family of products at different prices. Only a defect in the shared logic, the L2 slice or a memory controller or the scheduler that hands out work, actually kills a die. Memory gets the same treatment: the GH100 has six HBM sites on the package and the H100 uses five, so a package with one bad stack still ships.

The demo below lets you feel the arithmetic. It prints a wafer with dies of the size you choose, sprinkles defects at the rate you choose, and counts.

Try it: print a wafer

Two things to try. Set the die edge to 10 mm, the size of a phone chip, and nearly everything is green. Drag it to 29 mm, which is close to the reticle limit, and at the same defect rate half the wafer is red or ochre while only about sixty dies fit at all. Then set the defects to 0.30 per cm², which is roughly what a process looks like in its first year, and see what a big die costs before the foundry has had time to clean things up. That is why the biggest chips of each generation arrive a year or two after the first phone chips on the same node.

Two walls, one way out

Part 3 ended at the reticle wall: 858 mm² is the largest image the scanner can project, and it does not move. The demo above is the other wall, and it moves the wrong way, because yield punishes area long before the reticle does. For fifteen years the way around both was to shrink: the same design on the next process is a smaller die, so more fit and more are clean, and the freed area is spent on more SMs. Part 3 explained why that stopped paying off as it used to. When the process no longer shrinks fast enough and the die is already at the limit, there is exactly one move left, more than one die, and Part 3 described the bridge that makes two dies act as one. What it did not describe is the rest of the package the bridge sits in, which is where the memory story of Part 10 is finally decided.

Packaging: the die is not the chip you buy

A bare die cannot be soldered onto a circuit board. Its connection points are a few tens of micrometres apart and there are thousands of them, while a board's tracks are a hundred times coarser. The package is the translator between the two scales. The die is flipped face down and attached to a small, dense, multilayer board called the substrate through a grid of tiny solder bumps, and the substrate fans the connections out to a few thousand larger balls on its underside that meet the card. What you see when you look at a graphics card with the cooler off, a grey square with a smaller shiny rectangle in the middle, is the package with the die on top.

For a gaming GPU that is the whole story. The memory chips, the GDDR6X of Part 17, are separate packages soldered to the board around the GPU, and the wires between them run through the board, which is why they are limited to a few hundred bits wide and have to run very fast to compensate. Part 10 made the case that a data centre GPU wants the opposite: thousands of wires, running slower. Each HBM stack has 1,024 of them (2,048 for HBM4), and no board can route eight thousand tracks over a few millimetres. So the wiring has to be made the same way the chip was, with photolithography, on silicon.

That is what TSMC's CoWoS does, and the name spells out the layers: chip on wafer on substrate. A large, plain piece of silicon called an interposer, printed with a few coarse layers of wiring and nothing else, sits on the substrate. The GPU die and the HBM stacks sit on the interposer, and the tracks in the interposer connect them with the density only a wafer process can offer. The H100 uses this in its original form, one interposer of up to about three reticle fields in area, carrying the GPU and six HBM sites. Blackwell's two dies and eight stacks are larger than any interposer that could be made economically, so it uses a newer variant, CoWoS-L, in which small silicon bridges are embedded in an organic layer only where dies need to talk to each other. The 10 TB/s NV-HBI seam between the two Blackwell dies crosses one of those bridges. The interposer is also why the die-to-memory distance on an H100 is a few millimetres and why its bandwidth is three times a gaming card's: the numbers in Part 10 are, in the end, a packaging decision.

The HBM stack is packaging too. It is eight or twelve DRAM dies, each thinned to a fraction of a millimetre, stacked on top of a small logic die and connected vertically by thousands of holes drilled through the silicon and filled with copper, called through-silicon vias. That is how a memory device the size of a fingernail holds 24 or 36 GB and presents a thousand wires at its base. It is also why HBM is expensive and its supply is the number that decided how many Blackwells could be built: every stack is a small, difficult package in its own right, made by three companies in the world, and the GPU cannot ship without eight of them.

Do the arithmetic on a spec sheet and the packaging shows through. A B200 has 8 stacks × 8 dies × 3 GB = 192 GB. A B300 has the same eight stacks at twelve dies high, 288 GB. Rubin in Part 20 keeps 288 GB and doubles the wires per stack instead. "How much memory" has become "how many stacks, how many dies high, and how big is each die", and every one of those is a question about packaging rather than about the GPU.

Testing, and the card

Before the wafer is cut, every die on it is tested in place by lowering a bed of fine needles onto its pads and running it. That is where the binning happens: which SMs work, how fast the die runs at what voltage, which memory controllers are good. After dicing and packaging the finished part is tested again, then it goes to a board maker. A gaming card is a printed circuit board with the package in the middle, memory chips around it, the power stages of Part 21 along one edge, and the cooler on top. A data centre GPU is more often an SXM module, a bare board with the package and its on-package memory, designed to be plugged into a server tray beside seven others, with no fan and no display output, because the server supplies both the power and the cooling. Same die, same package, a different board, and a very different price.

What it all costs

A rough sense of the money helps make the rest of the series legible. A wafer on the 4N process that made Ada and Hopper was reported at somewhere between $16,000 and $20,000 by 2025, and 3 nm-class wafers a little more. About 90 AD102 candidates fit, and after yield and binning perhaps 60 to 70 ship as something. That is $250 to $300 of silicon in an RTX 4090 that sells for $1,600. For an H100, 65 candidates and maybe 40 sellable dies puts the GPU silicon at $400 to $500, before five HBM stacks, the interposer, and the package, in a product that sells for $25,000 to $30,000. The silicon is not where the price comes from. The price comes from the fact that the design cost billions, the fab that printed it cost twenty billion, the EUV machines inside it cost a couple of hundred million each, and the software of Part 12 has taken twenty years. A chip is the cheapest part of a chip.

Where this leaves us

You can now read the first few lines of any GPU spec sheet without a dictionary. Process is the recipe the wafer went through, with the label Part 3 warned you about. Die size is the area of one printed copy, and it fixes how many fit on a wafer and how many of those are clean. Transistors is the count in that area. SMs enabled out of SMs on the die is the yield story made visible. Anything above about 800 mm² is against the reticle limit, and anything larger than that is more than one die on a package. And the memory line, on a data centre part, is a description of what was glued beside the die on an interposer.

There is one more piece of background before the real chips, and it is the job the GPU was named for. The next part draws a frame.


Next: The Graphics Pipeline: what a frame is, why everything on screen is a triangle, the stages from vertices to pixels, why that job has exactly the shape that Part 8 described, and which parts of a modern GPU still do nothing else.