Fifteen parts in, the thing a graphics processor was built to do has appeared only in passing. Part 8 used "colouring the pixels of a frame" as the extreme case of independent work and said that graphics is why GPUs exist. Part 14 said that before 2006 a graphics chip was a pipeline of specialised stages and that the G80 replaced them with a pool of general lanes. Neither part showed you a frame. This one does, and for three reasons. The pipeline is where the SM's shape came from, so it explains the machine. A good part of an Ada die is units that do nothing but draw, so it explains the spec sheet. And the data centre chips in the next parts are defined partly by what they leave out, which only makes sense if you know what a gaming chip puts in.

This is not a course in 3D programming. It is one frame, followed from a list of corners to the screen, with the words you will meet on a spec sheet defined as they appear: frame, vertex, triangle, shader, rasteriser, texture unit, ROP.

A frame is a grid of numbers

A screen is a grid of dots called pixels, each showing one colour, and a colour is three numbers, the amount of red, green, and blue, usually a byte each. A 1080p screen is 1,920 by 1,080, which is 2.07 million pixels. A 4K screen is 3,840 by 2,160, which is 8.3 million. A frame is one complete set of those numbers, and a game produces a new one every time the screen refreshes: 60 times a second on an ordinary monitor, 120 or 240 on a gaming one. At 4K and 120 Hz that is a billion pixels a second, each of which needs its own colour worked out, with a hard deadline of eight milliseconds per frame. Miss it and the picture stutters.

The frame lives in memory, in a block called the frame buffer: 8.3 million pixels times four bytes is 33 MB, written once per frame at least, then read out by the display engine and sent down the cable. Keep that in mind, because the memory traffic of a frame turns out to be the reason a gaming card's memory looks the way it does.

Everything is a triangle

The scene behind the frame is not made of pixels. It is made of surfaces, and every surface, however curved it looks, is a mesh of flat triangles. A character in a modern game is a few hundred thousand of them, and a frame can contain millions. Triangles are the universal shape for a reason a hardware designer will appreciate: three points always lie in a plane, so a triangle is always flat, it is always convex, any polygon can be cut into triangles, and a machine that only has to draw one kind of shape can be built to draw it very fast.

Each corner of a triangle is a vertex, and a vertex is a handful of numbers: a position in three dimensions, a direction that says which way the surface faces there, called the normal, and a pair of coordinates saying which point of a picture should be pasted onto the surface at that corner. The picture is a texture, and textures are how a flat triangle comes to look like brick, skin, or rust. A scene, to the GPU, is a long list of vertices and a list of which threes make triangles. Everything from here on is about turning that list into pixels.

Six stages

1 Vertex shadermove and project 2 Assemble, culldrop what can't be seen 3 Rasterisewhich pixels? 4 Depth testis it in front? 5 Pixel shaderwhat colour? 6 Outputwrite the pixel SM lanes, a program per vertex fixed-function units, one per GPC SM lanes, a program per pixel ROPs Red stages run as programs on the lanes of Part 9. Grey stages are dedicated hardware. Millions of items flow through every frame.
The pipeline as it has stood since about 2001, with the programmable stages in red. Modern APIs add optional stages between 1 and 2 for tessellation and geometry, and a compute path that skips the pipeline entirely, which is how the AI work gets in.

Vertex shading. Every vertex arrives in the coordinates of its own model, a character standing at the origin facing forward. The first stage moves it into the world, then into the camera's view, then projects it onto the flat screen so that farther things look smaller. Each of those moves is a matrix times the position, and a 4 by 4 matrix times a 4-element vector is sixteen multiply-adds, the w × x + b of this whole series, done once per vertex with no vertex depending on any other. A frame with five million vertices is five million copies of the same short program, and that is the first place the shape of Part 8 shows up. It is a shader, a small program that runs once per item, and it runs on the lanes of Part 9.

Assembly and culling. The projected vertices are gathered back into triangles, and anything the camera cannot see is dropped: triangles behind the camera, triangles off the edge of the screen, which get clipped, and triangles facing away. A closed object like a cube has close to half its faces pointing away from the viewer, hidden by the near ones, and dropping them before any pixel work is done cuts the job nearly in half. That test is a single sign check on the triangle's winding on screen, and it happens in fixed hardware.

Rasterisation. Now the question changes from "where is this triangle" to "which pixels does it cover". The rasteriser walks the pixels inside the triangle's bounding box and, for each, tests the pixel's centre against the three edges. Each edge test is one more w × x + b, and a pixel is inside when all three agree. For every covered pixel it emits a fragment: the pixel's position, plus the vertex attributes blended across the triangle, so a fragment near one corner gets mostly that corner's texture coordinate and normal. This is the stage that turns geometry into pixel work, and it is where the counts explode. A few million triangles become tens of millions of fragments. On the AD102 die of the next part there is one raster engine per GPC, twelve on the full die, and they are dedicated hardware because the job never changes.

Depth test. Two triangles can cover the same pixel, a wall and the character in front of it. To settle which one wins, the GPU keeps a second grid beside the frame buffer, the depth buffer, with one number per pixel: the distance of whatever is drawn there so far. A fragment farther away than the stored depth is thrown away. A nearer one replaces it. This is how a nearer surface hides a farther one without anyone sorting the scene, and it is also why the same pixel can get shaded several times in a frame, a cost called overdraw.

Pixel shading. Every surviving fragment now needs a colour, and this is the second program. Sample the texture at the fragment's coordinates, which means fetching four or eight nearby texels and blending them, a job so regular and so memory-heavy that each SM has four dedicated texture units to do it. Take the surface normal and the direction to the light and compute their dot product: a surface facing the light is bright, a surface edge-on is dark, and that one multiply-add per pixel is most of what makes a rendered object look solid. Add a highlight, a shadow lookup, a reflection, a fog term. A pixel shader in a modern game runs hundreds of instructions, and it runs once for each of tens of millions of fragments, the same code on different inputs, which is the warp of Part 9 exactly.

Output. The finished colour is handed to a render output unit (ROP), which does the depth comparison for real, blends the new colour with the old one if the surface is partly transparent, and writes the result into the frame buffer. An RTX 4090 has 176 ROPs and can write about 440 billion pixels a second, which sounds absurd against a billion pixels of screen until you remember the overdraw. When the frame is complete, the display engine reads the buffer out and the pipeline starts again on the next one.

Try it: one frame, stage by stage

Below is a rasteriser small enough to watch. Two cubes, sixteen vertices, twenty-four triangles, drawn onto a screen of 48 by 36 pixels, which is what a real frame looks like if you zoom in far enough. Press Step to advance one stage at a time and watch the counters. Press Run and the pipeline keeps cycling while the cubes turn.

Watch the counters through one frame. Sixteen vertices become twenty-four triangles, about half are dropped unseen, and the dozen that remain turn into around eight hundred fragments, of which a hundred or so lose the depth test where the front cube overlaps the back one. Scale the screen up from 48 by 36 to 3,840 by 2,160 and the fragment count goes up by five thousand times while the vertex count does not change at all. That is the whole reason the pixel stage got the most hardware: the work multiplies as it moves down the pipeline.

Why this is the shape of Part 8

Look at what every stage has in common. Millions of vertices, then millions of triangles, then tens of millions of fragments, and at each stage the same short program is applied to each item without reference to any other item. There is no chain. The only place two items ever meet is the depth test, and that is settled by a fixed unit in a fixed order. A frame is, top to bottom, the embarrassingly parallel job that Part 8 described, and the GPU's design follows from it: many slow lanes rather than one fast one, because the deadline is on the frame, not on any single pixel.

The history in Part 14 reads differently once you have the stages in front of you. The GeForce 256 of 1999, which Nvidia advertised as the first GPU, was defined by its maker as "a single-chip processor with integrated transform, lighting, triangle setup/clipping, and rendering engines" that could handle ten million polygons a second: one dedicated circuit per stage, in a fixed order, with the vertex work done on the card for the first time instead of on the CPU. The GeForce 3 in 2001 made the vertex and pixel stages programmable, and shaders were born. That created a new problem. A frame heavy in geometry left the pixel units idle, and a frame heavy in pixels left the vertex units idle, and no chip designer could know in advance which kind of frame a game would draw. The G80's unified design in 2006 was the answer: one pool of general lanes, and a scheduler that hands them vertices when there are vertices and fragments when there are fragments. That pool is the SM. CUDA, three months later, was the observation that if the lanes can run any program on any item, the items do not have to be pixels.

What still does nothing but draw

Not everything moved into the pool. Some stages are so regular, and so heavy on memory traffic, that a dedicated circuit does them several times more efficiently than a program on the lanes could, and they have stayed in hardware. On the full AD102 die of Part 17 they are: twelve raster engines, one per GPC, 192 ROPs, 576 texture units, four per SM, 144 RT cores, a display engine that drives up to four monitors, and the video encoders and decoders that a streamer or a video call leans on. Add the two programmable stages and this is the list of what a graphics chip does, and it is the list that the H100 in Part 18 will delete almost entirely: no RT cores, no display, no encoders, with the same silicon spent on memory interfaces and matrix units instead. When Part 18 says the two chips spent nearly identical transistor budgets on different things, this part is the list of things.

Two additions have changed the pipeline in the last few years, and they get their proper treatment in Part 17 because Ada is where they matured. Ray tracing turns the pipeline's question around: instead of asking which pixels each triangle covers, it asks, for each pixel, what the eye would see along a line through it, and follows that line as it bounces. The RT core is a unit for the searching that involves. DLSS renders fewer pixels than the screen has and lets a neural network on the tensor cores fill in the rest, and in its latest form invents whole frames between the rendered ones. A modern frame is a hybrid: rasterised for the most part, ray-traced for reflections and shadows, and inferred for a good fraction of its pixels. The AI hardware on a gaming die is not a spare part. It is now in the pipeline.

The memory wall, seen from a frame

Run the numbers from the first section through the stages and the gaming card's memory system explains itself. A 4K frame buffer is 33 MB of colour and another 33 MB of depth. With overdraw of two or three, every pixel is depth-tested and written more than once, so a frame at 120 Hz moves tens of gigabytes a second through the ROPs before a single texture is read. Textures are the larger cost: a game carries gigabytes of them, and every fragment samples several, four or eight texels each, which is why the texture units exist and why the L2 of Part 17 grew sixteen times in one generation, to keep the most-used texels on the die. Add the vertex streams and the buffers that shadows and reflections are drawn into, and a frame is a memory workload before it is an arithmetic one. That is the same conclusion Part 7 reached about neural networks, arrived at from the other direction, and it is why a gaming card ships with a terabyte a second of GDDR6X.

Where this leaves us

A GPU was named for a job with a particular shape: a very large number of small, identical, independent calculations, repeated to a deadline, and heavier on memory than on arithmetic. Everything in the first thirteen parts, the lanes, the warps, the schedulers, the caches, the wide memory, is the machine that shape produced. The accident of the last decade is that neural networks have the same shape, which Part 8 said in the abstract and this part has now shown in the concrete. The next four parts are the real chips, and the tables in them will make a different kind of sense: the difference between a gaming die and a data centre die is, line by line, the list in this part, kept or deleted.


Next: Ada Lovelace: the RTX 40 series, 76 billion transistors on a gaming chip, a cache sixteen times bigger than its predecessor's, and the moment a graphics card's most interesting new parts were the AI ones.