Look at the sticker on a new laptop and you will find a line that did not exist three years ago: "NPU, 50 TOPS". Turn over an iPhone's spec sheet and there is a "16-core Neural Engine". Neither of these is a GPU, and yet both exist to run w × x + b in bulk, which is what this whole series has said a GPU is for. So what are they? And how can a chip that draws five watts advertise 50 TOPS when an H100 draws 700 W for 1,979 TFLOPS, which sounds like only forty times more for a hundred and forty times the power?
This part answers both questions, and in doing so it finally steps outside Nvidia. Everything from Part 3 to Part 21 has been about one company's GPUs, because that is where the ideas are cleanest and where the models you use are trained. But the machine that runs a neural network in your pocket is built on a different bet from the one in Part 8, and so is the machine Google trains on. Seeing where the GPU's bet stops paying is the best way to understand what the bet was.
The chip in your phone is not a card
Start with the most basic difference. A graphics card is a separate board with its own memory, plugged into a slot, talking to the CPU over PCIe, the slow road of Part 12. A phone has no slots. Everything sits on one piece of silicon, a system-on-chip (SoC): the CPU cores, a GPU, the NPU, the video and camera hardware, the radio, and one memory controller that they all share. The memory itself, a few LPDDR chips (the low-power cousin of the DRAM in Part 7), is soldered next to the SoC or stacked on top of it, and there is one pool of it. Apple calls this unified memory, and the idea is now standard: the CPU, the GPU and the NPU all read and write the same gigabytes, so a picture the camera writes can be read by the NPU and then by the GPU without ever being copied across a bus. The cudaMemcpy of Part 12 has no equivalent, because there is nowhere to copy to.
The price is bandwidth. A soldered pool of LPDDR delivers something like 100 to 230 GB/s (an Apple M5 quotes 153 GB/s, a Snapdragon X2 Elite up to 228), against 1 TB/s for the GDDR on a gaming card and 3 to 22 TB/s for the HBM on a data centre GPU. Hold that number, because it decides most of what follows.
The GPU inside the SoC is a real GPU, the same bet as Part 9 at a hundredth the scale: a handful of cores, each with a scheduler feeding a few dozen lanes in lockstep, built to colour pixels and increasingly to do matrix arithmetic. Apple, Arm (Mali), Qualcomm (Adreno), Intel (Xe) and AMD (RDNA) each design their own, and none of them speaks CUDA. But the SoC also carries a third arithmetic engine, and that one is new.
What an NPU is for
A neural processing unit is a block of the SoC that does one thing: run a trained neural network, forward only, as cheaply as possible. Not train one. Not draw anything. Take a network whose weights are already known, feed it an input, and produce an output, thousands of times an hour, for as long as the battery lasts.
Look at what a phone actually does with a neural network. It blurs the background of a video call, keeps the camera's portrait mode separating you from the wall, transcribes speech into captions, decides which photos have your dog in them, unlocks itself by looking at your face, corrects your typing, and, since 2024, answers questions with a small language model. Every one of those runs a fixed network, most of them run all day, and none of them can be allowed to drain the battery or warm your hand. That is a completely different requirement from a data centre GPU's. The GPU must be programmable, because nobody knows what kernel it will be asked to run next. The NPU only has to do w × x + b by the billion, at 8-bit precision, on networks small enough to fit in a few hundred megabytes, and it has to do it at a few watts.
Apple shipped the first one in a phone in 2017, the Neural Engine in the A11 chip, rated at 600 billion operations a second. Huawei's Kirin 970 arrived the same year with an NPU of its own, and within two years every phone chip had one. Laptops followed in 2024, when Microsoft announced that a PC could carry the "Copilot+" label only if it had an NPU of at least 40 TOPS, which made the unit standard on Windows machines overnight. The 2026 chips carry a good deal more than that, and the table further down lists them.
How it works: the systolic array
Here is the interesting part, because an NPU is not a small GPU. It has no warps, no warp scheduler, no branch divergence, and usually no cache. It is closer to the tensor core of Part 11 with everything else removed, and the tensor core's idea taken to its logical end.
Recall the tensor core's trick: deliver each element of a matrix tile once and broadcast it by wire to every multiplier that needs it, so the register-file traffic that limits ordinary lanes disappears. Now make that the whole chip. Lay out a square grid of multiply-add cells, say 32 by 32 or 256 by 256. Load the weights into the cells and leave them there. Then push the inputs in from the left edge, one row per step, so that each input value marches across its row of cells, being multiplied by a different weight in each, while the running sums march down each column and drop out at the bottom as finished outputs. Nothing is fetched twice. No cell ever asks for a number. The numbers arrive because the cell next door passed them on. The design is called a systolic array, from the rhythmic pumping of a heart, and it dates from a 1982 paper by H. T. Kung, thirty years before anyone needed it for this.
Try it: a systolic array, one step at a time
Below is a three-by-three array computing a layer with three inputs and three neurons: a weight matrix W times an input matrix X (three inputs, three examples at once). The weights are already loaded into the cells and never move. Press Step and watch the inputs enter from the left and the partial sums travel down, one cell per step. The counters underneath are the point: every number is read from memory exactly once, and every multiply-add happens without a single fetch.
Inputs X (flow in from the left)
Weights W (held in the cells)
Each cell shows the input passing through it (top), its fixed weight (middle), and the partial sum it has just passed down (bottom). Finished outputs collect below the grid.
Look at the counters when it finishes. Twenty-seven multiply-adds, eighteen numbers fetched, and the eighteen were fetched before and during the first few steps, never in the middle of the arithmetic. Now put the energy table from Part 21 next to that. An 8-bit multiply costs about 0.2 picojoules, while fetching its operand costs 5 pJ from a small on-chip memory and 640 pJ from DRAM. A design that fetches each number once and then reuses it hundreds of times by construction spends almost all of its energy on arithmetic, which is the whole reason a 50 TOPS engine fits in a five-watt budget. It has thrown away the machinery a GPU keeps for flexibility (instruction fetch, schedulers, register files, caches, the SIMT model of Part 8) and kept only the wires.
The trade is real. A systolic array is superb at dense matrix arithmetic of a shape it was built for and useless at anything else. A network with an odd layer, an unusual operation, or a tensor that does not tile onto the grid falls back to the CPU or GPU, and it is the job of a compiler (Apple's Core ML, Microsoft's Windows ML, Intel's OpenVINO, Qualcomm's AI Engine Direct) to carve a model into pieces the NPU can eat and pieces it cannot. That is the practical reason NPUs are used for the fixed, predictable workloads listed above and not, so far, for whatever a developer feels like running.
TOPS, and its asterisks
The unit on the laptop sticker needs unpacking, because it is the source of most confused comparisons. TOPS is trillions of operations per second, where a multiply and an add count as two operations, exactly the convention behind FLOPS. The difference is the word floating. NPU figures are almost always quoted for 8-bit integers (INT8), the cheapest arithmetic in the Horowitz table, and sometimes for INT4, and the vendor does not always say which. An H100's 1,979 TFLOPS of dense FP8 is, in the same units, 1,979 TOPS at 8 bits. So a 50 TOPS NPU has about a fortieth of the arithmetic of an H100, at about a hundredth of the power, which makes it two or three times more efficient per operation, by design, for the workloads it can run. The gap in absolute throughput is real, and the gap in memory bandwidth, twenty to a hundred times, is larger still.
Two asterisks, both familiar from Part 11. Some TOPS figures are quoted with structured sparsity, which doubles them. And Nvidia's "AI TOPS" figure for the RTX 5090 is an FP4 number with sparsity, which is why it reads as 3,352 against the card's roughly 830 TFLOPS of dense FP8. Whenever you see TOPS, ask three questions: how many bits, dense or sparse, and at what power.
The chips of 2026
| Chip family | NPU | Rated at | Where it lives | Notes |
|---|---|---|---|---|
| Apple M4 | 16-core Neural Engine | 38 TOPS | MacBook, iPad | Apple's figure. The 2025 M5 adds a "neural accelerator" inside each GPU core as well, so its GPU does matrix work too, and Apple has not quoted a TOPS figure for its Neural Engine |
| Apple A17 Pro and later | 16-core Neural Engine | 35 TOPS | iPhone | the Apple Intelligence models run here |
| Qualcomm Snapdragon X2 Elite | Hexagon | 80 TOPS | Windows laptops | up from 45 TOPS on the 2024 X Elite |
| Intel Core Ultra Series 3 (Panther Lake) | NPU 5 | 50 TOPS | Windows laptops | Intel's first chips on its 18A process, announced January 2026 |
| AMD Ryzen AI 300 | XDNA 2 | 50 TOPS | Windows laptops | |
| Microsoft Copilot+ requirement | any | 40 TOPS minimum | the label on the box | set in 2024 |
Every figure is the vendor's, at INT8 unless stated, and none of them has been through the kind of independent measurement that GPUs get. Read the table as a statement of what the 2026 sticker says, and note the shape of it: every laptop and phone processor now has one of these, from four companies with four different designs and four different software stacks, none compatible with the others. That is the opposite of the CUDA story in Part 12, and it is why most software still does not use the NPU it has.
Where the NPU stops
If the NPU is so efficient, why does the data centre not run on them? Three reasons, and the roofline of Part 10 gives all three.
Bandwidth. Generating text is memory-bound: every token reads every weight. An 8-billion-parameter model at 4 bits is about 4.5 GB. Through 150 GB/s of LPDDR that is a ceiling of around 30 tokens a second, however many TOPS the NPU has, and the NPU is usually not even the fastest reader of that memory. Through an H100's 3.35 TB/s the same model could be swept 700 times a second. The NPU's arithmetic is fine. Its road to memory is a lane where the GPU has a motorway, and for large models that is the only number that matters.
Capacity. A phone has 8 to 16 GB of memory for everything, and a laptop 16 to 64. A 70-billion-parameter model does not fit at any precision, and the frontier models are several times larger again. The NPU runs the small models that are made for it, and those are getting good, but the models you talk to in a browser live on HBM.
Training. Training needs the backward pass of ML Basics, which doubles the arithmetic and, worse, needs 16-bit floating point and enormous batches held in memory. A fixed INT8 array with a few megabytes of local SRAM cannot do it. Every NPU is a machine for the forward pass only.
So the NPU and the GPU sit at opposite ends of the same roofline. The NPU takes the small, known, memory-light workloads, runs them for pennies of energy, and hands everything else off. The GPU takes everything, at a hundred times the power. That division is stable, and it is the reason a laptop in 2026 has both.
The same idea at rack scale: the TPU
There is one more twist, and it is the one that matters for the data centre. The systolic array does not have to be small. Google built one for its data centres in 2015, called it the Tensor Processing Unit (TPU), and described it two years later in a paper that is the clearest statement of the NPU argument ever written. The first TPU was a 256 by 256 systolic array, 65,536 8-bit multiply-add cells, with 28 MiB of software-managed on-chip memory and no cache, on a card that fitted in a hard-drive slot. It reached 92 TOPS, ran fifteen to thirty times faster than the Nvidia K80 GPU of the day on Google's own inference workloads, and did so at thirty to eighty times the operations per watt. The paper's explanation is the one this part has been building: the TPU's "deterministic execution model" simply had none of the machinery, caches, out-of-order logic, warp scheduling, that a general-purpose processor spends its transistors and joules on.
Ten years on, the TPU has grown into a full competitor to the GPU for training as well as inference, with the same additions Nvidia made, HBM and a fast interconnect, bolted around the array. Google's seventh generation, Ironwood (2025), quotes 4,614 TFLOPS of FP8 per chip, 192 GB of HBM at 7.4 TB/s, an interconnect at 1.2 TB/s, and pods of 9,216 chips reaching 42.5 exaflops, with twice the performance per watt of the generation before. Put that beside the B200 column in Part 19's table and the numbers are of the same order, on a chip with no lanes, no warps, and no CUDA. Google's own models run on it, other labs have contracted for it at scale, and it is the strongest evidence that the GPU's design bet, general lanes with a matrix unit attached, is not the only one that works at the top end.
It is not alone. Amazon's Trainium, Cerebras's wafer-sized chip, Groq's LPU, and a growing list of startups are all variations on "a matrix engine, lots of on-chip memory, and as little else as possible". And AMD's MI300 and MI355X, and Intel's Gaudi, are GPU-shaped competitors to Nvidia's own, which this series will come to in a later part.
Which chip, for which job
| Job | Best fit | Why, in this series' terms |
|---|---|---|
| Blurring a video call, captions, camera tricks, autocorrect | NPU | small fixed network, all day, on a battery, the systolic array's home ground |
| A small language model on a laptop, one user | NPU or the SoC's GPU | memory-bound at 150 to 230 GB/s either way, so the NPU wins on battery and the GPU on flexibility |
| A large language model, one user, at home | Discrete GPU (RTX class) | capacity and bandwidth: 24 to 32 GB at 1 to 1.8 TB/s, Part 17 |
| Serving a model to many users | Data centre GPU or TPU | batching makes it compute-bound, and only HBM plus a fast interconnect keep up, Part 10 and 13 |
| Training a frontier model | Data centre GPU or TPU, tens of thousands of them | backward pass, 16-bit formats, terabytes of HBM, and the interconnect ladder of Part 13 |
| Graphics | GPU, integrated or discrete | the job it was named for, and nothing else has the pipeline |
Where this leaves us
An NPU is the tensor core with the GPU taken away: a fixed grid of multiply-add cells through which numbers flow, each fetched once, so that nearly every joule goes on arithmetic rather than on moving bytes. That is why it can run a network at a few watts and why it cannot run a large one at all, because the memory beside it is a lane where the GPU has a motorway. The TPU is the same grid at data centre size with HBM and an interconnect around it, and its existence shows that the GPU's particular bet, general lanes first and a matrix unit second, is one answer to w × x + b at scale rather than the only one.
You now have the whole landscape: the switch, the gate, the lane, the SM, the die, the package, the rack, the site, the watts that run it, and the two other shapes of chip that do the same calculation on different terms. The last part puts it all in one picture.
Next: The Whole Picture: the full stack from transistor to rack in one diagram, a guide to every line on a spec sheet, the three questions to ask of any chip, and a glossary to keep.