Chapter 0 · Introduction
Why kernels matter
0.1

Why kernels matter

A modern accelerator is not one fast processor. It is a very wide machine with a deep memory hierarchy, and almost every disappointing performance number comes from the same root cause: the arithmetic units finished early and spent the rest of their time waiting for data. A kernel that reads its inputs badly can be an order of magnitude slower than one that reads them well, running the exact same arithmetic on the exact same chip.

The gap this book is about. The top bar is what the hardware can do at peak; the bottom bar is what naive code gets on the same chip. The lengths are not measurements, the shape is the point: a kernel that reads its inputs badly can be an order of magnitude slower than one that reads them well.Illustrative numbers

That is why this field has a distinctive shape. The interesting work is rarely inventing new mathematics. It is arranging a known computation so that the data it needs arrives in the right place, in the right order, in large enough pieces, while the arithmetic units stay busy. Tiling, coalescing, pipelining, and quantization are all answers to that one question.

The gap has a specific arithmetic behind it. Every processor has a ratio between how fast it can compute and how fast it can feed itself, and accelerators have pushed that ratio steadily in favor of compute, because arithmetic has grown cheap faster than memory bandwidth has. A computation therefore has to do enough work per byte it loads to keep the arithmetic units busy, and a great many of the computations inference actually runs do not. As Foundations works out, a transformer matmul on an H100 needs a batch of roughly 280 tokens before it is compute-bound at all. Below that the chip is waiting on memory no matter how well the kernel is written, and the work of the kernel is to make sure it waits as little as it can.

The same logic scales outward. A single kernel worries about the trip from global memory to registers. An inference engine worries about the trip from one request to a batch of them. A distributed system worries about the trip across an interconnect. The unit of analysis changes; the reasoning does not.

That is what makes this an engineering discipline rather than a collection of tricks. The same model on the same hardware can serve very different numbers of tokens per second depending on how its kernels are written, how its requests are batched, and how its weights and its cache are laid out in memory. None of those are decisions about the model. They are all decisions about data movement, and they are made by people who understand the machine underneath.