How this book is organized
The book is ordered the way the problem is: from one inference request down to a single kernel, then back out to a fleet of machines. Prerequisites comes first and is the only chapter that assumes nothing: why a GPU is built the way it is, how to read the languages the listings use, and what a transformer actually computes. Foundations then covers the execution model and the compute-bound versus memory-bound boundary that everything later depends on. Kernel optimization applies it to the three computations that matter most in practice: matrix multiplication, low-precision arithmetic, and attention.
Programming models and profiling is about the tools you actually write and measure kernels with, since very little of this work is done in raw CUDA C++ any more. Inference engines and Distributed inference move up to the systems layer, where scheduling, batching, and placement decide throughput more than any single kernel does. Current hardware covers what the chips being deployed right now actually provide, read from their own specifications and tuning guides rather than from vendor peak numbers.
Frontier closes the book, and it is the one chapter held to a different standard on purpose. Everything before it rests on published documents and measurements that can be reproduced. Frontier collects what has not settled yet: the question of whether language models can write kernels worth deploying, the benchmarks that keep having to be redesigned once models learn to game them, and hardware that has been announced but not independently measured. It is there so that the rest of the book does not have to hedge.
Two appendices sit outside that order. The glossary collects the vocabulary, each entry quoted verbatim from the source that introduced it. The reading list collects the primary sources: papers, specifications, and repositories, grouped the same way the chapters are.