Chapter 8 · Frontier
Watchlist

This book has an evidentiary standard: an architecture gets covered in the hardware chapter once its ISA documents are public and its performance numbers can be reproduced by someone who does not work for the vendor. Until then it stays here, as a set of claims worth tracking rather than facts worth teaching. Everything in this section is in that waiting room.

Most of that waiting room has very little on the record, which is exactly why it is a watchlist. AMD’s next datacenter GPU generation, session-aware scheduling in inference engines, serving real-time voice and video, dedicated inference ASICs, and processing in memory are all areas where the interesting claims currently live in announcements, papers without public artifacts, or products without disclosed baselines. For each, the same three questions apply: is the hardware or system shipping, are the specifications published, and has anyone outside the organization reproduced the numbers? Pointers to whatever public material exists are kept current in the reading list, where the evidence can accumulate.

NVIDIA Rubin#

The entry with the most material on it is NVIDIA’s Rubin GPU, and it is worth reading closely precisely because none of it can be checked yet. The launch post describes two reticle-limited compute dies joined by , 336 billion transistors, 224 streaming multiprocessors and 896 Tensor Cores, a third-generation Transformer Engine rated at up to 50 petaflops of NVFP4, and up to 288 GB of HBM4 delivering up to 22 TB/s of peak bandwidth, which it puts at 2.8x Blackwell and Blackwell Ultra. The headline is up to 10x more agentic throughput per unit of energy than Blackwell. Every one of those is a vendor claim, and the throughput figure is drawn from an internal 2T mixture-of-experts workload. None of that makes them false. It makes them unverified, and the distinction is the whole point of this chapter.

The more useful half of the post is the half a reader of this book could go and measure, because it names mechanisms rather than multipliers. The Tensor Memory Accelerator gains inline descriptor updates: rather than edit a descriptor in memory per expert, “the kernel can keep one unified descriptor for tensors that share the same layout and override fields such as the memory pointer and stride directly in the TMA instruction at runtime,” which the post frames as an MoE scaling argument. Dependent kernels get finer-grained triggering, so consumer work can begin “earlier as required input data becomes available, rather than waiting for a larger set of producer work to complete,” where Blackwell’s programmatic dependent launch still waits on a bulk dependency. Attention is claimed to combine “activation sparsity with adaptive compression and improved softmax throughput”: the intermediate scores from the dense QK-transpose are loaded out of Tensor Memory into “a structured 2:4 sparse compressed form” so that softmax and the second attention GEMM run on the nonzeros, with exponential throughput per clock per SM tabulated at 2x FP32 and 4x BF16/FP16 against a Blackwell baseline. Weights get a 3-bit lookup-table format for the B operand, storing an index into a small table of representative values that the Tensor Cores resolve inline.

That list is the reason to keep Rubin here rather than to wave it off. TMA descriptor handling, kernel launch and transition latency, sparse matrix-multiply throughput, and achieved memory bandwidth are all things the earlier chapters have already taught you to profile, so each of these claims names a mechanism with a consequence someone outside NVIDIA could measure. That is more than most of this section can say. What would move Rubin into the hardware chapter is unchanged: shipped systems, published tuning and ISA documentation, and those measurements actually taken.