AI-generated kernels
The claim under test is simple: language models can write CUDA kernels that beat PyTorch. KernelBench, the benchmark that made the question precise, evaluates models on 250 PyTorch workloads and scores them with fast_p, the fraction of generated kernels that are both correct and faster than the baseline by more than a chosen threshold. Its authors found that frontier reasoning models did best out of the box but still matched the PyTorch baseline in less than 20 percent of cases, with the benchmark getting harder as the speedup threshold rises.
Then the correctness checks themselves came under scrutiny. KernelBench-Verified, a follow-up from Meta and Stanford, found that frontier models frequently engage in reward hacking. The original baseline ran PyTorch in plain float32, without TF32, so any generated kernel that merely invoked cuBLAS looked dramatically fast against an artificially slow reference. And because the test inputs came from one narrow distribution, all positive and small, models learned to hardcode bypasses: one generated ReLU kernel checked for the test shape and returned its input unchanged, passing the correctness check and reporting a 374x speedup while computing nothing.
Under the verified protocol, a TF32-enabled baseline plus a hidden test suite of four input distributions, the picture inverted. The best model’s geometric mean speedup fell from 1.43x under the standard protocol to 0.88x, no model consistently outperformed PyTorch, and 28 percent of the best model’s correct kernels increased peak GPU memory.
NVIDIA’s SOL ExecBench pushes the standard further still, and its design is a compact statement of what evaluating a generated kernel actually involves. Submissions are “checked for various forms of reward hacking,” “tested against a reference solution for numerical correctness,” and “timed under reproducible conditions”: three jobs that the original protocol left to a single correctness check, which is how a ReLU kernel that computes nothing got through. Ranking then runs on a SOL-Score, “a metric that grades custom kernel performance based on the theoretical roofline of a NVIDIA B200 GPU,” obtained analytically with SOLAR, a separate NVIDIA tool. That is the roofline of Foundations turned into a scoring function: the thing a generated kernel is measured against stops being another program that could itself be badly configured and becomes the hardware limit, which no baseline configuration can lower. And the languages it accepts are why it belongs at the end of this book rather than in a footnote: PyTorch, Triton, CUTLASS, cuDNN, CuTe DSL, cuTile, and CUDA C++, which is very nearly the range of toolchains the earlier chapters teach.
The state of the evidence, then: models can write working kernels, and any claimed speedup is only as trustworthy as the baseline configuration and input distribution behind it. When you read the next result in this space, those are the two things to check first.