Chapter 5 · Inference engines
Quantization

A decode step is memory bound because it spends most of its time streaming the model’s weights through memory rather than doing arithmetic on them, as established in the memory-bound and compute-bound boundary. Shrinking those weights shrinks that streaming cost directly. GPTQ does this with : it is a one-shot, post-training method based on approximate second-order information that compresses the weights of a model, including ones with 175 billion parameters, down to 3 or 4 bits each with negligible accuracy loss, and without any retraining. The cost of doing so is the part worth internalizing. Quantizing a 175 billion parameter model takes approximately four GPU hours, which buys the first execution of a model that size inside a single GPU for generative inference. Compression this cheap is a deployment decision, not a research project.

Where the saving lands is the more interesting half, and GPTQ is unusually direct about it. Because the compression touches only the weights, the arithmetic itself still runs at higher precision: the low-bit weights have to be dequantized back up before each matrix multiply. The paper states the consequence as a limitation of its own method, that it does not provide speedups for the actual multiplications, for lack of hardware support for mixed-precision operands such as FP16 times INT4 on mainstream architectures. Its bespoke GPU kernels are therefore described as leveraging compression for faster memory loading, and the end-to-end speedups they report, around 3.25 times on an A100 and 4.5 times on the more cost-effective A6000, are entirely a bandwidth result. A method that speeds nothing up arithmetically and still delivers those multipliers is about as clean a demonstration of this chapter’s premise as the literature offers: the benefit concentrates in the memory-bound decode step, where fewer bytes have to move, rather than in the compute-bound prefill step.

SmoothQuant instead quantizes both operands of the matrix multiply, . That is harder than quantizing weights alone because, as the paper puts it, weights are easy to quantize while activations are not: activation values in large language models develop large outliers concentrated in a handful of channels, and forcing those outliers into an 8-bit range destroys the precision available to every other value. SmoothQuant’s fix is a mathematically equivalent transformation, applied offline before serving, that migrates this quantization difficulty from the activations into the weights, making both sides easy enough to quantize to INT8 together. Because both operands end up in low precision, the matrix multiply itself can run on faster low-precision hardware paths, not just save memory traffic. The paper reports up to 1.56 times speedup and 2 times memory reduction with negligible accuracy loss, and one result that says more than either: it enables serving a 530 billion parameter model within a single 8-GPU node, the first time a model over 500 billion parameters fits in one.

How much difficulty to migrate is a knob, not a constant, and it is the part of SmoothQuant most often left out. A per-channel smoothing factor built from the activation maximum alone would push all the difficulty onto the weights, whose outlier channels then quantize badly; building it from the weight maximum alone pushes everything back onto the activations, which is where the problem started. Both extremes degrade accuracy, so the paper interpolates between them with a migration strength alpha, setting the smoothing factor for a channel to max(|X|)^alpha / max(|W|)^(1 - alpha). At an alpha of 0.5 the weights and activations of a channel end up sharing a similar maximum, and therefore a similar share of the difficulty, which the paper finds is a well-balanced point for models such as the OPT and BLOOM families. Models whose activations are worse behaved need more migration, not less: GLM-130B has roughly 30 percent outliers, and the paper raises alpha to around 0.75 for cases like it. Alpha is the dial that decides which operand absorbs the pain, and it is the reason a single quantization recipe does not transfer across model families untouched.

Weight-only versus weight-and-activation quantization. Two dataflows into the same matrix multiply. Weight-only quantization shrinks what streams from memory but dequantizes back to FP16 before the multiply, so the arithmetic stays high precision. Quantizing activations too lets the multiply itself run in INT8, at the cost of surviving activation outliers.Illustrative numbers

AWQ sits in the same weight-only family as GPTQ but reads the problem through the activations. Its core observation is that not all weights in a model are equally important: protecting only 1 percent of can greatly reduce quantization error, and the way to find those salient channels is to look at the activation distribution, not at the weights themselves. Keeping that 1 percent at higher precision would leave a hardware-inefficient mixed-precision layout, so AWQ instead derives an equivalent transformation that scales up the salient channels before quantizing, with the scale determined by activation statistics collected offline. Because the method relies on no backpropagation and no reconstruction, it does not overfit its calibration set, and it generalizes across domains and modalities, including instruction-tuned and multimodal models.

The word doing the work in that description is activation-aware, and the paper earns it with a control the field usually skips. The standard way to decide which weights matter is their own magnitude or L2 norm, so AWQ tries all three selection rules on the same model at the same bit-width. Under INT3 quantization with a group size of 128, OPT-6.7B moves from a WikiText perplexity of 10.86 in FP16 to 23.54 under plain round-to-nearest. Keeping 1 percent of channels in FP16 chosen by activation magnitude recovers almost all of it, to 11.39. Choosing the same 1 percent by weight magnitude gives 22.37, and choosing them at random gives 24.23. Selecting salient weights by looking at the weights is barely distinguishable from not selecting at all, which the paper states plainly as a similar marginal improvement to random selection. That is what makes the prefix in activation-aware load-bearing rather than decorative: salience is a property of what flows through a channel, and the weights alone do not reveal it.

Low-bit weights only pay off if the kernels serving them are fast, which is why the AWQ paper ships with TinyChat, an inference framework built for 4-bit models. With kernel fusion and platform-aware weight packing, TinyChat runs more than 3 times faster than the Huggingface FP16 implementation on both desktop and mobile GPUs, and the compression is what lets a 70B Llama-2 model be deployed on a mobile GPU at all.

Where quantization happens in a workflow#

Every method above runs before serving rather than during it, and the AWQ workflow in vLLM’s documentation makes the ordering plain. Quantizing is a separate program whose output is a directory on disk:

quantizing Mistral-7B, from the vLLM AutoAWQ documentation
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer

model_path = "mistralai/Mistral-7B-Instruct-v0.2"
quant_path = "mistral-instruct-v0.2-awq"
quant_config = {"zero_point": True, "q_group_size": 128, "w_bit": 4, "version": "GEMM"}

# Load model
model = AutoAWQForCausalLM.from_pretrained(
    model_path,
    low_cpu_mem_usage=True,
    use_cache=False,
)
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)

# Quantize
model.quantize(tokenizer, quant_config=quant_config)

# Save quantized model
model.save_quantized(quant_path)
tokenizer.save_pretrained(quant_path)

The configuration dictionary is the whole specification of the method: four-bit weights, a group size of 128, an asymmetric range through a zero point, and a named kernel version. The one call that does the work is model.quantize, and it is handed the tokenizer rather than only the weights, which follows from where AWQ’s scales come from: activation statistics collected offline, and producing activations means running text through the model. Everything after that call is file output. On the serving side the difference is one keyword, quantization="auto_awq", added to the same LLM constructor this chapter opened with. SGLang draws the line even more sharply for FP8: pass --quantization fp8 against an FP16 checkpoint, or load an FP8 checkpoint and pass nothing at all.

Which method is available at all is a hardware question before it is a quality question. vLLM publishes a compatibility matrix across Volta, Turing, Ampere, Ada, Hopper, AMD GPU, Intel GPU, x86 CPU and Arm CPU, and the rows do not agree: GPTQ is marked supported on Volta while AWQ starts at Turing, both run on Intel GPU and x86 CPU but not on AMD GPU, and the Marlin kernels that serve GPTQ, AWQ, FP8 and FP4 formats are NVIDIA-only from Turing onward. A checkpoint quantized for one fleet is not automatically servable on another, and the fast kernel path for a format is narrower than the format’s own support.

The libraries move faster than the methods: vLLM’s AutoAWQ page now carries a deprecation warning, the library having been adopted into llm-compressor, which vLLM recommends instead for FP8, INT8, INT4 and other formats. TensorRT-LLM shortens the workflow further by serving checkpoints somebody else quantized, so that trtllm-serve "nvidia/Qwen3-8B-FP8" is the whole deployment, with one documented caution: confirm the GPU supports FP8 first. For a growing share of deployments quantization is not a step the operator runs, it is a property the checkpoint already has.