Chapter 5 · Inference engines
Long context and multimodal inference
5.6

Long context and multimodal inference

Everything so far assumed a request fits on the devices serving it. Long inputs break that assumption twice: the KV cache for a single very long sequence can outgrow any one device, and prefill compute grows quadratically with input length. The memory side is exactly what the Ring Attention paper targets, motivated by videos, actions, and other long-form sequences and modalities whose token counts overwhelm a single accelerator. distributes blockwise computation of self-attention and feedforward across devices while fully overlapping the communication of key-value blocks with the computation of blockwise attention. The result is exact, not an approximation, and it scales the feasible sequence length by up to the number of devices in the ring, which the paper demonstrates at context sizes in the millions of tokens.

Ring attention. Four devices in a ring, the sequence split into blocks across them. Key-value blocks circulate device to device while each device keeps its own query block and computes blockwise attention against whichever KV block just arrived, so the communication is fully overlapped with compute and the feasible sequence length scales by up to the number of devices in the ring.Illustrative numbers

The compute side of long context is a prefill problem. The MInference paper measures the cost concretely: because attention is quadratic in the prompt length, an 8B parameter model takes 30 minutes to prefill a 1 million token prompt on a single A100. Its observation is that long-context attention matrices are not densely important; they exhibit three recurring structures, called A-shape, Vertical-Slash, and Block-Sparse. exploits this by assigning each attention head its best-fitting pattern offline, then building the sparse indices for that pattern dynamically at inference time and running optimized sparse kernels over them. Applied to existing models with no change to pre-training and no fine-tuning, this cuts prefill latency by up to 10 times on an A100 while maintaining accuracy across long-context benchmarks. Between them, the two techniques bracket the long-context problem: one spreads exact attention over more hardware, the other spends less compute per unit of hardware, and both leave the model’s output distribution intact enough to serve the same requests.

MInference recovers sparsity from a model trained with full attention; Native Sparse Attention, from DeepSeek, builds the sparsity in from the start. uses a dynamic hierarchical strategy that combines coarse-grained token compression, which preserves global context awareness, with fine-grained token selection, which preserves local precision. The design is hardware-aligned, balancing arithmetic intensity so the sparse kernels actually run fast on modern GPUs, and because it is trainable end to end it reduces pretraining computation as well as inference cost. A model pretrained with NSA maintains or exceeds its full-attention counterpart across general benchmarks, long-context tasks, and instruction-based reasoning, while reaching substantial speedups over full attention on 64k-length sequences across decoding, forward propagation, and backward propagation.

Multimodal requests in the same queue#

Multimodal requests arrive through the same serving stack. In vLLM a multimodal model takes its text prompt together with a separate multimodal data dictionary carrying the other modalities, and the documented input types cover images, video, and audio, with images passed as URLs, as image objects, or as pre-computed embeddings.

an image alongside a prompt, from the vLLM multimodal inputs documentation
from vllm import LLM

llm = LLM(model="llava-hf/llava-1.5-7b-hf")

# Refer to the HuggingFace repo for the correct format to use
prompt = "USER: <image>\nWhat is the content of this image?\nASSISTANT:"

# Load the image using PIL.Image
image = PIL.Image.open(...)

# Single prompt inference
outputs = llm.generate({
    "prompt": prompt,
    "multi_modal_data": {"image": image},
})

The prompt is still a string, and the image is not in it. A placeholder sits in the text where the model’s own Hugging Face format expects it, and the pixels travel beside the prompt under multi_modal_data. Everything downstream is machinery this chapter has already described: the request joins the same waiting queue, draws blocks from the same paged cache, and is admitted by the same iteration-level scheduler as a text-only one.

Where the extra input does perturb the scheduler is chunked prefill. vLLM carries a flag for the case where a chunk boundary would fall inside an image: with partial scheduling of a multimodal item disabled, a prompt of text tokens followed by image tokens runs as the text in one step and the whole image in the next, rather than the image being split across two. Sarathi-Serve’s near equal sized chunks and a multimodal item’s indivisibility are in direct tension, and the resolution is a switch. Prefix caching has to account for the extra inputs too: multimodal input hashes are one of the values the block hash includes beyond the tokens, so two requests carrying the same image behind the same text prefix land on the same blocks while two carrying different images do not.

Accepting media by URL also opens a network path a text-only engine never had. vLLM recommends restricting the domains the server may reach, warning that otherwise it can be pointed at arbitrary endpoints and used for server-side request forgery, which matters most when the engine runs inside a cluster with access to internal networks. SGLang has the same allowlist, checks the initial URL and every redirect destination against it, and caps remote media downloads at 64 MiB by default. Long context turned the KV cache into a distributed systems problem; multimodal input turns the front of the engine into an HTTP client.