All concepts

PagedAttention & Continuous Batching

Treat the KV cache like virtual memory — pages, not contiguous blocks — and the GPU serves several times more concurrent requests.

Advanced LLM Systems · Advanced · ~7 min

In plain English

A car park where every driver reserves twenty spaces in case they bring friends, and almost nobody does. Hand out one space at a time and the same car park holds far more cars.

Why it's worth your time

It is why the same GPU serves several times the traffic — and GPU count is the entire serving bill.

If you remember three things

  • KV cache as pages with a block table, like OS virtual memory
  • Waste falls from 'max length' to 'under one block'
  • Shared prefixes are shared blocks — that's prefix caching

Overview

Naive LLM serving reserves KV cache for each request's maximum possible length, contiguously. Since most generations stop far short of the maximum, the majority of reserved memory is never used — measured waste of 60-80%. PagedAttention borrows the operating-system solution: split the KV cache into fixed-size blocks, keep a per-sequence block table, and let a sequence's blocks live anywhere in GPU memory. Internal fragmentation drops to at most one block, sharing becomes trivial (a common prompt prefix is just blocks referenced by several sequences with copy-on-write), and the freed memory turns directly into concurrency. Continuous batching is the scheduler half: finished sequences leave the batch and waiting ones join every step, instead of the whole batch waiting for its slowest member.

In an interview

PagedAttention stores the KV cache in fixed-size blocks with a per-sequence block table, exactly like OS virtual memory. That removes the huge over-reservation of contiguous allocation, caps waste at under one block per sequence, and lets sequences share prefix blocks by reference. The freed memory becomes batch size, and continuous batching keeps that batch full by admitting new requests as finished ones leave — together, several times the throughput on the same GPU.

Production defaults

gpu_memory_utilization
0.90-0.95. Paging makes a high value safe; leaving it at 0.7 is throwing away concurrency
max_model_len
set from the p99 of real traffic, not the model's maximum. Every unused token is block pool you didn't get
Prefix caching
on, whenever a system prompt is shared. Usually the single largest throughput win available

What breaks

  • Throughput far below expectations — max_model_len is set to the theoretical maximum and starving the block pool. Size it from real traffic.
  • Requests keep being preempted — Sustained preemption means the pool is too small for the concurrency. Lower max_num_seqs or raise memory utilisation.

Watch it explained

What is vLLM? Efficient AI Inference for Large Language Models — IBM Technology, 4:57

Related