Treat the KV cache like virtual memory — pages, not contiguous blocks — and the GPU serves several times more concurrent requests.
A car park where every driver reserves twenty spaces in case they bring friends, and almost nobody does. Hand out one space at a time and the same car park holds far more cars.
It is why the same GPU serves several times the traffic — and GPU count is the entire serving bill.
Naive LLM serving reserves KV cache for each request's maximum possible length, contiguously. Since most generations stop far short of the maximum, the majority of reserved memory is never used — measured waste of 60-80%. PagedAttention borrows the operating-system solution: split the KV cache into fixed-size blocks, keep a per-sequence block table, and let a sequence's blocks live anywhere in GPU memory. Internal fragmentation drops to at most one block, sharing becomes trivial (a common prompt prefix is just blocks referenced by several sequences with copy-on-write), and the freed memory turns directly into concurrency. Continuous batching is the scheduler half: finished sequences leave the batch and waiting ones join every step, instead of the whole batch waiting for its slowest member.
PagedAttention stores the KV cache in fixed-size blocks with a per-sequence block table, exactly like OS virtual memory. That removes the huge over-reservation of contiguous allocation, caps waste at under one block per sequence, and lets sequences share prefix blocks by reference. The freed memory becomes batch size, and continuous batching keeps that batch full by admitting new requests as finished ones leave — together, several times the throughput on the same GPU.
What is vLLM? Efficient AI Inference for Large Language Models — IBM Technology, 4:57