Inside the KV Cache: The Life of a Gigabyte

From gpu_memory_utilization and max_model_len to weights, activations, CUDA graphs and the KV pool: where every gigabyte on the card actually goes.

In this post, I’ll follow the memory on a GPU running vLLM [1], gigabyte by gigabyte, from the moment vllm serve starts to the KV cache pool that ends up serving your requests. In particular, I’ll show how vLLM works out the size of that pool, and why the flags everyone tunes by feel (gpu_memory_utilization, max_model_len and max_num_seqs) all act on that one number.

It starts with what the KV cache is and then layers in detail, one term at a time, so you can build an accurate mental model of where the memory goes before we do any serious arithmetic.

This post is structured into five parts:

  1. Background: what the KV cache is, how vLLM hands it out in blocks, and why the pool’s size is the server’s capacity
  2. Startup: how vLLM sizes the pool, step by step, and what each flag changes along the way
  3. KV bytes per token: the number that turns gigabytes into tokens
  4. Scaling up: from one RTX 4090 to two H100s, and a calculator for your own setup
  5. Failure modes: five common OOMs, each traced back to the step that causes it

Background: the KV cache

We’ll use the following command as our running example, on a single RTX 4090:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --gpu-memory-utilization 0.9 \
  --max-num-batched-tokens 8192

This configuration is:

From here, we’ll gradually build up to a 70B model on two H100s. The hardware and the numbers get bigger, but the method never changes.

Before this command can serve a single request, vLLM has to decide how much of the card goes to the KV cache. So let’s start with what the KV cache is.

What the KV cache is

A transformer [3] generates text one token at a time. To produce each new token, every attention layer compares that token against all the tokens before it, and for that it needs each earlier token’s key and value vectors, at every layer.

Those vectors don’t change once they’re computed, so recomputing them on every step would redo the whole prompt for every token generated. Instead, the engine keeps them in GPU memory. That store is the KV cache, and it works in two phases:

So the cache grows by a fixed number of bytes per token, and the memory a server needs at any moment is the total number of tokens in flight across all of its requests: every prompt plus everything generated so far. How many bytes one token costs is part 3; for now, it’s enough that the cost is per token.

Blocks and the pool

The obvious way to store this would be one contiguous buffer per request, sized for the longest sequence it might reach. That wastes most of the buffer on requests that finish early. PagedAttention [4] borrows the fix from operating systems: split the memory into fixed-size blocks (16 tokens each by default) and give each request blocks only as its tokens arrive.

How does this work in vLLM?

  1. At startup, vLLM allocates one large pool of blocks and puts all of them on a free_block_queue (in vllm/v1/core/block_pool.py). The size of that pool is the subject of this whole post.
  2. Each engine step, the scheduler calls the KV cache manager’s allocate_slots for every request it wants to run. A request with 17 new tokens needs ceil(17 / 16) = 2 blocks, taken from the free queue.
  3. When a request finishes, its blocks go back on the queue.
  4. If a running request needs a new block and none is free, the scheduler preempts another running request: it frees that request’s blocks, moves it back to the waiting queue, and later recomputes its tokens. (With prefix caching, on by default, freed blocks keep their contents until they’re reused, so some of that work can be recovered. Under memory pressure, usually little of it is.)

Here is a toy example with a pool of just 6 blocks:

a toy pool: 6 blocks × 16 tokensstep t: A and B hold 40 tokens eachA 16/16A 16/16A 8/16B 16/16B 16/16B 8/16free blocks: 0step t+1: A needs a 4th blockA 16/16A 16/16A 16/16A 1/16freefreefree blocks: 2no block was free, so B was preempted:its blocks freed, its 40 tokens recomputed later
vllm/v1/core · v0.30.0
# block_pool.pyself.free_block_queue = FreeKVCacheBlockQueue(self.blocks)# sched/scheduler.pypreempted_req = self.running.pop()
Figure 1. A toy pool of 6 blocks. Requests A and B hold 40 tokens each, which is 3 blocks each, so the pool is full. When A generates its 49th token it needs a fourth block. None is free, so the scheduler preempts B (the most recently added running request), frees its 3 blocks and gives one to A. B has to recompute its 40 tokens later.

Two things follow from this, and the rest of the post leans on both:

So the question that matters is: how big is the pool? vLLM doesn’t let you set it directly. It works it out at startup. Next, let’s watch it do that.

Startup: how vLLM sizes the pool

The pool gets whatever memory is left once everything else is on the card. vLLM can’t know most of those other terms in advance, so it measures them one at a time while it starts up, and sizes the pool last.

How does this work in vLLM? The engine core runs these steps on each GPU worker, in this order (see _initialize_kv_caches in vllm/v1/engine/core.py):

  1. Load the model (load_model): build the architecture and load the weights. The CUDA context, about 0.6 GiB, is already there from when the worker initialised its GPU. That’s also when vLLM checks that the card has at least gpu_memory_utilization × total memory free, and refuses to start if it doesn’t.
  2. Profile a forward pass (determine_available_memory → profile_run): run a dummy batch of max_num_batched_tokens tokens through the model and record the peak memory. That peak is mostly activations. They’re freed afterwards, but the space is kept back, because real batches will need it.
  3. Estimate the CUDA graphs (profile_cudagraph_memory): unless you pass --enforce-eager, measure what capturing the CUDA graphs will cost, and hold that back too.
  4. Size the pool: the budget (total memory × gpu_memory_utilization) minus everything measured so far. vLLM logs this as “Available KV cache memory”.
  5. Check max_model_len (get_kv_cache_configs): make sure the pool can hold at least one request of max_model_len tokens, and stop with an error if it can’t. Otherwise, divide the pool into blocks and log its size in tokens.
  6. Allocate and capture (initialize_from_config, then compile_or_warm_up_model): allocate the blocks, capture the CUDA graphs into the space held for them, and start serving.

Here is the running example going through those steps:

  • Weights
  • Overheads
  • Measured, then freed
  • Sized, not yet allocated
  • KV cache pool
  1. Load the weights15.56 / 21.15 GiB

    The weights go onto the card first, next to the CUDA context the process created when it launched.

  2. Profile a forward pass17.72 / 21.15 GiB

    A dummy batch of max_num_batched_tokens (8,192) runs through the model. The activation peak is recorded, freed, and held back as headroom.

    Failure mode 2 happens here: the peak itself doesn't fit.

  3. Estimate the CUDA graphs18.32 / 21.15 GiB

    vLLM measures what capturing the graphs will cost and holds that back too.

  4. Size the pool21.15 / 21.15 GiB

    Whatever is left under the budget becomes the pool: 21.15 − 14.96 − 0.60 − 2.16 − 0.60 = 2.83 GiB.

  5. Check max_model_len21.15 / 21.15 GiB

    One request of max_model_len must fit in the pool. The default, 131,072 tokens, needs 16 GiB; 8,192 needs 1 GiB.

    Failure mode 1 happens here, with the default max_model_len.

  6. Allocate the blocks, capture the graphs21.15 / 21.15 GiB

    The pool is carved into 1,450 blocks of 16 tokens, the graphs are captured into the space held for them, and the server starts.

    Failure mode 3 comes later, while serving, if the space above the budget is too thin.

The track is the whole card (23.5 GiB); the tick is the budget, card × gpu_memory_utilization.

vllm serve · startup log, running example
$ vllm serve meta-llama/Llama-3.1-8B-Instruct --gpu-memory-utilization 0.9 --max-num-batched-tokens 8192INFO Model loading took 14.96 GiB memory and … secondsINFO Estimated CUDA graph memory: 0.60 GiB totalINFO Available KV cache memory: 2.83 GiBERROR ValueError: To serve at least one request with the model's max seq len (131072), (16.0 GiB KV cache is needed, which is larger than the available KV cache memory (2.83 GiB). Based on the available memory, the estimated maximum model length is 23184. Try increasing `gpu_memory_utilization` …# rerun with --max-model-len 8192INFO GPU KV cache size: 23,200 tokens, Maximum concurrency for 8,192 tokens per request: 2.83xINFO Graph capturing finished in … secs, took 0.60 GiB
Figure 2. The running example’s gigabytes during startup, one step per row. The pool is sized only after everything else has been measured, and the max_model_len check at step 5 is where the default of 131,072 tokens fails. The readout is the log you’d see, in the same order.

Our running example fails at step 5. Llama 3.1 advertises a 131,072-token context, and without a max_model_len vLLM uses that. One request that long needs 16 GiB of KV cache, and the pool is 2.83 GiB. That is the whole content of the error everyone pastes into GitHub issues. The error even tells you what would fit: about 23,000 tokens. We’ll come back to the fix in part 4.

What each flag actually changes

With the steps in place, we can say precisely what each flag does [5], and it’s often not what people assume:

FlagWhat people think it doesWhat it does
gpu_memory_utilizationHow much of the GPU the weights may useSets the budget in step 4 (the tick in Figure 2): the fraction of the card’s total memory the whole vLLM process may use, weights, activations, graphs and pool together. It is not a fraction of free memory: if something else already holds memory on the GPU and less than that much is free, vLLM refuses to start.
max_num_batched_tokensRarely touchedSizes the dummy batch in step 2, so it sets the activation peak. It’s the flag behind OOMs that happen before the pool even exists.
max_model_lenReserves memory per requestReserves nothing (remember, blocks are handed out as tokens arrive). It only sets the bar for the check in step 5.
max_num_seqsThe concurrency limit that controls memoryA scheduler cap on how many sequences run in one step. It reserves no KV cache. At startup it only sizes the sampler’s logits buffer (max_num_seqs × vocab × 4 bytes, about 0.5 GiB for Llama at 1,024 sequences). It’s a latency and throughput knob, not a memory one.

The useful consequence: the only big levers on the pool are the weights and the utilization fraction. max_model_len decides whether vLLM agrees to start, and max_num_seqs barely touches memory.

The terms of the subtraction

Putting the steps together, these are the terms that come out of the budget before the pool gets anything:

TermTypicalWhat moves it
Model weights60–90%params × bytes_per_param ÷ tensor_parallel_size, except that FP8 and INT4 quantise only the linear layers: the embedding table and output head stay 16-bit. The one term you can cut by a large factor, with quantisation or another GPU.
Activation peak1–4 GiBScales with max_num_batched_tokens and hidden size. Measured in step 2.
CUDA graphs0.5–2 GiBMeasured in step 3. --enforce-eager frees all of it and costs decode throughput, most at small batch sizes, so measure it on your traffic.
CUDA context~0.6 GiBPer GPU, per process. Not reclaimable, and easy to forget.
NCCL buffers~0.9 GiBOnly when tensor_parallel_size > 1. We’ll meet these in part 4.
KV cache poolthe restNot configured, computed. This is what serves your requests.

For the running example, the finished subtraction looks like this:

  • Weights
  • Overheads
  • KV cache pool
  • Reserve
RTX 4090 · Llama 3.1 8B, 16-bit weightsKV pool 2.83 GiB → 23,200 tokens
Show as a table
ScenarioWeightsOverheadsKV cache poolReserveTokens
RTX 4090 · Llama 3.1 8B, 16-bit weights14.963.362.832.3523,200
vllm serve · RTX 4090, running example
$ vllm serve meta-llama/Llama-3.1-8B-Instruct --gpu-memory-utilization 0.9 --max-model-len 8192 --max-num-batched-tokens 8192INFO Available KV cache memory: 2.83 GiBINFO GPU KV cache size: 23,200 tokens, Maximum concurrency for 8,192 tokens per request: 2.83x
Figure 3. Where the 23.5 GiB of an RTX 4090 go for the running example. The weights take almost two-thirds of the card; the KV cache pool, the only part that serves requests, is what’s left: 2.83 GiB. Hover a bar for the breakdown.

We’ve now seen every line of that box except the middle one. It’s the line that turns gigabytes into tokens, and the one people most often get wrong, so let’s look at it next.

KV bytes per token

The pool comes out of the subtraction in gigabytes, but requests are measured in tokens. The conversion rate between the two is how many bytes of KV cache one token costs:

KV per token = 2 (key and value)
             × num_layers
             × num_kv_heads
             × head_dim
             × bytes per value (2 for BF16)

vLLM does the same sum per block, with one more factor of block_size (16). The term that needs explaining is num_kv_heads.

In the original transformer, every attention head had its own keys and values. Grouped-query attention [6] lets a group of query heads share one set: the model still has many query heads, but only a few KV heads, and only those are cached. The query vectors themselves are recomputed on every step and never stored. Here is one token of the running example:

one layer of Llama 3.1 8B, one token32 query headsnot cachedKVKVKVKVKVKVKVKV8 KV heads, cached4 query heads share eachK: 128 × 2 BV: 128 × 2 Bone KV head = 512 B512 B × 8 KV heads = 4 KiB / layer4 KiB × 32 layers = 128 KiB / tokenwith num_attention_heads (32) instead:512 KiB / token, 4× too big (8× on 70B)
config.json · meta-llama/Llama-3.1-8B-Instruct
"hidden_size": 4096,"num_attention_heads": 32,"num_hidden_layers": 32,"num_key_value_heads": 8,# head_dim = hidden_size / num_attention_heads = 128
Figure 4. What one token costs in Llama 3.1 8B. The 32 query heads aren’t cached; only the 8 KV heads they share are, each storing 128 K values and 128 V values at 2 bytes. Stack 32 layers and one token costs 128 KiB. The readout shows where each number lives in the model’s config.json.

Here is the same sum for a few common models:

ModelLayersKV headshead_dimKV / tokenat 8kat 128k
Llama 3.1 8B328128128 KiB1.0 GiB16.0 GiB
Llama 3.1 70B808128320 KiB2.5 GiB40.0 GiB
Llama 3.1 405B1268128504 KiB3.9 GiB63.0 GiB
Qwen3 32B648128256 KiB2.0 GiBn/a
Mistral 7B v0.3328128128 KiB1.0 GiBn/a

Look at the 128k column. On Llama 3.1 70B, one full-length conversation needs 40 GiB of KV cache, half an H100 on top of the weights. Long context isn’t something you switch on; it’s capacity you have to pay for.

With the cost of a token in hand, we have everything we need. Let’s finish the running example, then scale it up.

Scaling up: from one 4090 to two H100s

The running example, in full

Here is the whole subtraction for our 4090, including the conversion to tokens:

Card: RTX 4090, 23.5 GiB reported ·   weights BF16   ·   util 0.90

budget          23.5 × 0.90            =  21.15 GiB
− weights       8.03e9 × 2 bytes       =  14.96 GiB
− activations   (8192 batched tokens)  =   2.16 GiB
− CUDA graphs                          =   0.60 GiB
− CUDA context                         =   0.60 GiB
                                          ─────────
KV pool                                =   2.83 GiB

tokens = 2.83 GiB ÷ 128 KiB            =  23,200 tokens

As we saw in Figure 2, this fails the max_model_len check out of the box. Set --max-model-len 8192 and it starts, with room for two full-length requests at a time. That works, but it isn’t much of a server. Here is what each fix buys, ranked by what it costs you:

FixPoolTokensWhat it costs
Baseline, max_model_len 81922.8 GiB23,2002 concurrent requests
--enforce-eager3.4 GiB28,100Some decode throughput. Cheap and reversible
--kv-cache-dtype fp8 [7]2.8 GiB46,400Halves KV bytes per token. A small quality effect at long context
FP8 weights [8]9.3 GiB76,400A quality change you should measure, usually small
util 0.90 → 0.954.0 GiB32,800Less headroom for fragmentation. See failure mode 3

Quantising the weights did more than make room: it multiplied the pool by 3.3. None of the other terms changed, so every gigabyte saved on weights went straight into the pool. Keep an eye on that pattern; it shows up again at every scale.

  • Weights
  • Overheads
  • KV cache pool
  • Reserve
RTX 4090 · Llama 3.1 8B, 16-bit weightsKV pool 2.83 GiB → 23,200 tokens
RTX 4090 · Llama 3.1 8B, FP8 weightsKV pool 9.33 GiB → 76,400 tokens
Show as a table
ScenarioWeightsOverheadsKV cache poolReserveTokens
RTX 4090 · Llama 3.1 8B, 16-bit weights14.963.362.832.3523,200
RTX 4090 · Llama 3.1 8B, FP8 weights8.463.369.332.3576,400
vllm serve · RTX 4090, FP8 weights
$ vllm serve meta-llama/Llama-3.1-8B-Instruct --quantization fp8 --gpu-memory-utilization 0.9 --max-model-len 8192 --max-num-batched-tokens 8192INFO Model loading took 8.46 GiB memory and … secondsINFO Available KV cache memory: 9.33 GiBINFO GPU KV cache size: 76,448 tokens, Maximum concurrency for 8,192 tokens per request: 9.33x
Figure 5. The same card with FP8 weights. Halving the bytes of the linear layers frees 6.5 GiB (the embedding table and output head stay 16-bit), and all of it lands in the KV pool, which grows 3.3× to 9.3 GiB. The overheads and the reserve don’t move.

70B at FP8 on one H100

Now let’s swap in a model nearly nine times the size, on a card a little over three times as big, and run the same subtraction:

Card: H100 80GB, 79.2 GiB reported ·  weights FP8  ·  util 0.90

budget          79.2 × 0.90            =  71.28 GiB
− weights       68.45e9 × 1 byte
                + 2.10e9 × 2 bytes     =  67.66 GiB
− activations                          =   3.92 GiB
− CUDA graphs + context                =   1.20 GiB
                                          ─────────
KV pool                                =  −1.50 GiB   →  won't start

The first line of the weights is the 68.45B parameters in the linear layers, at FP8. The second is the 2.10B in the embedding table and output head, which stay 16-bit. The weights alone fit on the card, which is why so many “can I run 70B on one H100?” answers say yes. But with the overheads they come to 72.8 GiB against a 71.3 GiB budget, and vLLM stops with “No available memory for the cache blocks”. Even at the 0.92 default, the pool is 0.08 GiB: 269 tokens.

Push utilization to 0.96 and the pool reaches 3.25 GiB, about 10,650 tokens. That’s one 8k request at a time with no headroom left, so in practice a 70B at FP8 doesn’t serve on one H100. There are two real options:

The same model at tensor_parallel_size 2

How does this work in vLLM? With --tensor-parallel-size 2, vLLM starts one worker process per GPU, and each one runs the startup steps from Figure 2 on its own card. Three things change per GPU:

Here is the subtraction for one of the two GPUs:

2 × H100 80GB  ·  weights FP8  ·  util 0.90  ·  tp=2

budget / GPU                           =  71.28 GiB
− weights / GPU   67.66 ÷ 2            =  33.83 GiB
− activations (sharded)                =   2.16 GiB
− CUDA graphs + context + NCCL        =   2.20 GiB
                                          ─────────
KV pool / GPU                          =  33.09 GiB

KV per token per GPU  320 KiB ÷ 2      =     160 KiB
tokens = 33.09 GiB ÷ 160 KiB           =  216,900 tokens

Two GPUs didn’t double the capacity. Compared with the most one card can give (10,650 tokens at 0.96), they multiplied it by twenty. The reason is the pattern from the 4090: the weights are a fixed cost on each card, and everything freed above that cost becomes pool. Halving the weights per GPU freed 34 GiB on each card, and halving the KV bytes per token made each of those gigabytes hold twice as many tokens.

It’s also why “does it fit on one GPU?” is the wrong question to ask when sizing hardware. The better one is: how many tokens of cache do I get per dollar?

  • Weights
  • Overheads
  • KV cache pool
  • Reserve
1 × H100 · Llama 3.1 70B, FP8, util 0.96KV pool 3.25 GiB → 10,650 tokens
2 × H100 · Llama 3.1 70B, FP8, util 0.90, tp=2 (per GPU)KV pool 33.09 GiB → 216,900 tokens
Show as a table
ScenarioWeightsOverheadsKV cache poolReserveTokens
1 × H100 · Llama 3.1 70B, FP8, util 0.9667.665.123.253.1710,650
2 × H100 · Llama 3.1 70B, FP8, util 0.90, tp=2 (per GPU)33.834.3633.097.92216,900
vllm serve · Llama 3.1 70B at FP8 on H100s
# one H100, pushed to 0.96$ vllm serve meta-llama/Llama-3.1-70B-Instruct --quantization fp8 --gpu-memory-utilization 0.96 --max-model-len 8192 --max-num-batched-tokens 8192INFO GPU KV cache size: 10,640 tokens, Maximum concurrency for 8,192 tokens per request: 1.30x# two H100s at 0.90$ vllm serve meta-llama/Llama-3.1-70B-Instruct --quantization fp8 --tensor-parallel-size 2 --gpu-memory-utilization 0.9 --max-model-len 8192 --max-num-batched-tokens 8192INFO GPU KV cache size: 216,848 tokens, Maximum concurrency for 8,192 tokens per request: 26.47x
Figure 6. Llama 3.1 70B at FP8, per GPU. One H100 only starts if you push utilization to 0.96, and even then the weights leave a 3.25 GiB pool: 10,650 tokens. Split across two at 0.90, each GPU carries half the weights and a 33 GiB pool, and the pair holds 216,900 tokens.

Now, where did the halving of KV bytes per token come from? Tensor parallelism splits the attention heads across the GPUs, and the KV heads go with them. That works as long as there are heads to split:

boxes = GPUs, cells = KV headsKV / token / GPUtp=101234567320 KiBtp=201234567160 KiBtp=40123456780 KiBtp=80123456740 KiBtp=16001122334455667740 KiBtp=16: every head on two GPUs, no KV gain
vllm/config/model.py · v0.30.0
def get_num_kv_heads(self, parallel_config, arch_config=None) -> int:    # If tensor parallelism is used, we divide the number of KV heads by    # the tensor parallel size. We will replicate the KV heads in the    # case where the number of KV heads is smaller than the tensor    # parallel size so each GPU has at least one KV head.    return max(1, total_num_kv_heads // parallel_config.tensor_parallel_size)
Figure 7. Llama 3.1 70B’s 8 KV heads across tensor-parallel GPUs. Up to tp=8 each GPU holds a share, and KV bytes per token per GPU halve with every doubling. At tp=16 there are more GPUs than heads, so vLLM stores each head on two GPUs and the per-GPU cost stays at 40 KiB (failure mode 4). The readout is the line in vLLM that decides it.

Size it yourself

Here is the same arithmetic as a calculator. Pick a model, a GPU and your flags, and it runs the subtraction, tells you whether vLLM will start, and prints the command:

KV cache sizing calculatorEstimates the pool vLLM will end up with, and whether your max_model_len will boot.

Drives the activation peak. Halve it first if you OOM during startup.

Starts, but only 2 requests at full context fit in the pool. It will serve very few users at once and idle expensively.
KV cache pool2.8 GiB / GPU
Budget (90% of 23.5 GiB)21.2 GiB
Weights / GPU15.0 GiB
Activation peak2.2 GiB
CUDA context0.6 GiB
CUDA graphs0.6 GiB
Capacity
Tokens in the pool23,208
KV per token (model)128 KiB
KV per token / GPU128 KiB
Blocks (16 tok)1,450
Concurrent @ 8,1922
Concurrent @ 2,04811
Largest safe max_model_len16,384
vllm serve Llama-3.1-8B \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.9 \
  --max-num-batched-tokens 8192

The calculator also has its own page, /tools/kv-calculator, for bookmarking or linking.

Estimates, not a simulator [9]. Activation and graph overheads are modelled from typical observed values and vary by version, attention backend and batch shape. Treat the result as the right order of magnitude and the right direction for each flag, then confirm against what vLLM logs at startup.

Failure modes, and which step each one breaks

With the full picture in place, each of the common failures maps to one step of Figure 2 or one term of the subtraction. Here are the five I see most often.

1. “KV cache is needed, which is larger than the available KV cache memory”

This is step 5: the pool is smaller than one request of max_model_len. It’s exactly what happened to our running example. Fix it in this order, cheapest first:

  1. Lower max_model_len to the p99 context you actually serve, not the model’s advertised maximum.
  2. Switch to --kv-cache-dtype fp8.
  3. Quantise the weights.
  4. Add a GPU.

Raising gpu_memory_utilization belongs last, not first, because it borrows from the headroom that failure mode 3 needs. If you only need the server to start, --max-model-len auto makes vLLM pick the largest length that fits. That tells you your capacity, though, not what your traffic needs.

2. OOM during startup, before the server is ready

This one happens in step 2, the profiling forward pass, so it isn’t about the pool at all: the pool doesn’t exist yet. The activation peak is too big. Halve --max-num-batched-tokens, then add --enforce-eager if you still need room. If you raised the batched-token budget to speed up prefill, this is the cost of that choice.

3. Starts fine, runs for twenty minutes, then OOMs

This usually means gpu_memory_utilization is at 0.95 or above. The pool was sized correctly, but there’s almost no memory left above the budget for allocator fragmentation and short-lived buffers, so eventually a particular batch shape asks for memory that no free region can satisfy. Drop to 0.88–0.90. If the model only fits at 0.95, it doesn’t really fit on this GPU, and the fix is one of the levers from part 4.

4. Works at tp=1, OOMs at tp=2

A new cost appears the moment you shard: NCCL communication buffers, roughly 0.5–1 GiB per GPU, taken from the same budget as the weights, so the pool shrinks by that much. A config sized to the last gigabyte on one GPU can fail on two for that reason alone.

There’s also a quieter limit: KV heads only split while tensor_parallel_size is at most the number of KV heads. Past that, every GPU holds a full copy of at least one head, so KV bytes per token per GPU stop shrinking and only the weight savings still grow the pool. With 8 KV heads, tp of 1, 2, 4 and 8 split the cache cleanly, but tp=16 buys much less than you’d expect (Figure 7).

5. No OOM, but throughput falls off a cliff under load

This isn’t a memory error but a memory symptom, and it’s Figure 1 happening at scale. The pool is full, so the scheduler preempts running requests, frees their blocks and later recomputes them, which means paying for the same prefill twice. Watch the preemption count in vLLM’s metrics: if it’s above zero under normal traffic, the pool is too small for your load. The fix is the same arithmetic, not a bigger max_num_seqs; raising max_num_seqs here usually makes it worse.

Putting it all together: a sizing checklist

  1. Compute KV/token = 2 × layers × kv_heads × head_dim × bytes. Use the KV heads.
  2. Compute the pool: the budget, minus the weights, minus about 3–5 GiB of overhead.
  3. Divide. That’s your token capacity, and it’s what your deployment is really made of.
  4. Set max_model_len from the p99 context you really serve, not the model’s maximum. This is the biggest lever most people never touch.
  5. Leave gpu_memory_utilization at or just below the default (0.92), and treat any need to raise it as a sign the hardware is undersized.
  6. Only then tune max_num_seqs, and tune it for latency, which is what it’s for.

Epilogue

We began with what the KV cache is and why vLLM hands it out in blocks, followed the running example through startup as vLLM sized its pool, worked out what one token costs, and scaled up to two H100s, where a second GPU bought twenty times the cache instead of two. Finally, we traced the common OOMs back to the step that causes each one.

This post skipped some things that change the arithmetic. E.g.:

The nice thing is that each of these changes one term of the subtraction, not its shape. Defaults move between vLLM versions and the overhead terms drift with the attention backend, but the method doesn’t. If a number here disagrees with what your server logs at startup, trust the server, then check which term you mis-estimated.

References

  1. vLLM, github.com/vllm-project/vllm
  2. “The Llama 3 Herd of Models”, arxiv.org/abs/2407.21783
  3. “Attention Is All You Need”, arxiv.org/abs/1706.03762
  4. “Efficient Memory Management for Large Language Model Serving with PagedAttention”, arxiv.org/abs/2309.06180
  5. vLLM docs: Engine Arguments, docs.vllm.ai/en/latest/configuration/engine_args.html
  6. “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”, arxiv.org/abs/2305.13245
  7. vLLM docs: Quantized KV Cache, docs.vllm.ai/en/latest/features/quantization/quantized_kvcache.html
  8. “FP8 Formats for Deep Learning”, arxiv.org/abs/2209.05433
  9. vLLM docs: Optimization and Tuning, docs.vllm.ai/en/latest/configuration/optimization.html
  10. “Mixtral of Experts”, arxiv.org/abs/2401.04088
  11. “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model”, arxiv.org/abs/2405.04434
  12. “Mistral 7B”, arxiv.org/abs/2310.06825
  13. “LoRA: Low-Rank Adaptation of Large Language Models”, arxiv.org/abs/2106.09685