vLLM KV cache calculator
Pick a model and a GPU; see whether vLLM will boot, how big the KV cache pool is, and how many requests fit.
KV cache sizing calculatorEstimates the pool vLLM will end up with, and whether your
max_model_len will boot.Drives the activation peak. Halve it first if you OOM during startup.
Starts, but only 2 requests at full context fit in the pool. It will serve very few users at once and idle expensively.
KV cache pool2.8 GiB / GPU
Budget (90% of 23.5 GiB)21.2 GiB
Weights / GPU15.0 GiB
Activation peak2.2 GiB
CUDA context0.6 GiB
CUDA graphs0.6 GiB
Capacity
Tokens in the pool23,208
KV per token (model)128 KiB
KV per token / GPU128 KiB
Blocks (16 tok)1,450
Concurrent @ 8,1922
Concurrent @ 2,04811
Largest safe max_model_len16,384
vllm serve Llama-3.1-8B \
--max-model-len 8192 \
--gpu-memory-utilization 0.9 \
--max-num-batched-tokens 8192Estimates, not a simulator: activation and CUDA graph overheads are modelled from typical values and vary by version and attention backend. The derivation behind every row is in Inside the KV Cache: The Life of a Gigabyte.