Inside the KV Cache: The Life of a Gigabyte
From gpu_memory_utilization and max_model_len to weights, activations, CUDA graphs and the KV pool: where every gigabyte on the card actually goes.
In this post, I’ll follow the memory on a GPU running vLLM [1], gigabyte by gigabyte, from the moment vllm serve starts to the KV cache pool that ends up serving your requests. In particular, I’ll show how vLLM works out the size of that pool, and why the flags everyone tunes by feel (gpu_memory_utilization, max_model_len and max_num_seqs) all act on that one number.
It starts with what the KV cache is and then layers in detail, one term at a time, so you can build an accurate mental model of where the memory goes before we do any serious arithmetic.
This post is structured into five parts:
- Background: what the KV cache is, how vLLM hands it out in blocks, and why the pool’s size is the server’s capacity
- Startup: how vLLM sizes the pool, step by step, and what each flag changes along the way
- KV bytes per token: the number that turns gigabytes into tokens
- Scaling up: from one RTX 4090 to two H100s, and a calculator for your own setup
- Failure modes: five common OOMs, each traced back to the step that causes it
Background: the KV cache
We’ll use the following command as our running example, on a single RTX 4090:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--gpu-memory-utilization 0.9 \
--max-num-batched-tokens 8192This configuration is:
- Llama 3.1 8B [2] with 16-bit weights (2 bytes per parameter)
- one RTX 4090, which CUDA reports as 23.5 GiB (tensor parallel size 1)
gpu_memory_utilizationat 0.90 (the default is 0.92 in v0.30; 0.90 leaves a little more headroom)max_num_batched_tokensat 8,192- no
max_model_len, so vLLM uses the model’s own maximum (we’ll see why that matters) - a standard dense transformer with grouped-query attention
From here, we’ll gradually build up to a 70B model on two H100s. The hardware and the numbers get bigger, but the method never changes.
Before this command can serve a single request, vLLM has to decide how much of the card goes to the KV cache. So let’s start with what the KV cache is.
What the KV cache is
A transformer [3] generates text one token at a time. To produce each new token, every attention layer compares that token against all the tokens before it, and for that it needs each earlier token’s key and value vectors, at every layer.
Those vectors don’t change once they’re computed, so recomputing them on every step would redo the whole prompt for every token generated. Instead, the engine keeps them in GPU memory. That store is the KV cache, and it works in two phases:
- Prefill: a forward pass over the whole prompt writes K and V for every prompt token.
- Decode: each step generates one token and appends that token’s K and V.
So the cache grows by a fixed number of bytes per token, and the memory a server needs at any moment is the total number of tokens in flight across all of its requests: every prompt plus everything generated so far. How many bytes one token costs is part 3; for now, it’s enough that the cost is per token.
Blocks and the pool
The obvious way to store this would be one contiguous buffer per request, sized for the longest sequence it might reach. That wastes most of the buffer on requests that finish early. PagedAttention [4] borrows the fix from operating systems: split the memory into fixed-size blocks (16 tokens each by default) and give each request blocks only as its tokens arrive.
How does this work in vLLM?
- At startup, vLLM allocates one large pool of blocks and puts all of them on a
free_block_queue(invllm/v1/core/block_pool.py). The size of that pool is the subject of this whole post. - Each engine step, the scheduler calls the KV cache manager’s
allocate_slotsfor every request it wants to run. A request with 17 new tokens needs ceil(17 / 16) = 2 blocks, taken from the free queue. - When a request finishes, its blocks go back on the queue.
- If a running request needs a new block and none is free, the scheduler preempts another running request: it frees that request’s blocks, moves it back to the waiting queue, and later recomputes its tokens. (With prefix caching, on by default, freed blocks keep their contents until they’re reused, so some of that work can be recovered. Under memory pressure, usually little of it is.)
Here is a toy example with a pool of just 6 blocks:
# block_pool.pyself.free_block_queue = FreeKVCacheBlockQueue(self.blocks)# sched/scheduler.pypreempted_req = self.running.pop()
Two things follow from this, and the rest of the post leans on both:
- Nothing is reserved per request. A request only holds blocks for tokens it actually has. Setting a large
max_model_lendoesn’t set any memory aside. - The pool is the server’s capacity. Its size in tokens caps how many tokens can be in flight across all requests at once. When it runs out, requests wait or get preempted, and preempted work is paid for twice.
So the question that matters is: how big is the pool? vLLM doesn’t let you set it directly. It works it out at startup. Next, let’s watch it do that.
Startup: how vLLM sizes the pool
The pool gets whatever memory is left once everything else is on the card. vLLM can’t know most of those other terms in advance, so it measures them one at a time while it starts up, and sizes the pool last.
How does this work in vLLM? The engine core runs these steps on each GPU worker, in this order (see _initialize_kv_caches in vllm/v1/engine/core.py):
- Load the model (
load_model): build the architecture and load the weights. The CUDA context, about 0.6 GiB, is already there from when the worker initialised its GPU. That’s also when vLLM checks that the card has at leastgpu_memory_utilization× total memory free, and refuses to start if it doesn’t. - Profile a forward pass (
determine_available_memory→profile_run): run a dummy batch ofmax_num_batched_tokenstokens through the model and record the peak memory. That peak is mostly activations. They’re freed afterwards, but the space is kept back, because real batches will need it. - Estimate the CUDA graphs (
profile_cudagraph_memory): unless you pass--enforce-eager, measure what capturing the CUDA graphs will cost, and hold that back too. - Size the pool: the budget (total memory ×
gpu_memory_utilization) minus everything measured so far. vLLM logs this as “Available KV cache memory”. - Check max_model_len (
get_kv_cache_configs): make sure the pool can hold at least one request ofmax_model_lentokens, and stop with an error if it can’t. Otherwise, divide the pool into blocks and log its size in tokens. - Allocate and capture (
initialize_from_config, thencompile_or_warm_up_model): allocate the blocks, capture the CUDA graphs into the space held for them, and start serving.
Here is the running example going through those steps:
- Weights
- Overheads
- Measured, then freed
- Sized, not yet allocated
- KV cache pool
- Load the weights15.56 / 21.15 GiB
The weights go onto the card first, next to the CUDA context the process created when it launched.
- Profile a forward pass17.72 / 21.15 GiB
A dummy batch of max_num_batched_tokens (8,192) runs through the model. The activation peak is recorded, freed, and held back as headroom.
Failure mode 2 happens here: the peak itself doesn't fit.
- Estimate the CUDA graphs18.32 / 21.15 GiB
vLLM measures what capturing the graphs will cost and holds that back too.
- Size the pool21.15 / 21.15 GiB
Whatever is left under the budget becomes the pool: 21.15 − 14.96 − 0.60 − 2.16 − 0.60 = 2.83 GiB.
- Check max_model_len21.15 / 21.15 GiB
One request of max_model_len must fit in the pool. The default, 131,072 tokens, needs 16 GiB; 8,192 needs 1 GiB.
Failure mode 1 happens here, with the default max_model_len.
- Allocate the blocks, capture the graphs21.15 / 21.15 GiB
The pool is carved into 1,450 blocks of 16 tokens, the graphs are captured into the space held for them, and the server starts.
Failure mode 3 comes later, while serving, if the space above the budget is too thin.
The track is the whole card (23.5 GiB); the tick is the budget, card × gpu_memory_utilization.
$ vllm serve meta-llama/Llama-3.1-8B-Instruct --gpu-memory-utilization 0.9 --max-num-batched-tokens 8192INFO Model loading took 14.96 GiB memory and … secondsINFO Estimated CUDA graph memory: 0.60 GiB totalINFO Available KV cache memory: 2.83 GiBERROR ValueError: To serve at least one request with the model's max seq len (131072), (16.0 GiB KV cache is needed, which is larger than the available KV cache memory (2.83 GiB). Based on the available memory, the estimated maximum model length is 23184. Try increasing `gpu_memory_utilization` …# rerun with --max-model-len 8192INFO GPU KV cache size: 23,200 tokens, Maximum concurrency for 8,192 tokens per request: 2.83xINFO Graph capturing finished in … secs, took 0.60 GiB
max_model_len check at step 5 is where the default of 131,072 tokens fails. The readout is the log you’d see, in the same order.Our running example fails at step 5. Llama 3.1 advertises a 131,072-token context, and without a max_model_len vLLM uses that. One request that long needs 16 GiB of KV cache, and the pool is 2.83 GiB. That is the whole content of the error everyone pastes into GitHub issues. The error even tells you what would fit: about 23,000 tokens. We’ll come back to the fix in part 4.
What each flag actually changes
With the steps in place, we can say precisely what each flag does [5], and it’s often not what people assume:
| Flag | What people think it does | What it does |
|---|---|---|
gpu_memory_utilization | How much of the GPU the weights may use | Sets the budget in step 4 (the tick in Figure 2): the fraction of the card’s total memory the whole vLLM process may use, weights, activations, graphs and pool together. It is not a fraction of free memory: if something else already holds memory on the GPU and less than that much is free, vLLM refuses to start. |
max_num_batched_tokens | Rarely touched | Sizes the dummy batch in step 2, so it sets the activation peak. It’s the flag behind OOMs that happen before the pool even exists. |
max_model_len | Reserves memory per request | Reserves nothing (remember, blocks are handed out as tokens arrive). It only sets the bar for the check in step 5. |
max_num_seqs | The concurrency limit that controls memory | A scheduler cap on how many sequences run in one step. It reserves no KV cache. At startup it only sizes the sampler’s logits buffer (max_num_seqs × vocab × 4 bytes, about 0.5 GiB for Llama at 1,024 sequences). It’s a latency and throughput knob, not a memory one. |
The useful consequence: the only big levers on the pool are the weights and the utilization fraction. max_model_len decides whether vLLM agrees to start, and max_num_seqs barely touches memory.
The terms of the subtraction
Putting the steps together, these are the terms that come out of the budget before the pool gets anything:
| Term | Typical | What moves it |
|---|---|---|
| Model weights | 60–90% | params × bytes_per_param ÷ tensor_parallel_size, except that FP8 and INT4 quantise only the linear layers: the embedding table and output head stay 16-bit. The one term you can cut by a large factor, with quantisation or another GPU. |
| Activation peak | 1–4 GiB | Scales with max_num_batched_tokens and hidden size. Measured in step 2. |
| CUDA graphs | 0.5–2 GiB | Measured in step 3. --enforce-eager frees all of it and costs decode throughput, most at small batch sizes, so measure it on your traffic. |
| CUDA context | ~0.6 GiB | Per GPU, per process. Not reclaimable, and easy to forget. |
| NCCL buffers | ~0.9 GiB | Only when tensor_parallel_size > 1. We’ll meet these in part 4. |
| KV cache pool | the rest | Not configured, computed. This is what serves your requests. |
For the running example, the finished subtraction looks like this:
- Weights
- Overheads
- KV cache pool
- Reserve
Show as a table
| Scenario | Weights | Overheads | KV cache pool | Reserve | Tokens |
|---|---|---|---|---|---|
| RTX 4090 · Llama 3.1 8B, 16-bit weights | 14.96 | 3.36 | 2.83 | 2.35 | 23,200 |
$ vllm serve meta-llama/Llama-3.1-8B-Instruct --gpu-memory-utilization 0.9 --max-model-len 8192 --max-num-batched-tokens 8192INFO Available KV cache memory: 2.83 GiBINFO GPU KV cache size: 23,200 tokens, Maximum concurrency for 8,192 tokens per request: 2.83x
We’ve now seen every line of that box except the middle one. It’s the line that turns gigabytes into tokens, and the one people most often get wrong, so let’s look at it next.
KV bytes per token
The pool comes out of the subtraction in gigabytes, but requests are measured in tokens. The conversion rate between the two is how many bytes of KV cache one token costs:
KV per token = 2 (key and value)
× num_layers
× num_kv_heads
× head_dim
× bytes per value (2 for BF16)vLLM does the same sum per block, with one more factor of block_size (16). The term that needs explaining is num_kv_heads.
In the original transformer, every attention head had its own keys and values. Grouped-query attention [6] lets a group of query heads share one set: the model still has many query heads, but only a few KV heads, and only those are cached. The query vectors themselves are recomputed on every step and never stored. Here is one token of the running example:
"hidden_size": 4096,"num_attention_heads": 32,"num_hidden_layers": 32,"num_key_value_heads": 8,# head_dim = hidden_size / num_attention_heads = 128
config.json.Here is the same sum for a few common models:
| Model | Layers | KV heads | head_dim | KV / token | at 8k | at 128k |
|---|---|---|---|---|---|---|
| Llama 3.1 8B | 32 | 8 | 128 | 128 KiB | 1.0 GiB | 16.0 GiB |
| Llama 3.1 70B | 80 | 8 | 128 | 320 KiB | 2.5 GiB | 40.0 GiB |
| Llama 3.1 405B | 126 | 8 | 128 | 504 KiB | 3.9 GiB | 63.0 GiB |
| Qwen3 32B | 64 | 8 | 128 | 256 KiB | 2.0 GiB | n/a |
| Mistral 7B v0.3 | 32 | 8 | 128 | 128 KiB | 1.0 GiB | n/a |
Look at the 128k column. On Llama 3.1 70B, one full-length conversation needs 40 GiB of KV cache, half an H100 on top of the weights. Long context isn’t something you switch on; it’s capacity you have to pay for.
With the cost of a token in hand, we have everything we need. Let’s finish the running example, then scale it up.
Scaling up: from one 4090 to two H100s
The running example, in full
Here is the whole subtraction for our 4090, including the conversion to tokens:
Card: RTX 4090, 23.5 GiB reported · weights BF16 · util 0.90
budget 23.5 × 0.90 = 21.15 GiB
− weights 8.03e9 × 2 bytes = 14.96 GiB
− activations (8192 batched tokens) = 2.16 GiB
− CUDA graphs = 0.60 GiB
− CUDA context = 0.60 GiB
─────────
KV pool = 2.83 GiB
tokens = 2.83 GiB ÷ 128 KiB = 23,200 tokensAs we saw in Figure 2, this fails the max_model_len check out of the box. Set --max-model-len 8192 and it starts, with room for two full-length requests at a time. That works, but it isn’t much of a server. Here is what each fix buys, ranked by what it costs you:
| Fix | Pool | Tokens | What it costs |
|---|---|---|---|
Baseline, max_model_len 8192 | 2.8 GiB | 23,200 | 2 concurrent requests |
--enforce-eager | 3.4 GiB | 28,100 | Some decode throughput. Cheap and reversible |
--kv-cache-dtype fp8 [7] | 2.8 GiB | 46,400 | Halves KV bytes per token. A small quality effect at long context |
| FP8 weights [8] | 9.3 GiB | 76,400 | A quality change you should measure, usually small |
| util 0.90 → 0.95 | 4.0 GiB | 32,800 | Less headroom for fragmentation. See failure mode 3 |
Quantising the weights did more than make room: it multiplied the pool by 3.3. None of the other terms changed, so every gigabyte saved on weights went straight into the pool. Keep an eye on that pattern; it shows up again at every scale.
- Weights
- Overheads
- KV cache pool
- Reserve
Show as a table
| Scenario | Weights | Overheads | KV cache pool | Reserve | Tokens |
|---|---|---|---|---|---|
| RTX 4090 · Llama 3.1 8B, 16-bit weights | 14.96 | 3.36 | 2.83 | 2.35 | 23,200 |
| RTX 4090 · Llama 3.1 8B, FP8 weights | 8.46 | 3.36 | 9.33 | 2.35 | 76,400 |
$ vllm serve meta-llama/Llama-3.1-8B-Instruct --quantization fp8 --gpu-memory-utilization 0.9 --max-model-len 8192 --max-num-batched-tokens 8192INFO Model loading took 8.46 GiB memory and … secondsINFO Available KV cache memory: 9.33 GiBINFO GPU KV cache size: 76,448 tokens, Maximum concurrency for 8,192 tokens per request: 9.33x
70B at FP8 on one H100
Now let’s swap in a model nearly nine times the size, on a card a little over three times as big, and run the same subtraction:
Card: H100 80GB, 79.2 GiB reported · weights FP8 · util 0.90
budget 79.2 × 0.90 = 71.28 GiB
− weights 68.45e9 × 1 byte
+ 2.10e9 × 2 bytes = 67.66 GiB
− activations = 3.92 GiB
− CUDA graphs + context = 1.20 GiB
─────────
KV pool = −1.50 GiB → won't startThe first line of the weights is the 68.45B parameters in the linear layers, at FP8. The second is the 2.10B in the embedding table and output head, which stay 16-bit. The weights alone fit on the card, which is why so many “can I run 70B on one H100?” answers say yes. But with the overheads they come to 72.8 GiB against a 71.3 GiB budget, and vLLM stops with “No available memory for the cache blocks”. Even at the 0.92 default, the pool is 0.08 GiB: 269 tokens.
Push utilization to 0.96 and the pool reaches 3.25 GiB, about 10,650 tokens. That’s one 8k request at a time with no headroom left, so in practice a 70B at FP8 doesn’t serve on one H100. There are two real options:
- INT4 weights: about 37 GiB (4-bit linear layers plus their group scales, 16-bit embeddings), which leaves a 29 GiB pool: about 95,000 tokens, or 11 concurrent requests at 8k. That’s a real server, with a quality question you have to answer by measuring.
- A second GPU, which is next.
The same model at tensor_parallel_size 2
How does this work in vLLM? With --tensor-parallel-size 2, vLLM starts one worker process per GPU, and each one runs the startup steps from Figure 2 on its own card. Three things change per GPU:
- The weights are split, so each GPU loads half.
- The attention heads are split too, so each GPU caches only its share of every token’s KV: half the bytes per token (more on this in Figure 7).
- NCCL buffers appear. The GPUs combine their partial results twice in every layer, and the communication buffers take about 0.9 GiB per card, out of the same budget.
Here is the subtraction for one of the two GPUs:
2 × H100 80GB · weights FP8 · util 0.90 · tp=2
budget / GPU = 71.28 GiB
− weights / GPU 67.66 ÷ 2 = 33.83 GiB
− activations (sharded) = 2.16 GiB
− CUDA graphs + context + NCCL = 2.20 GiB
─────────
KV pool / GPU = 33.09 GiB
KV per token per GPU 320 KiB ÷ 2 = 160 KiB
tokens = 33.09 GiB ÷ 160 KiB = 216,900 tokensTwo GPUs didn’t double the capacity. Compared with the most one card can give (10,650 tokens at 0.96), they multiplied it by twenty. The reason is the pattern from the 4090: the weights are a fixed cost on each card, and everything freed above that cost becomes pool. Halving the weights per GPU freed 34 GiB on each card, and halving the KV bytes per token made each of those gigabytes hold twice as many tokens.
It’s also why “does it fit on one GPU?” is the wrong question to ask when sizing hardware. The better one is: how many tokens of cache do I get per dollar?
- Weights
- Overheads
- KV cache pool
- Reserve
Show as a table
| Scenario | Weights | Overheads | KV cache pool | Reserve | Tokens |
|---|---|---|---|---|---|
| 1 × H100 · Llama 3.1 70B, FP8, util 0.96 | 67.66 | 5.12 | 3.25 | 3.17 | 10,650 |
| 2 × H100 · Llama 3.1 70B, FP8, util 0.90, tp=2 (per GPU) | 33.83 | 4.36 | 33.09 | 7.92 | 216,900 |
# one H100, pushed to 0.96$ vllm serve meta-llama/Llama-3.1-70B-Instruct --quantization fp8 --gpu-memory-utilization 0.96 --max-model-len 8192 --max-num-batched-tokens 8192INFO GPU KV cache size: 10,640 tokens, Maximum concurrency for 8,192 tokens per request: 1.30x# two H100s at 0.90$ vllm serve meta-llama/Llama-3.1-70B-Instruct --quantization fp8 --tensor-parallel-size 2 --gpu-memory-utilization 0.9 --max-model-len 8192 --max-num-batched-tokens 8192INFO GPU KV cache size: 216,848 tokens, Maximum concurrency for 8,192 tokens per request: 26.47x
Now, where did the halving of KV bytes per token come from? Tensor parallelism splits the attention heads across the GPUs, and the KV heads go with them. That works as long as there are heads to split:
def get_num_kv_heads(self, parallel_config, arch_config=None) -> int: # If tensor parallelism is used, we divide the number of KV heads by # the tensor parallel size. We will replicate the KV heads in the # case where the number of KV heads is smaller than the tensor # parallel size so each GPU has at least one KV head. return max(1, total_num_kv_heads // parallel_config.tensor_parallel_size)
Size it yourself
Here is the same arithmetic as a calculator. Pick a model, a GPU and your flags, and it runs the subtraction, tells you whether vLLM will start, and prints the command:
max_model_len will boot.Drives the activation peak. Halve it first if you OOM during startup.
vllm serve Llama-3.1-8B \
--max-model-len 8192 \
--gpu-memory-utilization 0.9 \
--max-num-batched-tokens 8192The calculator also has its own page, /tools/kv-calculator, for bookmarking or linking.
Estimates, not a simulator [9]. Activation and graph overheads are modelled from typical observed values and vary by version, attention backend and batch shape. Treat the result as the right order of magnitude and the right direction for each flag, then confirm against what vLLM logs at startup.
Failure modes, and which step each one breaks
With the full picture in place, each of the common failures maps to one step of Figure 2 or one term of the subtraction. Here are the five I see most often.
1. “KV cache is needed, which is larger than the available KV cache memory”
This is step 5: the pool is smaller than one request of max_model_len. It’s exactly what happened to our running example. Fix it in this order, cheapest first:
- Lower
max_model_lento the p99 context you actually serve, not the model’s advertised maximum. - Switch to
--kv-cache-dtype fp8. - Quantise the weights.
- Add a GPU.
Raising gpu_memory_utilization belongs last, not first, because it borrows from the headroom that failure mode 3 needs. If you only need the server to start, --max-model-len auto makes vLLM pick the largest length that fits. That tells you your capacity, though, not what your traffic needs.
2. OOM during startup, before the server is ready
This one happens in step 2, the profiling forward pass, so it isn’t about the pool at all: the pool doesn’t exist yet. The activation peak is too big. Halve --max-num-batched-tokens, then add --enforce-eager if you still need room. If you raised the batched-token budget to speed up prefill, this is the cost of that choice.
3. Starts fine, runs for twenty minutes, then OOMs
This usually means gpu_memory_utilization is at 0.95 or above. The pool was sized correctly, but there’s almost no memory left above the budget for allocator fragmentation and short-lived buffers, so eventually a particular batch shape asks for memory that no free region can satisfy. Drop to 0.88–0.90. If the model only fits at 0.95, it doesn’t really fit on this GPU, and the fix is one of the levers from part 4.
4. Works at tp=1, OOMs at tp=2
A new cost appears the moment you shard: NCCL communication buffers, roughly 0.5–1 GiB per GPU, taken from the same budget as the weights, so the pool shrinks by that much. A config sized to the last gigabyte on one GPU can fail on two for that reason alone.
There’s also a quieter limit: KV heads only split while tensor_parallel_size is at most the number of KV heads. Past that, every GPU holds a full copy of at least one head, so KV bytes per token per GPU stop shrinking and only the weight savings still grow the pool. With 8 KV heads, tp of 1, 2, 4 and 8 split the cache cleanly, but tp=16 buys much less than you’d expect (Figure 7).
5. No OOM, but throughput falls off a cliff under load
This isn’t a memory error but a memory symptom, and it’s Figure 1 happening at scale. The pool is full, so the scheduler preempts running requests, frees their blocks and later recomputes them, which means paying for the same prefill twice. Watch the preemption count in vLLM’s metrics: if it’s above zero under normal traffic, the pool is too small for your load. The fix is the same arithmetic, not a bigger max_num_seqs; raising max_num_seqs here usually makes it worse.
Putting it all together: a sizing checklist
- Compute
KV/token = 2 × layers × kv_heads × head_dim × bytes. Use the KV heads. - Compute the pool: the budget, minus the weights, minus about 3–5 GiB of overhead.
- Divide. That’s your token capacity, and it’s what your deployment is really made of.
- Set
max_model_lenfrom the p99 context you really serve, not the model’s maximum. This is the biggest lever most people never touch. - Leave
gpu_memory_utilizationat or just below the default (0.92), and treat any need to raise it as a sign the hardware is undersized. - Only then tune
max_num_seqs, and tune it for latency, which is what it’s for.
Epilogue
We began with what the KV cache is and why vLLM hands it out in blocks, followed the running example through startup as vLLM sized its pool, worked out what one token costs, and scaled up to two H100s, where a second GPU bought twenty times the cache instead of two. Finally, we traced the common OOMs back to the step that causes each one.
This post skipped some things that change the arithmetic. E.g.:
- Mixture-of-experts [10]: only a couple of experts are active per token, but all of them stay resident, so the weight term uses total parameters while the compute uses active ones.
- Multi-head latent attention [11], which caches a compressed latent instead of full K and V and changes the per-token number by roughly an order of magnitude.
- Sliding-window attention [12], which caps KV per layer instead of letting it grow with context.
- LoRA adapters [13], which add their own resident weights per adapter.
- Quantised-KV quality, which deserves measurement rather than a rule of thumb.
The nice thing is that each of these changes one term of the subtraction, not its shape. Defaults move between vLLM versions and the overhead terms drift with the attention backend, but the method doesn’t. If a number here disagrees with what your server logs at startup, trust the server, then check which term you mis-estimated.
Get notified when I publish a new post.
Or follow the RSS feed.
References
- vLLM, github.com/vllm-project/vllm
- “The Llama 3 Herd of Models”, arxiv.org/abs/2407.21783
- “Attention Is All You Need”, arxiv.org/abs/1706.03762
- “Efficient Memory Management for Large Language Model Serving with PagedAttention”, arxiv.org/abs/2309.06180
- vLLM docs: Engine Arguments, docs.vllm.ai/en/latest/configuration/engine_args.html
- “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”, arxiv.org/abs/2305.13245
- vLLM docs: Quantized KV Cache, docs.vllm.ai/en/latest/features/quantization/quantized_kvcache.html
- “FP8 Formats for Deep Learning”, arxiv.org/abs/2209.05433
- vLLM docs: Optimization and Tuning, docs.vllm.ai/en/latest/configuration/optimization.html
- “Mixtral of Experts”, arxiv.org/abs/2401.04088
- “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model”, arxiv.org/abs/2405.04434
- “Mistral 7B”, arxiv.org/abs/2310.06825
- “LoRA: Low-Rank Adaptation of Large Language Models”, arxiv.org/abs/2106.09685