Skip to content
AI•4 min read

The 5 vLLM Flags That Actually Matter

vLLM has over a hundred configuration flags. Five decide how much context you hold, how much you spend, and whether your agent works. The rest is fine tuning for when you already measure. This chapter opens those five with the real number each one moves.

The mental rule before we start: vLLM splits your VRAM into two piles. One fixed pile for the model's weights, one variable pile for the KV cache (your sessions' context). Almost every flag is a way to cut that cake.

1. --max-model-len 65536

The maximum context window the server will accept, in tokens.

With the KV cache in FP8, each context token costs 48 KB of VRAM (the full math is in 48 KB per token). A 64K session costs 65,536 × 48 KB ≈ 3.1 GB. A 128K one, double.

What it breaks if you raise it without thinking: startup dies with No available memory for the cache blocks, because you promised more memory than is left after the weights. What it breaks if you lower it too much: your agents with long sessions die with "context length exceeded".

In my PoC: 65536. With ~175K tokens of KV cache on the L40S, that fits 2 full 64K sessions.

2. --gpu-memory-utilization 0.92

The percentage of VRAM vLLM may use. With 48 GB, 0.92 is ~44 GB for weights + KV cache + activations.

  • 31 GB goes to the 30B model's weights in FP8.
  • What's left (~13 GB before overhead) is your KV cache: the ~150-200K tokens you saw in chapter 4's log.

If you hit OOM at startup, lower it to 0.88 before cutting context. And don't set it to 1.0: the runtime needs margin for CUDA and buffers, and a startup with 0 MB free is a startup that fails.

3. --kv-cache-dtype fp8

Stores the KV cache in 8 bits instead of 16. Doubles your context capacity at the cost of a little precision: 96 KB per token in FP16 becomes 48 KB in FP8.

For agents reading code, the quality impact is practically invisible; the capacity impact is double the sessions. This flag is the reason the guide demands an Ada or newer GPU: Ampere cards (A40, A100, A6000) don't support FP8 in hardware and startup fails. If that's you, drop it and move on: you lose half the context, but it works.

4. --enable-prefix-caching

Reuses the KV cache of already-processed prefixes. If all your requests start with the same system prompt and the same history (which is exactly what an agent does on every turn), vLLM doesn't re-process that text: it reuses it.

This is the best-return flag on the list. My PoC measurement, same scenario of 10 users with 16K context:

ScenarioTotal timeTTFTAggregate throughput
Without shared prefix16.1 s9,652 ms199 tok/s
With shared prefix5.2 s3,773 ms740 tok/s

The full breakdown is in Prefix caching: same load, 3× faster. In an agentic workload, where every turn re-sends the history, this flag dominates the latency bill.

5. --enable-auto-tool-choice --tool-call-parser qwen3_coder

Without these two flags, your agent doesn't work. The model generates tool calls as text and the client receives prose instead of executing them. It's the quietest bug on the list: everything "seems" to work until the agent tries to edit a file.

The parser depends on the model: qwen3_coder for the Qwen3 Coder family; if vLLM complains about the name, vllm serve --help | grep -A5 tool-call-parser and try hermes. Every model family has its own.

The full command, with what each flag controls

nohup vllm serve Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 \
  --served-model-name qwen3-coder-30b \
  --max-model-len 65536 \
  --gpu-memory-utilization 0.92 \
  --kv-cache-dtype fp8 \
  --enable-prefix-caching \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --api-key $VLLM_API_KEY \
  --host 0.0.0.0 --port 8000 \
  > /workspace/vllm.log 2>&1 &

Every flag is explained above. The short version: the alias is for clients, the window is per session, the percentage splits VRAM, FP8 charges 48 KB per token, prefix caching reuses repeated context, and the last two wake the agents up.

Checkpoint and the typical mistake

Checkpoint: you understand every flag in your startup command and the number it controls. The test: could you say how many 64K sessions you can hold before running vllm serve?

The typical mistake: copying flags from a tutorial without checking your VRAM. This guide's numbers correspond to 48 GB; on a 24 GB GPU, the same command won't start. Always nvidia-smi first, then the command.

Next chapter

Now that the server is properly configured, we measure it the right way: TTFT, tokens/s, and concurrency, with the same scripts I used in my PoC.

How many KV cache tokens did your log show? Leave it in the comments, it's the number that defines your real capacity.


> More posts