Skip to content
AI•4 min read

Measure Your vLLM: TTFT, tokens/s and Concurrency Without Fooling Yourself

A server that "works" says nothing. The engineer's question is how much: how long until the first token, how many tokens per second, how many people at once. This chapter leaves you measuring your vLLM with the same scripts I used in my PoC, plus the lesson I learned by getting it wrong.

The 4 numbers that matter

MetricWhat it measuresWhat it means for the user
TTFTTime to first tokenHow "instant" the response feels
Decode tok/sGeneration speedFluidity; above ~30 tok/s it reads comfortably
Prefill tok/sSpeed processing the promptHow much sending huge context costs
Aggregate throughputTokens/second for ALL usersThe server's real capacity

And one reading rule: look at p50 and p90, not averages. The average hides the user who waits twice as long.

1 user: benchmark_openai.py

This script measures one request end to end against any OpenAI-compatible endpoint:

# pod terminal
python benchmark_openai.py \
  --base-url http://localhost:8000/v1 \
  --model qwen3-coder-30b \
  --prompt "$(cat prompt_largo.txt)" \
  --api-key $VLLM_API_KEY \
  --out results_1user.csv

The CSV it produces has one row per request with these columns:

ttft_ms,total_s,tok_s,prompt_tokens,output_tokens,ok,error
2541.2,3.31,120.3,15269,93,True,

That row is real, from my PoC: 15,269 prompt tokens processed and 93 output tokens in 3.31 seconds, with a TTFT of 2,541 ms. Sit with that ugly TTFT for a moment, because it's trap number one of this chapter and we uncover it in the network section.

Note that the prompt is long (15K tokens). If you measure with "hi", everything is instant prefill and the numbers are noise. Use a prompt the size of your real workload.

Many users: load_test.py

A 1-user benchmark doesn't answer the business question: how many can it hold at once?

# pod terminal
python load_test.py \
  --base-url http://localhost:8000/v1 \
  --model qwen3-coder-30b \
  --concurrency 10 \
  --context-tokens 16000 \
  --max-tokens 300 \
  --api-key $VLLM_API_KEY \
  --session-mode unique

My real results on the L40S (through the gateway), so you have a floor to compare against:

ScenarioTotal timeTTFT p50tok/s per userAggregate
1 user, 16K3.3 s2,541 ms120.3120 tok/s
5 users, 16K9.1 s5,435 ms27.3166 tok/s
10 users, 16K (unique sessions)16.1 s9,652 ms15.9199 tok/s
10 users, 16K (shared prefix)5.2 s3,773 ms73.6740 tok/s

The reading: aggregate throughput climbs with concurrency (199 tok/s), but per-user tok/s drops (15.9 each). The GPU doesn't do magic, it splits. If your SLO is "40 tok/s per person at peak", that table tells you how many fit.

The bug that inflates your numbers

The 740 tok/s row is a trap, and also a preview. My first load_test.py sent the same prompt to every user. With prefix caching on, vLLM processed that context ONCE and all 10 sessions reused it: I was measuring the memory of one session multiplied by magic.

The fix was --session-mode unique (a distinct marker per session), and the honest 16.1 s row came out of it. The full post-mortem is in The error in my own benchmark.

The nuance that makes it interesting: on a real team, agents DO share prefixes (the same giant system prompt). So the "shared prefix" scenario isn't a trap if your workload looks like that; it's a trap when you measure to decide capacity and your real workload shares nothing.

The network trap

That 2,541 ms TTFT with 1 user was absurd for an L40S. Measuring segment by segment, I found this:

SegmentTime
The model responds56 ms
The gateway adds15 ms
The public network to the pod~1,000 ms

94% of the wait was the internet, not the system. If you measure against the pod's public URL, you're measuring RunPod, not vLLM. To compare hardware, measure through an SSH tunnel or inside the pod. The full case, with methodology, is in The bottleneck wasn't the GPU: it was the network.

Checkpoint and the typical mistake

Checkpoint: your own 1-user CSV and a 10-user one, with your real prompt. You can now answer the question that opens this chapter with your own numbers.

The typical mistake: benchmarking with a 10-token prompt and a single run. Short prompt = you're measuring Python's startup, not the GPU. One run = you're measuring luck. Three runs minimum, prompt at real size, and look at p90.

For the theory of why percentiles and not averages, I have Measuring inference without lying to yourself. And vLLM exposes all of this in /metrics, Prometheus format: curl localhost:8000/metrics | grep -i ttft.

Next chapter

Your server serves. Now we turn it into a team service: per-person keys, budgets, per-request logs, and the cost conversation you need to have before going to production.

What numbers did your 10 concurrent users give you? That's the number that decides everything ahead. Leave it in the comments.


> More posts