Skip to content
AI•3 min read

Your First vLLM Server (hello world)

The pod is running. Now the important part: install vLLM, download a 30B model, and turn on your first inference server on the internet. It's 20 minutes of real work; the rest is download wait.

We'll serve Qwen3-Coder-30B-A3B-Instruct-FP8, the same model from my PoC: a MoE with 30B total parameters and ~3B active per token (I explained how that works in MoE from the inside), in FP8 to fit the L40S's 48 GB, Apache-2.0 licensed.

1. Prepare the environment

Copy block by block into the pod's web terminal:

# pod terminal
pip install -q uv
uv venv /root/venv-vllm --python 3.12
source /root/venv-vllm/bin/activate
uv pip install vllm hf_transfer

vllm --version    # write it down: you'll use it to reproduce everything
nvidia-smi        # you should see 1x L40S with 48 GB

The venv lives in /root (container disk) because it needs real POSIX permissions; the weights will go to the volume. If the why is fuzzy, it's chapter 3's lesson.

2. Download the model (~31 GB)

# pod terminal
export HF_HUB_ENABLE_HF_TRANSFER=1
hf download Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8
du -sh /workspace/hf    # should be ~31 GB

HF_HOME already points to /workspace/hf thanks to chapter 3's .bashrc. The download takes 5-12 minutes depending on your network. hf_transfer speeds it up with parallel downloads.

3. Make up your API key

# pod terminal
echo 'export VLLM_API_KEY=sk-vllm-change-this-123' >> ~/.bashrc
source ~/.bashrc

Why protect a server you just created? Because the pod URL is public: anyone with your Pod ID can reach it. vLLM starts with --api-key, no way around it.

4. Start vLLM

# pod terminal
mkdir -p /workspace/poc && cd /workspace/poc

nohup vllm serve Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 \
  --served-model-name qwen3-coder-30b \
  --max-model-len 65536 \
  --gpu-memory-utilization 0.92 \
  --kv-cache-dtype fp8 \
  --enable-prefix-caching \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --api-key $VLLM_API_KEY \
  --host 0.0.0.0 --port 8000 \
  > /workspace/vllm.log 2>&1 &

tail -f /workspace/vllm.log

We dissect the flags in chapter 5; for now, three notes:

  • nohup ... & sends the process to the background with a log file, so it doesn't die when you close the browser.
  • --host 0.0.0.0 is what makes the server reachable from the public URL. Without it, 502.
  • --served-model-name defines the name clients will call the model by. It doesn't have to match the Hugging Face name.

Wait 3-5 minutes and leave tail with Ctrl-C when you see Uvicorn running on http://0.0.0.0:8000.

Write down the most important number in the log:

# pod terminal
grep -i "kv cache size" /workspace/vllm.log

On an L40S you should see between 150K and 200K tokens. That number is the context memory left after loading the weights, and it decides how many concurrent sessions you can hold. The per-token math is in 48 KB per token.

5. Test that it responds

# pod terminal
curl -s localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer $VLLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-coder-30b","messages":[{"role":"user","content":"Say hello in 3 words"}]}' \
  | head -c 400

If you see JSON with content, the model is serving. Now from your Mac, through the public URL:

# your machine
curl -N https://YOUR_POD_ID-8000.proxy.runpod.net/v1/chat/completions \
  -H "Authorization: Bearer sk-vllm-change-this-123" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-coder-30b","messages":[{"role":"user","content":"Count from 1 to 5"}],"stream":true}' \
  | head -20

You see the data: {...} fragments arrive one by one, same as Ollama. Same language, 48 GB of hardware on the other side.

Checkpoint and the typical mistake

Checkpoint: your own vLLM server, 30B parameters, reachable from the internet, speaking OpenAI-compatible. ~20 minutes of work, about $2 of GPU between download and startup.

The typical mistake: calling the model by its Hugging Face name (Qwen/Qwen3-...) instead of the --served-model-name (qwen3-coder-30b). The server answers "model not found" and it looks like everything broke, when it's just the alias.

Two more symptoms you'll see: the proxy URL gives 502 (the service hasn't finished starting, check vllm.log), or hf: command not found (you forgot source /root/venv-vllm/bin/activate).

Next chapter

The 5 flags in that command decide your capacity, your cost, and whether your agent works. We open them one by one, with the real number each one moves.

Did it run on the first try, or did a flag eat you? Tell me which one in the comments.


> More posts