When It Explodes: The 7 Errors of Serving vLLM on a Rented GPU
Your server will fail. That's not pessimism, it's statistics: you rented an ephemeral container, downloaded 31 GB of weights, and started three services that met each other today. The difference between losing 10 minutes and losing an afternoon is knowing the symptoms.
These 7 errors happened to me, in a real environment, serving this guide's model. Each one comes with its exact error message, its cause and its fix, so next time the diagnosis takes minutes.
1. Operation not permitted (os error 1) when installing
Symptom: installing vLLM or any chmod blows up with that error.
Cause: you installed on the network volume (/workspace). Network volumes don't support POSIX operations like chmod.
Fix: binaries and venvs on the container disk (/root/venv-vllm), caches included (UV_CACHE_DIR=/root/.cache/uv). Weights on the volume. The full rule is in chapter 3.
2. config.json is not a valid JSON (or weights at 0 bytes)
Symptom: the download "finishes" but the model doesn't load, or the error mentions invalid JSON.
Cause: the network volume doesn't preserve the symlinks Hugging Face's cache uses internally. Some files ended up as broken or empty links.
Fix: download again with the modern CLI (hf download ...) and verify with du -sh that it weighs what's expected (~31 GB in our case). If your CLI version allows it, download to a flat folder with --local-dir.
3. vLLM dies at startup with no apparent reason
Symptom: vllm serve starts and crashes, or a fresh start fails with odd memory errors.
Cause: an orphaned VLLM::EngineCore process from a previous attempt is still holding the VRAM. The pod is a container: when you kill the parent, sometimes the child survives.
Fix:
# pod terminal
pkill -9 -f VLLM::EngineCore
nvidia-smi # until you see 0 MiB used
Don't relaunch until nvidia-smi shows free VRAM. Relaunching on top of a zombie is the recipe for a confusing OOM.
4. address already in use on port 8000 or 4000
Symptom: the service won't start because the port is taken.
Cause: almost always the template's Jupyter (not 8888, but if you reassigned it) or an old LiteLLM you didn't kill properly.
Fix: identify and kill the process holding the port. And if the missing port isn't even exposed, that's chapter 3's error: ports are declared when creating the pod.
5. /key/generate answers Not Found
Symptom: LiteLLM runs, but creating keys returns 404.
Cause: LiteLLM started without its database layer (chapter 7, step 1), or the virtual key endpoint doesn't exist because you installed the base package.
Fix: install the full extra (litellm[proxy,extra_proxy]), check service postgresql status and that DATABASE_URL is exported before starting.
6. prisma-client-py: not found in LiteLLM
Symptom: LiteLLM won't start, complaining about Prisma.
Cause: the Prisma client was generated with an absolute path from install time.
Fix: export the venv to the PATH before starting: export PATH=/root/venv-litellm/bin:$PATH. It's the same venv you activated, but LiteLLM spawns subprocesses that don't inherit the activation.
7. The LiteLLM UI looks dead (empty or redirecting to an internal IP)
Symptom: the UI loads broken or tries to redirect to an IP that isn't yours.
Cause: RunPod's proxy serves the pod under a different host, and LiteLLM generates URLs with its internal host without trusting the proxy.
Fix: adjust the asset prefix (/litellm-asset-prefix) and export FORWARDED_ALLOW_IPS='*' before starting the gateway.
Bonus: the two memory and hardware errors
| Message | Cause | Fix |
|---|---|---|
OOM / No available memory for the cache blocks | You promised more KV cache than is left | --gpu-memory-utilization 0.88 or --max-model-len 32768 |
Errors mentioning fp8 at startup | Ampere GPU without FP8 support | Drop --kv-cache-dtype fp8; you lose half the context, but it starts |
The general diagnostic method
Behind the 7 there's a method worth more than the fixes:
- The log first.
tail -f /workspace/vllm.logbefore touching anything. 80% of the answers are written there. nvidia-smialways. VRAM at 0 when you thought everything died is the signature of the zombie process.- One change at a time. If you change three things and it works, you don't know which one. If you change three and it breaks, neither.
- Write down the exact symptom. The literal error message is what saves you next time. My list of 7 became a runbook, an idempotent startup script, and two posts.
Checkpoint: your own runbook.md with the symptoms YOU have seen. Start today, with this chapter as the seed.
The typical mistake at this level: redeploying blindly (destroying the pod and starting from zero) as soon as something fails. Sometimes that's right (the cloud is disposable), but destroying without writing down the cause is buying the same lesson twice.
Guide wrap-up
You made it to the end: you understood what serving a model is, ran one locally, rented a GPU, started vLLM, configured its flags, measured with real numbers, put a gateway with keys and budgets in front, and survived the errors. That is, in small, the daily job of an inference engineer.
My real PoC posts go deeper into each measurement: start with The plan: serving my own model to cut the bill and follow the series from there.
What error hit you that isn't on the list? Leave it in the comments: this chapter gets updated with readers' cases.