Skip to content
Guide8/8 chapters

vLLM from zero: your first inference server

From zero to your own vLLM server on a rented GPU: 8 hands-on chapters, for under $5, with measured numbers at every step.

Guide: a closed list on a single topic, start to finish.

Updated:

Progress100%

A hands-on guide to go from zero inference knowledge to your own vLLM server on a rented GPU: for under $5, with measured numbers at every step.

What you'll build

  • A local model running on your machine with Ollama
  • A vLLM server with a 30B model in FP8 on a 48 GB GPU
  • A gateway with per-person keys, budgets, and logs
  • Your own TTFT, tokens/s, and concurrency benchmarks

8 chapters, each with commands that ran on a real pod and a time/cost checkpoint. Start at chapter 1 and don't skip steps: each one builds on the previous.

> View all guides and series