Skip to content
AI•5 min read

What Serving an AI Model Actually Means (and Why vLLM Exists)

Serving an AI model means turning a file of weights into an API that answers questions. That's the whole job of an inference engineer, and this guide takes you from zero to your own vLLM server for under $5.

This is chapter 1 of the guide. We install nothing yet: first the full map, because every mistake you'll make later lives in one of these pieces.

A model is a file, not a program

When you download "a 7B model", you download a folder with billions of numbers: the weights. They come as multi-GB safetensors files, plus the tokenizer and a config.json describing the architecture.

That file does nothing on its own. It's closer to a library (.dll / .so) than to an executable: it needs a program that loads it into memory, feeds it text, and reads what comes back. That program can be a 20-line script or a production server.

And here comes the first classic misconception: "7B" is not 7 GB. It's ~7 billion parameters. In FP16 each parameter takes 2 bytes, so that model is ~14 GB of weights alone. We'll do the exact math in chapter 5, but hold on to one idea: memory is the boss.

Tokens: the only currency

The model doesn't see words. It sees tokens: chunks of text averaging 3-4 characters. "inference" may be 2 tokens; "Tlaxcala", 4. And the model does exactly one thing, millions of times:

Given this text, what is the most likely next token?

Generating a response is a loop: predict a token, append it to the context, predict the next one. Everything we measure in this guide (TTFT, tokens/s) is a measurement of that loop.

Loading weights into a script is not serving

Everyone's first attempt looks like this:

# naive_serve.py
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct")

while True:
    prompt = input("> ")
    inputs = tokenizer(prompt, return_tensors="pt")
    output = model.generate(**inputs, max_new_tokens=200)
    print(tokenizer.decode(output[0]))

It works. And it's perfect for learning why it doesn't work:

  • One request at a time. While one generates, the others wait with the GPU paid and idle.
  • No standard API. Every client would have to speak "my homegrown stdin protocol".
  • No streaming. The user stares at an empty screen for 20 seconds.
  • No memory management. Every request re-processes its full prompt from scratch.

An inference server solves those four things: a request queue, batching (packing several prompts into the same GPU pass), token-by-token streaming, and an API that is already the de facto standard.

The API that changed everything: OpenAI-compatible

In 2023 OpenAI defined the format of its API (/v1/chat/completions, messages, choices, usage) and the rest of the world adopted it. Today almost everything speaks it: opencode, Cursor, LangChain, your 10-line script.

And this has a very practical consequence: if your server speaks OpenAI-compatible, you can swap the model or the backend and no client notices. The chapter 7 example (a gateway with per-person keys) only exists because of this.

Why vLLM and not another one

An inference server manages the most expensive resource in the system: the memory taken by each request's context. vLLM (born at UC Berkeley in 2023) won that match with a technique called PagedAttention: it treats context memory like pages, the same way an OS manages RAM, instead of reserving giant contiguous blocks that go to waste.

"vLLM achieves near-zero waste in KV cache memory... and delivers up to 24x higher throughput compared to HuggingFace Transformers." - Source: vLLM blog, PagedAttention

The 24× was the original paper's number; nobody promises that against everything today, but the idea won: vLLM is the de facto open source standard for serving LLMs, Apache-2.0 licensed, with support for MoE, FP8, and the models we'll use in this guide.

"vLLM is a fast and extensible library for LLM inference and decoding." - Source: vLLM documentation

There are serious alternatives (TensorRT-LLM, SGLang, llama.cpp, Ollama). In chapter 2 we'll use Ollama precisely because it's the easiest entry point; vLLM is the next level: the one you size, measure, and put in production.

Your first contact: reading a real response

Nothing to install. You only need a free account at OpenRouter (adding a card is optional for :free models), an API key, and curl. Any model with the :free suffix from openrouter.ai/models works:

# terminal
curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.2-3b-instruct:free",
    "messages": [
      { "role": "user", "content": "Explain in one sentence what an inference server is" }
    ],
    "max_tokens": 60,
    "temperature": 0.2
  }'

The response (trimmed) looks like this:

{
  "model": "meta-llama/llama-3.2-3b-instruct:free",
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "An inference server exposes an AI model through an API to process requests and return predictions."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 24,
    "completion_tokens": 21,
    "total_tokens": 45
  }
}

Note I use max_tokens, not max_input_tokens: that doesn't exist. That parameter limits the output; the input is limited by the model's context window. And temperature: 0.2 makes the response nearly deterministic, ideal for testing.

The fields that actually matter:

  • choices[0].message.content: the answer. Obvious, but it's the only field most people read.
  • finish_reason: "stop" means the model finished on its own; "length" means you cut it with max_tokens. Telling them apart saves you hours of "the model cuts itself off".
  • usage: the bill. prompt_tokens is what your question cost; completion_tokens, the answer.

Checkpoint and the typical mistake

Checkpoint: ~10 minutes, $0. You can now read an OpenAI-compatible response field by field, which is 80% of what we'll do with vLLM in chapters 4-7.

The typical mistake at this level: ignoring usage. Every turn of a conversation re-sends the entire history, so prompt_tokens grows with each question. On an API you pay per token; on your own server it becomes memory. I took that mechanism apart in 48 KB per token and it's the technical reason behind half a chapter of this guide.

If any term slipped by (KV cache, MoE, FP8), I keep an AI dictionary in Spanish for that. And if you want the full discipline view, what Inference Engineering is is the deep read.

Next chapter

We install Ollama, run your first model on your machine, for free, and measure your first response times. For the first time you'll see with your own eyes how long a token takes.

Did you already know the OpenAI-compatible format, or had you been using it without knowing it was a standard? Tell me in the comments.


> More posts