Your First Local Model in 10 Minutes (Ollama)
Before renting a single GPU, let's run a model on your own machine. Free, no account, 10 minutes. The goal isn't production: it's understanding the full request-to-tokens cycle, which is exactly what vLLM will do later with serious hardware.
Ollama is the standard entry point to local models: it downloads weights, quantizes them, and serves them through an OpenAI-compatible API on localhost:11434. It's not for production, and chapter 4 will show you why. It's your lab.
"Ollama... Get up and running with large language models locally." - Source: Ollama on GitHub
Install and run your first model
# macOS
brew install ollama
# Linux
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
Now the first model. The sizing rule is simple: the model must fit in half your RAM. Start with a tiny one:
# terminal
ollama run qwen2.5:0.5b
That downloads ~400 MB and opens a chat in your terminal. Ask it something, watch the speed, exit with /bye.
The suffix is the scale: 0.5b = 500 million parameters, 7b = 7 billion. Now the big one:
# terminal
ollama run qwen2.5:7b
It downloads ~4.7 GB. The weights come quantized to 4 bits (GGUF format), which is why a 7B model fits in ~5 GB instead of the 14 GB it would take in FP16. I did the quantization math with numbers in Quantization with numbers.
The same language from chapter 1
Ollama is already serving an API on port 11434. And not just any API: the OpenAI-compatible one. Recognize this command:
# terminal
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen2.5:7b",
"messages": [
{ "role": "user", "content": "Explain in one sentence what a token is" }
],
"max_tokens": 60
}'
Same endpoint, same messages format, same usage as the OpenRouter curl from chapter 1. The backend changed and your client didn't notice. That decoupling is exactly what we'll exploit with vLLM.
Streaming: watching tokens arrive
Run the same request with "stream": true and -N (disables curl buffering):
# terminal
curl -N http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen2.5:7b","messages":[{"role":"user","content":"Count from 1 to 10"}],"stream":true,"max_tokens":60}'
You don't get one JSON at the end: you get dozens of data: {...} fragments, one per token. Each fragment carries a delta with the new piece of text. That's what ChatGPT does when it "types". In chapter 4 we'll measure when the first one arrives.
Measure your first timings
Keeping timings in your head doesn't work. This script measures TTFT (time to first token) and tokens/s against any OpenAI-compatible endpoint:
# measure.py
import json, sys, time
import requests
URL = "http://localhost:11434/v1/chat/completions"
MODEL = sys.argv[1] if len(sys.argv) > 2 else "qwen2.5:7b"
PROMPT = sys.argv[2] if len(sys.argv) > 2 else sys.argv[1]
def measure(prompt: str) -> None:
start = time.perf_counter()
ttft = None
tokens = 0
with requests.post(URL, json={
"model": MODEL,
"messages": [{"role": "user", "content": prompt}],
"stream": True,
"max_tokens": 200,
}, stream=True) as r:
for line in r.iter_lines():
if not line or not line.startswith(b"data: "):
continue
payload = line[6:]
if payload == b"[DONE]":
break
chunk = json.loads(payload)
delta = chunk["choices"][0]["delta"].get("content")
if delta:
if ttft is None:
ttft = time.perf_counter() - start
tokens += 1
total = time.perf_counter() - start
print(f"TTFT: {ttft*1000:6.0f} ms | tokens: {tokens:3d} | tok/s: {tokens/total:6.1f} | total: {total:.2f}s")
if __name__ == "__main__":
measure(PROMPT)
Note that I count fragments, not exact tokens: almost always 1 delta = 1 token, and for comparing models it's good enough. Run it on both models:
# terminal
pip install requests
python measure.py qwen2.5:0.5b "Explain what quantization is in 3 sentences"
python measure.py qwen2.5:7b "Explain what quantization is in 3 sentences"
My numbers on a MacBook Pro M4 Max, so you can compare against something real:
| Model | TTFT | Decode |
|---|---|---|
| qwen2.5:0.5b | ~1,000-1,200 ms | 308-326 tok/s |
| qwen2.5:7b | ~2,700-3,200 ms | 72-79 tok/s |
| qwen3-coder:30b | 4,007 ms cold / 34.9 ms warm | 91-93 tok/s |
Two readings: decode gets ~4× slower when you multiply parameters by 14, and TTFT depends on whether the model is already warm in memory (34.9 ms vs 4 seconds). The first request pays for the load; the rest fly.
What NOT to expect from Ollama
Ollama serves one request at a time and lives on your laptop. For you alone, perfect. For a team, no: no batching, no context memory management, no keys. When you want 10 people working at the same time, that's vLLM's job on a real GPU.
Checkpoint and the typical mistake
Checkpoint: $0, ~15 minutes. You have a model running, streaming understood, and your first 2 measurements.
The typical mistake: asking an 8 GB laptop to run a 7B model. RAM fills up, the system starts swapping, everything freezes. Respect the half-RAM rule; if your machine is small, stick with 0.5b or 3b and keep following the guide anyway.
If you want the full local path (LM Studio, MLX, when vLLM on Mac), I covered it in Local inference in 2026. And to not fool yourself with benchmarks, measuring inference without lying to yourself.
Next chapter
We rent your first GPU: RunPod account, the right L40S (and why NOT an A100), disks, SSH, and the golden rule for not burning money with the pod running.
What numbers did your two models give you? Drop them in the comments: comparing real hardware is the best part.