Skip to content
AI•4 min read

LiteLLM: Keys, Budgets and Your vLLM's Bill

Your vLLM is on the internet with a single shared API key. That's fine for you alone. For a team it's a mess: nobody knows who spends what, and revoking someone's access means changing everyone's key.

The fix is a gateway. Between clients and your GPU we add LiteLLM: virtual keys per person, budgets, and every request logged with its owner. In my PoC, this was the piece that delivered the most value, even before any savings math.

   Your team (opencode)  ──https──►  pod
                                     ├─ LiteLLM  :4000  (keys, logs, budgets)
                                     │     └──localhost──►
                                     └─ vLLM     :8000   (the model)
                                          + Postgres :5432 (gateway data)

Nobody talks to the GPU directly. If tomorrow you swap vLLM for another backend, you change one config line and your team never notices.

1. Postgres (where keys live)

# pod terminal
apt-get update -qq && apt-get install -y -qq postgresql
service postgresql start
su - postgres -c "psql -c \"ALTER USER postgres PASSWORD 'litellm';\""
su - postgres -c "createdb litellm"

Without a database, LiteLLM starts in memory mode: keys vanish on restart and the UI doesn't work.

2. LiteLLM in its own venv

# pod terminal
uv venv /root/venv-litellm --python 3.12
source /root/venv-litellm/bin/activate
uv pip install 'litellm[proxy]'

Separate environments on purpose. LiteLLM and vLLM demand different versions of the same libraries; mix them in one venv and you break vLLM.

3. The configuration

# pod terminal
mkdir -p /workspace/poc && cat > /workspace/poc/litellm-config.yaml <<'EOF'
model_list:
  - model_name: qwen3-coder-30b
    litellm_params:
      model: openai/qwen3-coder-30b
      api_base: http://localhost:8000/v1
      api_key: os.environ/VLLM_API_KEY
      input_cost_per_token: 0
      output_cost_per_token: 0

litellm_settings:
  drop_params: true

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
EOF

Three decisions in the config: model_name is the alias your clients will use (it must match exactly what you put in opencode in step 5). api_base points to your local vLLM, which stays private. And per-token costs are 0 because your GPU charges by the hour, not per token; the real cost model comes at the end.

Also create the master key if you didn't in chapter 4:

# pod terminal
echo 'export LITELLM_MASTER_KEY=sk-master-change-this-456' >> ~/.bashrc
source ~/.bashrc

4. Start the gateway

# pod terminal
cd /workspace/poc
export DATABASE_URL="postgresql://postgres:litellm@localhost:5432/litellm"
export UI_USERNAME=admin
export UI_PASSWORD=change-this
export STORE_MODEL_IN_DB=true

nohup litellm --config litellm-config.yaml --host 0.0.0.0 --port 4000 \
  > /workspace/litellm.log 2>&1 &

tail -f /workspace/litellm.log

The first start takes 1-2 minutes preparing the database. Leave tail with Ctrl-C when you see Uvicorn running on http://0.0.0.0:4000.

5. Create your first virtual key

# pod terminal
curl -s -X POST localhost:4000/key/generate \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"key_alias":"your-name","models":["qwen3-coder-30b"],"max_budget":5}'

It returns {"key":"sk-..."}. That's one person's key, with a $5 budget. For another person, the same command with another alias: each one shows up separately with their tokens and spend, and gets revoked in seconds.

Open the UI at https://YOUR_POD_ID-4000.proxy.runpod.net/ui (login with UI_USERNAME/UI_PASSWORD): that's where Virtual Keys, Logs (each request with its owner, tokens and latency) and Usage live.

6. The test that proves everything

# WITHOUT key → should answer 401
curl -s -o /dev/null -w "%{http_code}\n" -X POST \
  https://YOUR_POD_ID-4000.proxy.runpod.net/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-coder-30b","messages":[{"role":"user","content":"hi"}]}'

# WITH key → should answer 200
curl -s -o /dev/null -w "%{http_code}\n" -X POST \
  https://YOUR_POD_ID-4000.proxy.runpod.net/v1/chat/completions \
  -H "Authorization: Bearer sk-YOUR-KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-coder-30b","messages":[{"role":"user","content":"hi"}]}'

401 without key, 200 with key. Your GPU is no longer public.

7. Connect your agent

On your machine, point opencode (or any OpenAI-compatible client) at the gateway:

// ~/.config/opencode/opencode.json
{
  "provider": {
    "poc": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "PoC RunPod",
      "options": {
        "baseURL": "https://YOUR_POD_ID-4000.proxy.runpod.net/v1",
        "apiKey": "sk-YOUR-VIRTUAL-KEY"
      },
      "models": {
        "qwen3-coder-30b": {
          "name": "Qwen3-Coder 30B FP8 (L40S)",
          "limit": { "context": 65536, "output": 8192 }
        }
      }
    }
  },
  "model": "poc/qwen3-coder-30b"
}

The limit.context: 65536 isn't decorative: it tells the client to compact the context before hitting the server's limit. Without it, the session dies when it grows.

The cost conversation

Now, the question that holds this guide up: when does this make sense?

  • Your GPU charges by the hour (~$1.09/h ≈ $796/month always on), not per token.
  • Subscriptions charge per person (~$200/month each).
  • Self-hosted cost is per capacity: one server serves many, so cost per person drops as the team grows.

With my numbers: at 10 engineers the savings are marginal; at 30 it reaches 24-37% (with reserved contract); at 100, 50-62%. And there's an argument that wins even with zero savings: knowing who uses what, with how many tokens and what result. My formal recommendation to my company was don't migrate yet, and why, in The report that said "don't migrate yet".

The missing piece for the full case is routing: the cheap and repetitive to your vLLM, the hard to the premium model. I covered it in Model routing and costs.

Checkpoint and the typical mistake

Checkpoint: a gateway with per-person keys, budgets and logs; your agent connected to your own GPU. Plus the cost table to decide if it's worth it on your team.

The typical mistake: a model_name mismatch between LiteLLM's config and your client. The error says "model not found" and it looks like the server, when it's an alias not matching. Second classic: installing LiteLLM in vLLM's venv.

Next chapter (the last one)

Everything we built will fail at some point. The final chapter is the rescue runbook: the 7 errors I hit, their exact symptoms and their fixes.

Would you put a gateway on a single-person server? It changed my whole sales pitch. Tell me in the comments.


> More posts