Skip to content
Guide8/8 chapters

My First Self-Hosted Inference Server

The complete PoC: what I measured, what I broke and what it cost to serve an open-source model on a rented GPU.

Guide: a closed list on a single topic, start to finish.

Updated:

Progress100%

Serving your own model stopped being a lab topic: today you can rent a GPU for under a dollar an hour and an open 30B model fits in 48 GB of VRAM. The question is no longer is it possible, but is it worth it for you.

This series is the complete PoC we ran to answer that with numbers: two engineers, one rented L40S, a MoE code model and one rule — nothing gets decided without measuring. Eight chapters with the stack (opencode -> LiteLLM -> vLLM), the hardware and its real costs, the bugs that cost us hours, the latency and throughput measurements, the memory math and the final report that —spoiler— recommended not migrating yet, math in hand.

You'll find: setups you can copy, the 7 bugs of a rented environment with symptom, cause and fix, and real numbers of our own (TTFT, tokens/s, VRAM and cost) instead of someone else's benchmarks. And also what we did not measure: model quality, private-network latency and behavior with 30 concurrent users. That honesty is part of the value: a PoC that says "not yet" saves you an expensive decision.

Start with chapter 1 (the plan) and follow the order: the series is meant as a journey, from "can we?" to "now what?".

> View all guides and series