Plan: serving my own model to cut the bill
Chapter 01 // 3 min
The complete PoC: what I measured, what I broke and what it cost to serve an open-source model on a rented GPU.
Guide: a closed list on a single topic, start to finish.
Updated:
Serving your own model stopped being a lab topic: today you can rent a GPU for under a dollar an hour and an open 30B model fits in 48 GB of VRAM. The question is no longer is it possible, but is it worth it for you.
This series is the complete PoC we ran to answer that with numbers: two engineers, one rented L40S, a MoE code model and one rule — nothing gets decided without measuring. Eight chapters with the stack (opencode -> LiteLLM -> vLLM), the hardware and its real costs, the bugs that cost us hours, the latency and throughput measurements, the memory math and the final report that —spoiler— recommended not migrating yet, math in hand.
You'll find: setups you can copy, the 7 bugs of a rented environment with symptom, cause and fix, and real numbers of our own (TTFT, tokens/s, VRAM and cost) instead of someone else's benchmarks. And also what we did not measure: model quality, private-network latency and behavior with 30 concurrent users. That honesty is part of the value: a PoC that says "not yet" saves you an expensive decision.
Start with chapter 1 (the plan) and follow the order: the series is meant as a journey, from "can we?" to "now what?".
Chapter 01 // 3 min
Chapter 02 // 3 min
Chapter 03 // 3 min
Chapter 04 // 3 min
Chapter 05 // 3 min
Chapter 06 // 3 min
Chapter 07 // 2 min
Chapter 08 // 3 min