[ Scrolls & chronicles ]Migration from Codea Bien
Blog
What I broke, what I measured and what I learned.
- OCT 2026AI #vllm #servir modelos de ia
What Serving an AI Model Actually Means (and Why vLLM Exists)
Serving an AI model means turning a file of weights into an API that answers questions. That's the whole job of an inference engineer, and this guide takes you from zero to your ow…
5 min // 250 xp
- OCT 2026AI #ollama #modelo local
Your First Local Model in 10 Minutes (Ollama)
Before renting a single GPU, let's run a model on your own machine. Free, no account, 10 minutes. The goal isn't production: it's understanding the full request-to-tokens cycle, wh…
4 min // 200 xp
- OCT 2026AI #runpod #alquilar gpu
Rent Your First GPU Without Fear (~$1 per hour)
It's time for real hardware. We'll rent a 48 GB GPU, run our first commands, and shut it down without leaving money behind. Total chapter budget: ~$3. You buy nothing.
3 min // 150 xp
- OCT 2026AI #vllm flags #gpu-memory-utilization
The 5 vLLM Flags That Actually Matter
vLLM has over a hundred configuration flags. Five decide how much context you hold, how much you spend, and whether your agent works. The rest is fine tuning for when you already m…
4 min // 200 xp
- OCT 2026AI #vllm serve #primer servidor vllm
Your First vLLM Server (hello world)
The pod is running. Now the important part: install vLLM, download a 30B model, and turn on your first inference server on the internet. It's 20 minutes of real work; the rest is d…
3 min // 150 xp
- OCT 2026AI #vllm troubleshooting #errores gpu alquilada
When It Explodes: The 7 Errors of Serving vLLM on a Rented GPU
Your server will fail. That's not pessimism, it's statistics: you rented an ephemeral container, downloaded 31 GB of weights, and started three services that met each other today.…
4 min // 200 xp
- OCT 2026AI #medir vllm #benchmark ttft
Measure Your vLLM: TTFT, tokens/s and Concurrency Without Fooling Yourself
A server that "works" says nothing. The engineer's question is how much: how long until the first token, how many tokens per second, how many people at once. This chapter leaves yo…
4 min // 200 xp
- OCT 2026AI #litellm gateway #virtual keys
LiteLLM: Keys, Budgets and Your vLLM's Bill
Your vLLM is on the internet with a single shared API key. That's fine for you alone. For a team it's a mess: nobody knows who spends what, and revoking someone's access means chan…
4 min // 200 xp
- OCT 2026AI #antigravity sdk #gemma 4
Antigravity SDK with Local Models: Your GPU, Your Tokens, Zero Cloud
At 12:14 AM on a Sunday, my content pipeline ran out of quota. It wasn't a bug. The API returned a 429 halfway through the batch and the job died with 18 posts half-generated. I sa…
9 min // 450 xp
- OCT 2026AI #stitch cli #google stitch
Stitch CLI: Google Brings UI Design to the Terminal (and to Your Agent)
Google has introduced Stitch CLI, a command-line tool (@google/stitch) that brings Stitch's generative UI design to the terminal and to your coding agent. The official promise: "yo…
11 min // 550 xp