Skip to content

[ Scrolls & chronicles ]Migration from Codea Bien

Blog

What I broke, what I measured and what I learned.

  • AI #vllm #servir modelos de ia

    What Serving an AI Model Actually Means (and Why vLLM Exists)

    Serving an AI model means turning a file of weights into an API that answers questions. That's the whole job of an inference engineer, and this guide takes you from zero to your ow…

    5 min // 250 xp

  • AI #ollama #modelo local

    Your First Local Model in 10 Minutes (Ollama)

    Before renting a single GPU, let's run a model on your own machine. Free, no account, 10 minutes. The goal isn't production: it's understanding the full request-to-tokens cycle, wh…

    4 min // 200 xp

  • AI #runpod #alquilar gpu

    Rent Your First GPU Without Fear (~$1 per hour)

    It's time for real hardware. We'll rent a 48 GB GPU, run our first commands, and shut it down without leaving money behind. Total chapter budget: ~$3. You buy nothing.

    3 min // 150 xp

  • AI #vllm flags #gpu-memory-utilization

    The 5 vLLM Flags That Actually Matter

    vLLM has over a hundred configuration flags. Five decide how much context you hold, how much you spend, and whether your agent works. The rest is fine tuning for when you already m…

    4 min // 200 xp

  • AI #vllm serve #primer servidor vllm

    Your First vLLM Server (hello world)

    The pod is running. Now the important part: install vLLM, download a 30B model, and turn on your first inference server on the internet. It's 20 minutes of real work; the rest is d…

    3 min // 150 xp

  • AI #vllm troubleshooting #errores gpu alquilada

    When It Explodes: The 7 Errors of Serving vLLM on a Rented GPU

    Your server will fail. That's not pessimism, it's statistics: you rented an ephemeral container, downloaded 31 GB of weights, and started three services that met each other today.…

    4 min // 200 xp

  • AI #litellm gateway #virtual keys

    LiteLLM: Keys, Budgets and Your vLLM's Bill

    Your vLLM is on the internet with a single shared API key. That's fine for you alone. For a team it's a mess: nobody knows who spends what, and revoking someone's access means chan…

    4 min // 200 xp