Guide8/8 chapters
Inference Engineering: Zero To Production
Correr modelos de verdad: local, cuantización, VRAM, serving, medición y costos.
Guide: a closed list on a single topic, start to finish.
Updated:
Progress100%
Local inference in 2026: from LM Studio to vLLM (with VRAM math)
Chapter 02 // 3 min
Quantization with numbers: quality vs VRAM vs speed
Chapter 03 // 3 min
What hardware do you need? GPU, VRAM and cloud
Chapter 04 // 3 min
Serving in production: batching, concurrency and autoscaling
Chapter 05 // 3 min
The real cost of your AI feature: API vs self-host vs edge
Chapter 07 // 3 min
Serving multiple models in production: routing, versioning and canary
Chapter 08 // 3 min