Rent Your First GPU Without Fear (~$1 per hour)
It's time for real hardware. We'll rent a 48 GB GPU, run our first commands, and shut it down without leaving money behind. Total chapter budget: ~$3. You buy nothing.
The provider will be RunPod, but what you learn here works the same on Lambda, Vast.ai, or whoever: they all sell the same thing, a container with a GPU and terminal access.
What a "pod" actually is
A pod is a container with a GPU inside. You rent it by the hour, it starts in 1-2 minutes, and you access it through a web terminal or SSH. It has two disks you must distinguish from minute one:
| Disk | What goes there | Survives restart |
|---|---|---|
| Container disk (ephemeral) | Binaries, venvs, install caches | NO |
| Volume (persistent) | Model weights, your configs | YES |
The most expensive lesson of this chapter lives in that table. RunPod pods are containers without Docker inside: everything installs natively.
Picking the right GPU
For this guide you need an Ada or newer GPU, because the model we'll serve uses FP8 and Ampere cards don't support it in hardware.
| GPU | VRAM | Works? |
|---|---|---|
| L40S | 48 GB | Yes, the recommended one |
| L40 / RTX 6000 Ada | 48 GB | Yes, valid alternative |
| H100 | 80 GB | Yes, way too expensive to learn |
| A40 / A100 / A6000 | 40-80 GB | NO: Ampere, no FP8 |
I paid the L40S at ~$1.09/h. Prices move; check the listing that day and write it down in your runbook.
Create the pod, step by step
- Create an account at runpod.io and add $20 in Billing. Twenty dollars runs this whole guide twice.
- Pods → Deploy.
- GPU: L40S, quantity 1.
- Template: RunPod PyTorch 2.x (the one with CUDA 12.x).
- Open Edit Template and set:
Container Disk: 30 GB
Volume Disk: 80 GB → Volume Mount Path: /workspace
Expose HTTP Ports: 8000,4000
Expose TCP Ports: 22
Ports must be declared now. Adding them later forces you to recreate the container, which wipes the container disk with all your venvs. The /workspace volume survives, but reinstalling is wasted time. If you ever want dashboards, also expose 3000,9090.
- Deploy On-Demand → wait for Running (1-2 min).
- Connect → Start Web Terminal. It's a terminal in your browser, zero setup.
Your public URL and your Pod ID
On the pod card you'll find your Pod ID (something like a1b2c3d4e5). Your services get public URLs:
https://<POD_ID>-8000.proxy.runpod.net ← whatever you expose on 8000
https://<POD_ID>-4000.proxy.runpod.net ← whatever you expose on 4000
Security note: anyone with that ID can reach those URLs. That's why in chapter 4 the server starts with --api-key and in chapter 7 we add per-person keys. Security isn't an extra, it's the architecture.
Check Connect and confirm you see HTTP Service [Port 8000] and [Port 4000]. The PyTorch template only exposes 8888 (Jupyter): if that's all you see, the ports weren't declared. Plan B without recreating: kill Jupyter (pkill -f jupyter) and use 8888 for your service.
Prepare the disk (the expensive lesson)
# pod terminal
df -h / /workspace
Rule of this guide: weights on /workspace, binaries on the container disk. Two reasons I learned by breaking it:
- Network volumes don't support
chmod: installing there blows up withOperation not permitted (os error 1). - On pod restart,
/root(container disk) gets wiped and/workspacesurvives. If the weights are on the volume, you don't re-download them (~10 minutes you don't want to pay twice).
Save the variables in .bashrc, which lives in /workspace and therefore persists:
# pod terminal
echo 'export HF_HOME=/workspace/hf' >> ~/.bashrc
source ~/.bashrc
The golden money rule
| State | What it charges |
|---|---|
| Pod Running | Full GPU per hour (~$1.09/h = ~$26/day if you forget it) |
| Pod Stop | Volume only: ~$8/month for 80 GB |
| Pod Destroy | Nothing. Re-downloading the model: ~10 min |
- Coming back this week → Stop.
- Not coming back this week → Destroy without guilt. The cloud is disposable, that's the point.
Checkpoint: pod Running, ~$3 spent, and you know how to get in, check disks, and shut down without burning money. The typical mistake at this level isn't technical: it's leaving the pod running overnight.
I covered the 7 errors of rented environments in The 7 bugs of the rented environment, and the VRAM-per-price analysis in What hardware do you need?.
Next chapter
We install vLLM inside the pod, download a 30B model, and turn on your first inference server on the internet.
Renting already? Tell me which GPU and at what price you found it, prices move more than you think.