Building a homelab LLM router with LiteLLM and Headroom
TL;DR: Built a homelab LLM router out of two independent Docker Compose stacks: one for local model serving, one for routing. The routing stack, litellm-headroom/, runs a LiteLLM proxy that gives every client (agents, IDEs, scripts) one OpenAI-compatible endpoint that routes to local silicon, other homelab GPU boxes over Tailscale, or cloud APIs. A Headroom sidecar compresses only the requests actually headed to OpenAI/Anthropic, decided by where the request is actually going rather than whatever alias I happened to name the model. Postgres backs per-bot LiteLLM API keys so I can tell which agent is burning tokens and cap it before it blows a budget. Every secret lives in one Bitwarden note and reaches the containers as Docker Compose file-based secrets instead of environment: values docker inspect could read; direnv handles the Bitwarden unlock automatically so plain docker compose up still works. Full writeup and working compose files in homelab-llm-router. ...