TL;DR: Built a homelab LLM router out of two independent Docker Compose stacks: one for local model serving, one for routing. The routing stack, litellm-headroom/, runs a LiteLLM proxy that gives every client (agents, IDEs, scripts) one OpenAI-compatible endpoint that routes to local silicon, other homelab GPU boxes over Tailscale, or cloud APIs. A Headroom sidecar compresses only the requests actually headed to OpenAI/Anthropic, decided by where the request is actually going rather than whatever alias I happened to name the model. Postgres backs per-bot LiteLLM API keys so I can tell which agent is burning tokens and cap it before it blows a budget. Every secret lives in one Bitwarden note and reaches the containers as Docker Compose file-based secrets instead of environment: values docker inspect could read; direnv handles the Bitwarden unlock automatically so plain docker compose up still works. Full writeup and working compose files in homelab-llm-router.
This stack is backend-agnostic. It doesn’t care whether the local model behind it is served by vLLM, Ollama, or nothing at all if you’re only routing to cloud APIs. Mine happens to be a vLLM setup on an AMD Radeon AI PRO R9700, and that write-up covers the GPU/Docker side. Everything here is about the router in front of it.
Two independent Compose stacks, not one
The local model server and the routing layer live in separate Compose stacks, vllm/ and litellm-headroom/, in the same repo. They scale on different axes. A GPU box hosts exactly one model at a time; there’s no benefit to running two inference containers on a single card, just memory contention. The routing layer is CPU-bound and doesn’t need to live on the same machine as the GPU at all. It can run on the same box, on a separate always-on service host, or in front of several GPU machines. docker-compose assumes single-host, so keeping that flexibility meant two folders instead of one file:
vllm/— local model serving, bound to127.0.0.1:8000only. Copy the folder to any other homelab GPU box and it works, no changes beyond.env.litellm-headroom/— LiteLLM Proxy, the single OpenAI-compatible endpoint every client talks to, plus a Headroom compression sidecar and a Postgres database. LiteLLM’smodel_listmaps model names to backends: the local machine, other homelab GPU boxes, cloud APIs like OpenAI and Anthropic.
Because the two stacks can live on different physical hosts, the local backend is always addressed by its Tailscale MagicDNS hostname in litellm-config.yaml, never a docker-network container name. That’s true even when both stacks run on the same machine. There’s no same-host shortcut once they’re separate Compose projects. Exposure to the tailnet goes through tailscale serve --bg --https=443 http://localhost:4000 rather than publishing a raw port. Tailscale handles TLS, and only the LiteLLM proxy is reachable from other tailnet devices. The local backend itself never is.
Compressing only the calls that cost money
Once local inference and cloud API calls both flow through the same LiteLLM endpoint, compression is an obvious next move. Headroom shrinks tool outputs, logs, and JSON before they reach the model. That matters a lot for a coding agent’s cloud token bill and not at all for local inference on hardware I already own. Local inference isn’t billed per token, and compressing prior turns would actively hurt it by busting the local server’s own prefix cache and forcing full reprocessing on every request. So Headroom needed to run conditionally: compress requests headed to OpenAI/Anthropic, skip everything else.
The obvious way to decide that is pattern-matching the caller-facing model_name alias (gpt-4o, claude-sonnet, whatever). But those are strings I made up and could rename any time I feel like it. Better to key off the resolved outbound destination instead: litellm_params.model and api_base, which only get populated after LiteLLM’s router has already picked a real deployment for the request.
_CLOUD_MODEL_PREFIXES = ("gpt-", "openai/", "claude-", "anthropic/")
def _is_cloud_bound(litellm_params: dict) -> bool:
model = litellm_params.get("model", "") or ""
api_base = litellm_params.get("api_base", "") or ""
if model.startswith(_CLOUD_MODEL_PREFIXES):
return True
if "api.openai.com" in api_base or "api.anthropic.com" in api_base:
return True
return False
That lives in a LiteLLM CustomLogger callback (async_pre_call_hook), registered in litellm-config.yaml, that calls Headroom’s /v1/compress endpoint only when _is_cloud_bound comes back true. Every local backend passes through untouched no matter what I name its alias. It also fails open: if Headroom is unreachable, the request goes through uncompressed instead of blocking. A compression layer going down shouldn’t take down model access with it.
Headroom runs as its own sidecar container instead of an in-process LiteLLM dependency, same reason the two stacks got split. Independent lifecycle, and no risk of its ML-heavy dependency set (fastembed, ONNX, tree-sitter) fighting with LiteLLM’s own.
Per-bot API keys, backed by Postgres
The router serves more than one caller: several agents and automations, not just me typing curl commands by hand. That’s a real visibility problem. If everything shares one LITELLM_MASTER_KEY, there’s no way to tell which bot is burning tokens, and no way to cap a runaway one before it blows a monthly cloud budget. LiteLLM’s answer is virtual keys. POST /key/generate issues a distinct key per bot with its own max_budget, and /spend/logs / /key/info report usage per key.
Here’s the catch: that virtual-key store needs a real database to survive a restart, and LiteLLM’s Prisma-backed schema only supports Postgres. No SQLite option, even though Prisma technically supports SQLite elsewhere. So litellm-headroom/docker-compose.yml runs a third service, postgres:16-alpine, with its data directory bind-mounted to a host folder (./data/postgres, gitignored) instead of a Docker-managed volume. Same “just a directory I can back up” simplicity SQLite would’ve given me, just running as its own process instead of an embedded file. DATABASE_URL joined the Bitwarden-secrets pattern the same way every other secret in this stack does.
Secrets: Bitwarden, direnv, then Compose file-based secrets
Not .env. The obvious approach, API keys in a .env file loaded via env_file:, has two real problems: the file sits on disk in plaintext, and even shell-exported environment variables get exposed by docker inspect <container> once they land in a environment: block, regardless of where the value came from. Not great for OpenAI/Anthropic keys with live billing attached.
What ended up in the repo:
- One Bitwarden Secure Note, with custom fields
OPENAI_API_KEY,ANTHROPIC_API_KEY,LITELLM_MASTER_KEY,HEADROOM_PROXY_TOKEN, andPOSTGRES_PASSWORD. Not separate Login items, which is where I started before realizing Bitwarden’s Login item type forces every secret into an identically-namedpasswordfield. Confusing fast once you’ve got more than one. - direnv unlocks the vault, reads those fields, and exports them into the shell environment automatically whenever I
cdintolitellm-headroom/. A wrapper script would do the same job, but it means remembering to run a repo-specific command instead of thedocker compose upmuscle memory that works everywhere. direnv keeps that muscle memory intact. After onedirenv allow, plaindocker compose up -d/docker compose downwork exactly as documented anywhere, and Bitwarden’s master password/2FA gate is still there, just triggered oncdinstead of by name. This is the right call for now, while I’m still running the stack by hand during setup. It’s the wrong shape oncelitellm-headroom/lands on its own always-on host, since direnv only fires on an interactive shellcd. Nothing happens on boot, nothing restarts on crash, nojournalctl-style logging. The eventual fix is a systemd unit withExecStartPredoing the same Bitwarden fetch into a tmpfs-backedEnvironmentFile,ExecStart/ExecStopwrappingdocker compose up/down,WantedBy=multi-user.targetfor auto-start. Still haven’t figured out how the unit itself authenticates to Bitwarden without an interactivebw unlockprompt. That’s a problem for when the migration actually happens. docker-compose.ymldeclares them as Compose file-based secrets, sourced from that shell environment but mounted into the containers as read-only files at/run/secrets/*, not as container env vars.docker inspectshows nothing.- Small entrypoint shims (
scripts/litellm-entrypoint.sh,scripts/headroom-entrypoint.sh) run first inside each container, read those files, export them as env vars for that one process, then exec the real entrypoint. LiteLLM’s config format only knows how to read secrets viaos.environ/OPENAI_API_KEY, not from a file path, so something has to bridge the gap. The LiteLLM shim also assemblesDATABASE_URLfrom the Postgres password secret, since LiteLLM’s config wants one combined connection string, not separate host/user/password fields.
Actually running this against a live Bitwarden setup, instead of trusting the design on paper, turned up two real silent-failure bugs. Both in the “what happens when a secret can’t be fetched” path. Exactly the path that shouldn’t fail quietly.
- In
.envrcitself: a failedbw get itemcall (item doesn’t exist yet, wrong name, whatever) left the fetched JSON empty, and every downstream field-lookup call’sreturn 1only exited that shell function, not the whole.envrc. A missing Bitwarden item silently exported empty strings for every secret instead of stopping direnv from loading at all. Fixed by checkingbw get item’s exit code directly and propagating each field lookup’s failure with|| return 1instead of a barereturn 1inside the function. - In the container entrypoints: Compose’s
secrets:block has no equivalent of${VAR:?message}for the environment variable backing a secret. I confirmed this by testing it directly. IfOPENAI_API_KEY(or any of the others) is unset whendocker compose upruns, for any reason,.envrcfailed, or the stack got started from a shell that never sourced direnv (cron, a raw SSH command), Compose just creates the secret as an empty file. No warning.docker compose configstill exits0. Both entrypoint shims now check for a missing or empty secret file and exit1with a clear message before the process starts, instead of running LiteLLM or Headroom with a blank credential.
I verified the whole pipeline end to end rather than trusting any of it on paper: started the stack with fake keys, confirmed docker inspect litellm --format '{{json .Config.Env}}' showed nothing sensitive, confirmed the secret files existed under /run/secrets/, and confirmed a real chat completion request round-tripped through LiteLLM and correctly attempted to call OpenAI with the injected fake key, failing only because the key itself was fake. Then separately confirmed a missing Bitwarden item now fails loudly instead of quietly exporting blank secrets.
Where this leaves things
The repo (vllm/docker-compose.yml, litellm-headroom/docker-compose.yml, litellm-headroom/litellm-config.yaml, litellm-headroom/callbacks/headroom_conditional.py, litellm-headroom/.envrc, and the entrypoint shims) is at rayjanwilson/homelab-llm-router.
Next up: bringing other homelab GPU boxes online behind the same LiteLLM router. Each one just needs Tailscale, tailscale serve on its own local backend, and one more entry in litellm-config.yaml’s model_list. Then issuing real per-bot virtual keys as agents actually come online, and eventually the systemd migration once litellm-headroom/ lands on its own dedicated box. Separately, still want to get the NAS on one of those boxes onto Tailscale, though that’s unrelated to this stack (raw SMB/NFS over Tailscale directly, not through tailscale serve, which only proxies HTTP).