TL;DR: Replaced Ollama with vLLM on an AMD Radeon AI PRO R9700 (gfx1201), running in Docker via a community image built specifically for this chip. The hard part wasn’t the model, it was getting vLLM to run as a non-root user instead of root. Three separate, unrelated permission traps had to be fixed together before it would even start: numeric vs. named GPU group IDs, an unreadable /root, and Docker silently root-owning auto-created bind-mount directories. Working docker-compose.yml and setup script in homelab-llm-router.

Why replace Ollama

Ollama’s fine for casual local model use, but it doesn’t give you an OpenAI-compatible proxy layer, doesn’t expose the serving control vLLM does (continuous batching, tensor parallelism, quantization backends), and doesn’t fit a setup where I want one endpoint routing between local GPUs, other homelab machines, and cloud APIs. vLLM does all three, and it’s what I’m switching to.

The hardware problem: gfx1201 is not a first-class citizen yet

The R9700 is RDNA4, reporting as gfx1201 under ROCm. AMD’s official rocm/vllm-dev images target MI300/MI350 first. gfx1201 support exists but lags, and there are real regressions reported on the community forums. Rather than fight the official image, I used kyuz0/vllm-therock-gfx1201, a community image built and actively maintained specifically for this chip. It includes AITER unified-attention patches that fix bugs stock vLLM hits on RDNA4.

Getting vLLM to run as a non-root user

The upstream vLLM image runs as root by default. No USER directive. That means every model weight downloaded into the mounted Hugging Face cache ends up owned by root on the host, which is annoying the moment you want to browse or manage that cache yourself. I figured this would be a one-line user: "1000:1000" in the compose file. It was not. Three separate, unrelated permission problems had to be solved together, and I only found all three by actually running it against the real hardware instead of reasoning about it in the abstract.

Trap 1: GPU device groups have to be numeric GIDs, not names

/dev/kfd and /dev/dri/*, the GPU device nodes, are owned by the host’s video and render groups. The obvious fix is group_add: [video, render]. That’s wrong. Those group names resolve to different GIDs inside the container image than on the host:

host:      video=983  render=987
container: video=39   render=105

As root none of this matters, since root bypasses permission checks entirely, so the bug stays invisible until you actually try to drop privileges. Linux permission checks are GID-number-based, not name-based. Fix is adding the host’s actual numeric GIDs as supplementary groups:

group_add:
  - "${VIDEO_GID}"   # host's numeric GID for `video`, not the string "video"
  - "${RENDER_GID}"  # same for `render`

vllm/scripts/up.sh resolves these automatically with getent group video/render before every run, so it’s not something you hardcode once and hope stays valid forever.

Trap 2: /root is unreadable to non-root, even with a correctly-owned mount inside it

The natural place to mount the Hugging Face cache is /root/.cache/huggingface, matching where the image’s default $HOME points. Running the container as UID 1000 hit this immediately:

OSError: PermissionError at /root/.cache/huggingface/token when downloading Qwen/Qwen2.5-0.5B-Instruct.

The mounted directory itself was correctly owned by UID 1000. The problem was one level up:

$ docker exec vllm-r9700 stat /root
Access: (0550/dr-xr-x---)  Uid: (0/root)  Gid: (0/root)

/root in this image is 0550, root-only, no group or other access at all. A non-root user can’t even traverse into /root to reach what’s mounted inside it, no matter what permissions that inner directory has. The fix was mounting everything under /home/vllm instead, which is 0755 and world-traversable. A completely different directory tree, not just a different owner.

Trap 3: Docker auto-creates missing bind-mount directories as root

Switching to /home/vllm fixed the Hugging Face download, but the container still crashed, this time in a completely different subsystem: Triton’s JIT kernel compiler cache.

PermissionError: [Errno 13] Permission denied: '/.triton'

Setting HOME=/home/vllm explicitly (needed anyway, the image doesn’t set a sane $HOME for an arbitrary non-root UID) fixed that path, but revealed a second layer of the same bug right behind it:

PermissionError: [Errno 13] Permission denied: '/home/vllm/.cache/vllm'

Here’s what was actually going on. /home/vllm was never itself an explicit bind mount. Only /home/vllm/.cache/huggingface, a subpath inside it, was mounted. Docker auto-creates missing intermediate directories for a bind mount, but it creates them owned by root, regardless of what user the container runs as. So /home/vllm existed, but was root-owned. Nothing else that wanted to write there, Triton’s ~/.triton cache, vLLM’s own ~/.cache/vllm model-info cache, could touch it, even though the one specific subpath explicitly mounted worked fine.

Two-part fix:

  1. Mount an actual host directory at /home/vllm itself, not just a subpath inside it, so the whole tree has correct ownership from the start.
  2. Pre-create that host directory (and its .cache subdirectory, itself an intermediate path for the nested Hugging Face mount) with mkdir -p, owned by the invoking user, before docker compose up ever runs. That way Docker never gets the chance to auto-create any part of it as root in the first place.
# vllm/scripts/up.sh, before `docker compose up`
VLLM_HOME_DIR="${VLLM_HOME_DIR:-$HOME/.cache/homelab-llm-router/vllm-home}"
mkdir -p "$VLLM_HOME_DIR/.cache"
# vllm/docker-compose.yml
environment:
  - HOME=/home/vllm
  - HF_HOME=/home/vllm/.cache/huggingface
volumes:
  - ${VLLM_HOME_DIR:-${HOME}/.cache/homelab-llm-router/vllm-home}:/home/vllm
  - ${HF_CACHE_DIR:-${HOME}/.cache/huggingface}:/home/vllm/.cache/huggingface

What made this hard to debug

Each trap threw a completely different-looking error, from a different subsystem, at a different point in startup. A subprocess pydantic.ValidationError masking a swallowed inner traceback. A bare PermissionError on a dotfile. Then the same error again, one layer deeper. None of them mentioned Docker, mounts, or root. What worked every time was the same instinct: stop guessing and actually reproduce the failure in isolation. For trap 3 that meant calling vLLM’s internal model-registry inspection function directly in a one-off Python invocation inside the container, rather than trying to read tea leaves out of vLLM’s full startup log.

I verified the final result the same way, not by reading the compose file and calling it correct. Deleted the model cache, ran the whole stack fresh through actual docker compose up, confirmed zero errors in the logs, sent a real inference request and got a real generated response back, and checked on the host that every downloaded weight file was owned by my own user, not root.

Where this leaves things

vLLM is up, serving one local model, bound to 127.0.0.1:8000 and never exposed beyond this host. The vllm/ folder in rayjanwilson/homelab-llm-router is self-contained and portable: copy it to any other homelab GPU box, no changes beyond .env. Ollama is stopped, disabled, and masked via systemd on this machine.

I put a LiteLLM proxy in front of this vLLM instance so every client (agents, IDEs, scripts) has one OpenAI-compatible endpoint instead of talking to vLLM directly, with routing to cloud APIs and other homelab GPU boxes on top of that. That’s its own story, with its own set of problems (compressing only the requests that actually cost money, per-bot API keys, a secrets pipeline that had two real bugs in it) — see Building a homelab LLM router with LiteLLM and Headroom.