A plain `docker run -v ./data:/app/data` bind mount owned by the host
user had its uid taken over (chown to 999) and its contents left
unwritable, because the adoption probe only looked at the generations
subdir, which `mkdir -p` had just created root-owned. Probe /app/data
first and fall back to generations; re-own only what the old uid owned;
chown a dir only when root created it; and let a read-only HF cache
mount through (the app only warns about it) instead of aborting on the
bare chown error.
Docker creates /home/voicebox/.cache (parent of the huggingface-cache
volume mountpoint) owned by root on every container create, and a
fresh named volume is root-owned too. The app runs as the voicebox
user (uid 999), so anything that writes a cache outside the HF mount
(torch hub, spacy, ...) fails with Permission denied.
The entrypoint already runs as root before dropping privileges via
gosu — create/chown the cache dirs there (non-recursive, instant).
Co-Authored-By: Claude Fable 5 <[email protected]>
* Fix ROCm setup for Linux AMD GPUs
- Ensure Docker ROCm builds resolve PyTorch packages from the ROCm wheel index so later dependency installs do not replace them with CUDA wheels.
- Move ROCm device group handling to a runtime entrypoint that joins the groups owning /dev/kfd and /dev/dri, avoiding distro-specific render/video GID defaults.
- Leave HSA_OVERRIDE_GFX_VERSION unset by default in the ROCm compose overlay so newer RDNA GPUs can use native ROCm detection.
- Add Linux GPU detection to the Unix setup recipe so AMD systems install ROCm torch wheels and NVIDIA systems install CUDA wheels before backend dependencies.
* docs(changelog): add Linux ROCm setup entry
* fix(setup): pin ROCm torch wheels and prefer NVIDIA over amdgpu
- Install torch/torchaudio from the ROCm index only, before the pooled
requirements install, so a plain PyPI (CUDA) wheel can't outrank +rocm
- Detect NVIDIA before AMD and gate ROCm on /dev/kfd, so hybrid
AMD+NVIDIA hosts get CUDA instead of ROCm