Add Docker support and update dependencies

- Introduced Docker support with CPU-only and GPU-enabled configurations via Dockerfiles and docker-compose files.
- Added a .dockerignore file to exclude unnecessary files from Docker images.
- Updated bun.lock and package.json to include new dependencies for icon handling.
- Enhanced README with Docker usage instructions and deployment options.
- Refactored components to utilize new icon libraries for improved UI consistency.
This commit is contained in:
Jamie Pine
2026-02-02 02:19:05 -08:00
parent 4c4b3e5463
commit f090759d8f
61 changed files with 5982 additions and 394 deletions
+1 -1
View File
@@ -40,7 +40,7 @@
{
"group": "Getting Started",
"icon": "rocket",
"pages": ["overview/introduction", "overview/installation", "overview/quick-start"]
"pages": ["overview/introduction", "overview/installation", "overview/docker", "overview/quick-start"]
},
{
"group": "Features",
+403
View File
@@ -0,0 +1,403 @@
---
title: "Docker Deployment"
description: "Run Voicebox in Docker with the web UI for server deployments"
---
## Overview
Voicebox is available as Docker images that include both the backend API and web UI. Run the full Voicebox experience in a container with a single command.
**What's included:**
- FastAPI backend with all TTS/Whisper capabilities
- Complete web UI (same React app as the desktop version)
- Provider download system (downloads PyTorch on first use)
- Multi-architecture support (amd64, arm64)
## Quick Start
<Tabs>
<Tab title="NVIDIA GPU">
```bash
docker run --gpus all -p 8000:8000 \
-v voicebox-data:/app/data \
ghcr.io/jamiepine/voicebox:latest-cuda
```
Open http://localhost:8000 in your browser.
</Tab>
<Tab title="CPU Only">
```bash
docker run -p 8000:8000 \
-v voicebox-data:/app/data \
ghcr.io/jamiepine/voicebox:latest
```
Open http://localhost:8000 in your browser.
</Tab>
<Tab title="Docker Compose">
Clone the repo or download `docker-compose.yml`:
```bash
# CUDA variant (default)
docker compose up -d
# CPU-only variant
docker compose -f docker-compose.cpu.yml up -d
```
Open http://localhost:8000 in your browser.
</Tab>
</Tabs>
<Note>
On first launch, you'll be prompted to download a TTS provider (PyTorch CPU ~300MB or PyTorch CUDA ~2.4GB). This happens once and is cached in the `huggingface-cache` volume.
</Note>
## Available Images
Images are automatically built and published to GitHub Container Registry on each release.
| Image | Description | Platforms |
|-------|-------------|-----------|
| `ghcr.io/jamiepine/voicebox:latest` | Latest CPU-only release | linux/amd64, linux/arm64 |
| `ghcr.io/jamiepine/voicebox:0.1.13` | Specific version (CPU) | linux/amd64, linux/arm64 |
| `ghcr.io/jamiepine/voicebox:latest-cuda` | Latest with NVIDIA GPU support | linux/amd64 |
| `ghcr.io/jamiepine/voicebox:0.1.13-cuda` | Specific version (CUDA) | linux/amd64 |
<Tip>
Pin to a specific version in production to avoid unexpected updates:
```yaml
image: ghcr.io/jamiepine/voicebox:0.1.13-cuda
```
</Tip>
## Docker Compose Examples
### GPU Deployment (Recommended)
```yaml
version: '3.8'
services:
voicebox:
image: ghcr.io/jamiepine/voicebox:latest-cuda
container_name: voicebox
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- voicebox-data:/app/data
- huggingface-cache:/root/.cache/huggingface
environment:
- GPU_MEMORY_FRACTION=0.8
- LOG_LEVEL=info
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
volumes:
voicebox-data:
huggingface-cache:
```
### CPU Deployment
```yaml
version: '3.8'
services:
voicebox:
image: ghcr.io/jamiepine/voicebox:latest
container_name: voicebox
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- voicebox-data:/app/data
- huggingface-cache:/root/.cache/huggingface
environment:
- LOG_LEVEL=info
volumes:
voicebox-data:
huggingface-cache:
```
## Volume Mounts
<CardGroup cols={2}>
<Card title="voicebox-data" icon="database">
Stores voice profiles, generated audio, and database
</Card>
<Card title="huggingface-cache" icon="download">
Caches downloaded TTS/Whisper models (saves re-downloading)
</Card>
</CardGroup>
<Warning>
Always mount `/app/data` to preserve your voice profiles and generations across container restarts.
</Warning>
## Environment Variables
Configure Voicebox behavior with environment variables:
| Variable | Default | Description |
|----------|---------|-------------|
| `GPU_MEMORY_FRACTION` | `0.9` | Fraction of GPU memory to use (0.0-1.0) |
| `LOG_LEVEL` | `info` | Logging level: `debug`, `info`, `warning`, `error` |
| `DATA_DIR` | `/app/data` | Directory for profiles and generations |
Example:
```bash
docker run -e GPU_MEMORY_FRACTION=0.8 \
-e LOG_LEVEL=debug \
-p 8000:8000 \
ghcr.io/jamiepine/voicebox:latest-cuda
```
## Cloud Deployment
### AWS EC2
<Steps>
<Step title="Launch GPU Instance">
Use g4dn.xlarge or p3.2xlarge with NVIDIA GPU
</Step>
<Step title="Install Docker & NVIDIA Container Toolkit">
```bash
# Install Docker
curl -fsSL https://get.docker.com -o get-docker.sh
sudo sh get-docker.sh
# Install NVIDIA Container Toolkit
distribution=$(. /etc/os-release;echo $ID$VERSION_ID)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo systemctl restart docker
```
</Step>
<Step title="Deploy">
```bash
docker run -d --gpus all -p 8000:8000 \
-v voicebox-data:/app/data \
--restart unless-stopped \
ghcr.io/jamiepine/voicebox:latest-cuda
```
</Step>
</Steps>
### DigitalOcean
<Steps>
<Step title="Create GPU Droplet">
```bash
doctl compute droplet create voicebox \
--size gpu-h100x1-80gb \
--image ubuntu-22-04-x64 \
--region nyc3
```
</Step>
<Step title="SSH and Deploy">
```bash
ssh root@<droplet-ip>
curl -fsSL https://get.docker.com | sh
docker run -d --gpus all -p 80:8000 \
ghcr.io/jamiepine/voicebox:latest-cuda
```
</Step>
</Steps>
### Fly.io
Create `fly.toml`:
```toml
app = "voicebox"
[build]
image = "ghcr.io/jamiepine/voicebox:latest"
[[services]]
http_checks = []
internal_port = 8000
protocol = "tcp"
[[services.ports]]
port = 80
handlers = ["http"]
[[services.ports]]
port = 443
handlers = ["tls", "http"]
[mounts]
source = "voicebox_data"
destination = "/app/data"
```
Deploy:
```bash
fly launch
fly deploy
```
## Updates
Docker images are automatically built and published on each GitHub release.
<Tabs>
<Tab title="Latest Tag">
Always get the newest version:
```bash
docker pull ghcr.io/jamiepine/voicebox:latest
docker compose up -d
```
</Tab>
<Tab title="Pinned Version">
Update to a specific version:
```yaml
services:
voicebox:
image: ghcr.io/jamiepine/voicebox:0.1.13-cuda
```
```bash
docker compose pull
docker compose up -d
```
</Tab>
<Tab title="Automatic Updates">
Use Watchtower for automatic updates:
```yaml
services:
voicebox:
image: ghcr.io/jamiepine/voicebox:latest-cuda
# ... other config ...
watchtower:
image: containrrr/watchtower
volumes:
- /var/run/docker.sock:/var/run/docker.sock
command: --interval 3600 # Check hourly
```
</Tab>
</Tabs>
## GPU Requirements
### NVIDIA GPU
Requires:
- **Docker version:** 19.03+
- **NVIDIA Driver:** 450.80.02+
- **NVIDIA Container Toolkit:** Installed and configured
Verify GPU access:
```bash
docker run --rm --gpus all nvidia/cuda:12.1.1-base-ubuntu22.04 nvidia-smi
```
If this works, Voicebox will detect and use your GPU automatically.
### AMD GPU (ROCm)
AMD GPU support via ROCm is not currently available in pre-built images. If you need ROCm support, build a custom image using the ROCm base.
## Troubleshooting
### GPU Not Detected
<Accordion title="Check NVIDIA Docker">
```bash
# Verify NVIDIA Container Toolkit is installed
docker run --rm --gpus all nvidia/cuda:12.1.1-base-ubuntu22.04 nvidia-smi
```
If this fails, reinstall NVIDIA Container Toolkit.
</Accordion>
<Accordion title="Insufficient GPU Memory">
Reduce GPU memory usage:
```bash
docker run -e GPU_MEMORY_FRACTION=0.5 \
--gpus all -p 8000:8000 \
ghcr.io/jamiepine/voicebox:latest-cuda
```
Or use CPU-only mode:
```bash
docker run -p 8000:8000 \
ghcr.io/jamiepine/voicebox:latest
```
</Accordion>
<Accordion title="Port Already in Use">
Change the host port:
```bash
docker run -p 8080:8000 ghcr.io/jamiepine/voicebox:latest
```
Then open http://localhost:8080
</Accordion>
<Accordion title="Permission Errors">
Run with specific user:
```bash
docker run --user $(id -u):$(id -g) \
-v $(pwd)/data:/app/data \
ghcr.io/jamiepine/voicebox:latest
```
</Accordion>
## Building From Source
If you need to customize the Docker image:
```bash
# Clone the repo
git clone https://github.com/jamiepine/voicebox.git
cd voicebox
# Build web UI
bun install
cd web && bun run build && cd ..
# Build Docker image
docker build -t voicebox:custom .
# Or CUDA variant
docker build -f Dockerfile.cuda -t voicebox:custom-cuda .
```
## Next Steps
<CardGroup cols={2}>
<Card title="API Reference" icon="code" href="/api/overview">
Integrate Voicebox into your applications
</Card>
<Card title="Remote Mode" icon="server" href="/overview/remote-mode">
Connect desktop app to Docker backend
</Card>
</CardGroup>
+5 -2
View File
@@ -7,13 +7,16 @@ description: "Download and install Voicebox on macOS, Windows, or Linux"
Voicebox is available for macOS and Windows, with Linux builds coming soon.
<CardGroup cols={2}>
<CardGroup cols={3}>
<Card title="macOS" icon="apple">
Download for Apple Silicon or Intel Macs
</Card>
<Card title="Windows" icon="windows">
Download MSI installer or Setup executable
</Card>
<Card title="Docker" icon="docker" href="/overview/docker">
Run with web UI in a container
</Card>
</CardGroup>
### macOS
@@ -61,7 +64,7 @@ Voicebox is available for macOS and Windows, with Linux builds coming soon.
### Linux
<Note>
Linux builds are coming soon. Currently blocked by GitHub runner disk space limitations.
Linux desktop builds are coming soon. For server deployments, use [Docker](/overview/docker).
</Note>
## First Launch
+72 -193
View File
@@ -1,24 +1,31 @@
# Docker Deployment Guide
**Status:** In Development for v0.2.0
**Requested By:** Reddit community ([thread](https://reddit.com/r/LocalLLaMA/...))
**Status:** Implemented
**Images:** `ghcr.io/jamiepine/voicebox`
## Overview
Docker support makes Voicebox easier to deploy, especially for:
Voicebox is available as Docker images with the full web UI included. Images are automatically built and published to GitHub Container Registry on each release.
- **Consistent Environments**: Same setup across dev/staging/prod
- **GPU Passthrough**: Easy NVIDIA/AMD GPU access
**What's included:**
- FastAPI backend with all TTS/Whisper capabilities
- Complete web UI (same React app as the Tauri desktop version)
- Provider download system (downloads TTS providers on first use, just like desktop)
- Multi-architecture support (amd64, arm64 for CPU variant)
Docker support is ideal for:
- **Server Deployments**: Run on headless Linux servers
- **Multi-User Setups**: Isolate instances per user/team
- **GPU Passthrough**: Easy NVIDIA GPU access
- **Consistent Environments**: Same setup across dev/staging/prod
- **Cloud Platforms**: Deploy to AWS, GCP, Azure, DigitalOcean
- **Multi-User Setups**: Isolate instances per user/team
## Quick Start
### Using Pre-Built Images (Recommended)
```bash
# CPU-only version
# CPU-only version (supports amd64 and arm64)
docker run -p 8000:8000 -v voicebox-data:/app/data \
ghcr.io/jamiepine/voicebox:latest
@@ -26,184 +33,80 @@ docker run -p 8000:8000 -v voicebox-data:/app/data \
docker run --gpus all -p 8000:8000 -v voicebox-data:/app/data \
ghcr.io/jamiepine/voicebox:latest-cuda
# AMD GPU version (experimental)
docker run --device=/dev/kfd --device=/dev/dri -p 8000:8000 \
-v voicebox-data:/app/data \
ghcr.io/jamiepine/voicebox:latest-rocm
# Specific version (pinned for stability)
docker run -p 8000:8000 -v voicebox-data:/app/data \
ghcr.io/jamiepine/voicebox:0.1.13
```
Then open: `http://localhost:8000`
The web UI will load automatically. On first use, you'll be prompted to download a TTS provider (PyTorch CPU ~300MB or PyTorch CUDA ~2.4GB).
### Using Docker Compose (Easiest)
Create `docker-compose.yml`:
Use the provided `docker-compose.yml` (CUDA) or `docker-compose.cpu.yml` in the repository root:
```yaml
version: '3.8'
```bash
# CUDA (default)
docker compose up -d
services:
voicebox:
image: ghcr.io/jamiepine/voicebox:latest-cuda
ports:
- "8000:8000"
volumes:
- voicebox-data:/app/data
- huggingface-cache:/root/.cache/huggingface
environment:
- GPU_MEMORY_FRACTION=0.8 # Use 80% of GPU memory
- TTS_MODE=local
- WHISPER_MODE=local
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
volumes:
voicebox-data:
huggingface-cache:
# Or CPU-only
docker compose -f docker-compose.cpu.yml up -d
```
Run:
```bash
docker compose up -d
To pin to a specific version, edit the compose file:
```yaml
services:
voicebox:
image: ghcr.io/jamiepine/voicebox:0.1.13-cuda # Pinned version
```
## Building From Source
### Basic Dockerfile
```dockerfile
# Dockerfile
FROM python:3.11-slim
WORKDIR /app
# Install system dependencies
RUN apt-get update && apt-get install -y \
git \
build-essential \
ffmpeg \
&& rm -rf /var/lib/apt/lists/*
# Copy application
COPY backend/ /app/backend/
COPY requirements.txt /app/
# Install Python dependencies
RUN pip install --no-cache-dir -r requirements.txt
RUN pip install --no-cache-dir git+https://github.com/QwenLM/Qwen3-TTS.git
# Create data directory
RUN mkdir -p /app/data
# Expose port
EXPOSE 8000
# Run server
CMD ["uvicorn", "backend.main:app", "--host", "0.0.0.0", "--port", "8000"]
```
See `Dockerfile` and `Dockerfile.cuda` in the repository root.
Build and run:
```bash
# Build web UI first
bun install
cd web && bun run build && cd ..
# Build CPU image
docker build -t voicebox .
docker run -p 8000:8000 -v $(pwd)/data:/app/data voicebox
docker run -p 8000:8000 -v voicebox-data:/app/data voicebox
# Or build CUDA image
docker build -f Dockerfile.cuda -t voicebox:cuda .
docker run --gpus all -p 8000:8000 -v voicebox-data:/app/data voicebox:cuda
```
### Multi-Stage Build (Optimized)
### Architecture
Smaller image size by separating build and runtime:
The Docker images include:
- **Backend**: FastAPI server with TTS/Whisper endpoints
- **Web UI**: Pre-built React app served as static files from the backend
- **Provider System**: Downloads PyTorch CPU/CUDA providers on first use (same UX as desktop app)
```dockerfile
# Dockerfile.optimized
# Stage 1: Build dependencies
FROM python:3.11-slim AS builder
WORKDIR /build
RUN apt-get update && apt-get install -y \
git build-essential && \
rm -rf /var/lib/apt/lists/*
COPY backend/requirements.txt .
RUN pip install --no-cache-dir --target=/build/packages \
-r requirements.txt
RUN pip install --no-cache-dir --target=/build/packages \
git+https://github.com/QwenLM/Qwen3-TTS.git
# Stage 2: Runtime
FROM python:3.11-slim
WORKDIR /app
# Install only runtime dependencies
RUN apt-get update && apt-get install -y \
ffmpeg \
&& rm -rf /var/lib/apt/lists/*
# Copy installed packages from builder
COPY --from=builder /build/packages /usr/local/lib/python3.11/site-packages/
# Copy application code
COPY backend/ /app/backend/
# Create data directory
RUN mkdir -p /app/data
EXPOSE 8000
CMD ["uvicorn", "backend.main:app", "--host", "0.0.0.0", "--port", "8000"]
```
Build:
```bash
docker build -f Dockerfile.optimized -t voicebox:slim .
```
Images are automatically built on release and tagged with both version number and `latest`.
## GPU Support
### NVIDIA GPUs (CUDA)
**Dockerfile:**
```dockerfile
FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04
# Install Python
RUN apt-get update && apt-get install -y \
python3.11 python3-pip git ffmpeg && \
rm -rf /var/lib/apt/lists/*
WORKDIR /app
# Install PyTorch with CUDA support
COPY backend/requirements.txt .
RUN pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
# Install other dependencies
RUN pip3 install -r requirements.txt
RUN pip3 install git+https://github.com/QwenLM/Qwen3-TTS.git
COPY backend/ /app/backend/
EXPOSE 8000
CMD ["uvicorn", "backend.main:app", "--host", "0.0.0.0", "--port", "8000"]
```
The CUDA image includes PyTorch with CUDA 12.1 support:
**Run with GPU:**
```bash
docker run --gpus all -p 8000:8000 \
-v voicebox-data:/app/data \
voicebox:cuda
ghcr.io/jamiepine/voicebox:latest-cuda
```
**Docker Compose with GPU:**
```yaml
services:
voicebox:
image: voicebox:cuda
image: ghcr.io/jamiepine/voicebox:latest-cuda
deploy:
resources:
reservations:
@@ -213,47 +116,9 @@ services:
capabilities: [gpu]
```
### AMD GPUs (ROCm) - Experimental
### AMD GPUs (ROCm)
**Dockerfile:**
```dockerfile
FROM rocm/dev-ubuntu-22.04:6.0
# Install Python
RUN apt-get update && apt-get install -y \
python3.11 python3-pip git ffmpeg && \
rm -rf /var/lib/apt/lists/*
WORKDIR /app
# Install PyTorch with ROCm support
COPY backend/requirements.txt .
RUN pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/rocm6.0
# Install other dependencies
RUN pip3 install -r requirements.txt
RUN pip3 install git+https://github.com/QwenLM/Qwen3-TTS.git
# Set ROCm environment variables
ENV HSA_OVERRIDE_GFX_VERSION=10.3.0
ENV ROCM_PATH=/opt/rocm
COPY backend/ /app/backend/
EXPOSE 8000
CMD ["uvicorn", "backend.main:app", "--host", "0.0.0.0", "--port", "8000"]
```
**Run with AMD GPU:**
```bash
docker run --device=/dev/kfd --device=/dev/dri \
--group-add video --ipc=host --cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-p 8000:8000 -v voicebox-data:/app/data \
voicebox:rocm
```
**Note:** ROCm support varies by GPU model. Works best on Linux. See [AMD ROCm docs](https://rocm.docs.amd.com) for compatibility.
ROCm support is not currently available in pre-built images. If you need ROCm, build a custom image using the ROCm base and PyTorch ROCm builds.
## Volume Mounts
@@ -734,13 +599,27 @@ docker logs -f voicebox
docker compose logs -f voicebox
```
## Next Steps
## Updates
- [ ] Publish official images to GitHub Container Registry
- [ ] Add Kubernetes Helm charts
- [ ] Create Docker Desktop extension
- [ ] Add automated vulnerability scanning
- [ ] Support ARM64 builds for Raspberry Pi / Apple Silicon
Docker images are automatically built and published on each GitHub release. To update:
```bash
# Pull latest
docker pull ghcr.io/jamiepine/voicebox:latest
docker compose up -d
# Or pin to a specific version
docker pull ghcr.io/jamiepine/voicebox:0.1.13
```
For automatic updates, use [Watchtower](https://containrrr.dev/watchtower/).
## Future Enhancements
- Kubernetes Helm charts
- Docker Desktop extension
- Automated vulnerability scanning
- ROCm image variant
## Contributing