Update TTS Provider Architecture status to v0.1.13

This commit is contained in:
Jamie Pine
2026-01-31 02:11:41 -08:00
parent 2bc243f93e
commit cb541521d2
+188 -156
View File
@@ -1,6 +1,6 @@
# TTS Provider Architecture # TTS Provider Architecture
**Status:** Planned for v0.2.0 **Status:** Planned for v0.1.13
**Created:** 2025-01-31 **Created:** 2025-01-31
**Problem:** GitHub 2GB release limit + poor UX for frequent updates requiring 2.4GB re-downloads **Problem:** GitHub 2GB release limit + poor UX for frequent updates requiring 2.4GB re-downloads
@@ -14,6 +14,7 @@ Split the monolithic backend into modular components:
2. **TTS Providers** (downloadable plugins): Separate executables for model inference 2. **TTS Providers** (downloadable plugins): Separate executables for model inference
This architecture solves: This architecture solves:
- ✅ GitHub 2GB release artifact limit - ✅ GitHub 2GB release artifact limit
- ✅ Frequent app updates without re-downloading large ML models - ✅ Frequent app updates without re-downloading large ML models
- ✅ User choice of compute backend (CPU/GPU/Cloud) - ✅ User choice of compute backend (CPU/GPU/Cloud)
@@ -68,6 +69,7 @@ This architecture solves:
### Current Architecture Issues ### Current Architecture Issues
**Monolithic Binary:** **Monolithic Binary:**
- CPU version: ~295MB - CPU version: ~295MB
- CUDA version: ~2.37GB - CUDA version: ~2.37GB
- GitHub releases: 2GB file size limit (BLOCKED) - GitHub releases: 2GB file size limit (BLOCKED)
@@ -75,6 +77,7 @@ This architecture solves:
- Poor UX: update app → restart → download CUDA update → restart again - Poor UX: update app → restart → download CUDA update → restart again
**User Pain Points:** **User Pain Points:**
1. Cannot release CUDA version on GitHub (over 2GB) 1. Cannot release CUDA version on GitHub (over 2GB)
2. Every app update forces 2.4GB re-download for GPU users 2. Every app update forces 2.4GB re-download for GPU users
3. No flexibility (can't use OpenAI, remote servers, etc.) 3. No flexibility (can't use OpenAI, remote servers, etc.)
@@ -91,6 +94,7 @@ This architecture solves:
**Size:** ~150-200MB **Size:** ~150-200MB
**Includes:** **Includes:**
- Tauri runtime + React UI - Tauri runtime + React UI
- FastAPI backend (pure Python, no PyTorch) - FastAPI backend (pure Python, no PyTorch)
- Whisper model (tiny, ~50MB) - Whisper model (tiny, ~50MB)
@@ -99,6 +103,7 @@ This architecture solves:
- Provider management system - Provider management system
**Does NOT include:** **Does NOT include:**
- PyTorch (CPU or CUDA) - PyTorch (CPU or CUDA)
- TTS models (Qwen3-TTS) - TTS models (Qwen3-TTS)
- Heavy ML dependencies - Heavy ML dependencies
@@ -113,6 +118,7 @@ This architecture solves:
**Size:** ~300MB **Size:** ~300MB
**Includes:** **Includes:**
- PyTorch CPU build - PyTorch CPU build
- Qwen3-TTS package - Qwen3-TTS package
- Transformers - Transformers
@@ -129,6 +135,7 @@ This architecture solves:
**Size:** ~2.4GB **Size:** ~2.4GB
**Includes:** **Includes:**
- PyTorch CUDA build (cu121) - PyTorch CUDA build (cu121)
- Qwen3-TTS package - Qwen3-TTS package
- CUDA runtime, cuDNN, cuBLAS - CUDA runtime, cuDNN, cuBLAS
@@ -146,6 +153,7 @@ This architecture solves:
**Size:** ~800MB **Size:** ~800MB
**Includes:** **Includes:**
- MLX framework - MLX framework
- MLX-optimized Qwen3-TTS - MLX-optimized Qwen3-TTS
- Metal acceleration - Metal acceleration
@@ -161,11 +169,13 @@ This architecture solves:
**Size:** 0MB **Size:** 0MB
**How it works:** **How it works:**
- User provides URL to their own TTS server - User provides URL to their own TTS server
- Backend proxies requests to that server - Backend proxies requests to that server
- Implements API spec from `EXTERNAL_PROVIDERS.md` - Implements API spec from `EXTERNAL_PROVIDERS.md`
**Use cases:** **Use cases:**
- AMD GPU users running their own server - AMD GPU users running their own server
- Team deployments with shared GPU server - Team deployments with shared GPU server
- Cloud hosting (Modal, RunPod, Replicate) - Cloud hosting (Modal, RunPod, Replicate)
@@ -178,11 +188,13 @@ This architecture solves:
**Size:** 0MB **Size:** 0MB
**How it works:** **How it works:**
- User provides OpenAI API key - User provides OpenAI API key
- Backend wraps OpenAI Audio API - Backend wraps OpenAI Audio API
- Voice profiles map to OpenAI voices - Voice profiles map to OpenAI voices
**Benefits:** **Benefits:**
- Zero local compute - Zero local compute
- Pay-per-use - Pay-per-use
- Instant setup - Instant setup
@@ -200,22 +212,26 @@ All TTS providers must implement these endpoints:
Generate speech from text. Generate speech from text.
**Request:** **Request:**
```json ```json
{ {
"text": "Hello world!", "text": "Hello world!",
"voice_prompt": { /* voice prompt object */ }, "voice_prompt": {
"language": "en", /* voice prompt object */
"seed": 12345, },
"model_size": "1.7B" "language": "en",
"seed": 12345,
"model_size": "1.7B"
} }
``` ```
**Response:** **Response:**
```json ```json
{ {
"audio": "base64-encoded-audio", "audio": "base64-encoded-audio",
"sample_rate": 24000, "sample_rate": 24000,
"duration": 2.5 "duration": 2.5
} }
``` ```
@@ -224,13 +240,17 @@ Generate speech from text.
Create voice prompt from reference audio. Create voice prompt from reference audio.
**Request:** (multipart/form-data) **Request:** (multipart/form-data)
- `audio`: Audio file - `audio`: Audio file
- `reference_text`: Transcript - `reference_text`: Transcript
**Response:** **Response:**
```json ```json
{ {
"voice_prompt": { /* serialized prompt */ } "voice_prompt": {
/* serialized prompt */
}
} }
``` ```
@@ -239,13 +259,14 @@ Create voice prompt from reference audio.
Health check. Health check.
**Response:** **Response:**
```json ```json
{ {
"status": "healthy", "status": "healthy",
"provider": "pytorch-cuda", "provider": "pytorch-cuda",
"version": "1.0.0", "version": "1.0.0",
"model": "Qwen3-TTS-12Hz-1.7B-Base", "model": "Qwen3-TTS-12Hz-1.7B-Base",
"device": "cuda:0" "device": "cuda:0"
} }
``` ```
@@ -254,13 +275,14 @@ Health check.
Model status. Model status.
**Response:** **Response:**
```json ```json
{ {
"model_loaded": true, "model_loaded": true,
"model_size": "1.7B", "model_size": "1.7B",
"available_sizes": ["0.6B", "1.7B"], "available_sizes": ["0.6B", "1.7B"],
"gpu_available": true, "gpu_available": true,
"vram_used_mb": 1234 "vram_used_mb": 1234
} }
``` ```
@@ -434,6 +456,7 @@ class ProviderInstaller:
``` ```
**Provider Storage Location:** **Provider Storage Location:**
- Windows: `%APPDATA%/voicebox/providers/` - Windows: `%APPDATA%/voicebox/providers/`
- macOS: `~/Library/Application Support/voicebox/providers/` - macOS: `~/Library/Application Support/voicebox/providers/`
- Linux: `~/.local/share/voicebox/providers/` - Linux: `~/.local/share/voicebox/providers/`
@@ -448,133 +471,133 @@ class ProviderInstaller:
```tsx ```tsx
export function ProviderSettings() { export function ProviderSettings() {
const [selectedProvider, setSelectedProvider] = useState<ProviderType>('auto'); const [selectedProvider, setSelectedProvider] =
const { data: installedProviders } = useQuery({ useState<ProviderType>("auto");
queryKey: ['providers', 'installed'], const {data: installedProviders} = useQuery({
queryFn: () => apiClient.getInstalledProviders() queryKey: ["providers", "installed"],
}); queryFn: () => apiClient.getInstalledProviders(),
});
return ( return (
<Card> <Card>
<CardHeader> <CardHeader>
<CardTitle>TTS Provider</CardTitle> <CardTitle>TTS Provider</CardTitle>
<CardDescription> <CardDescription>Choose how Voicebox generates speech</CardDescription>
Choose how Voicebox generates speech </CardHeader>
</CardDescription> <CardContent>
</CardHeader> <RadioGroup
<CardContent> value={selectedProvider}
<RadioGroup value={selectedProvider} onValueChange={setSelectedProvider}> onValueChange={setSelectedProvider}
>
{/* Auto-detect */}
<div className="flex items-center space-x-2">
<RadioGroupItem value="auto" id="auto" />
<Label htmlFor="auto">
<div className="font-medium">Auto-detect (Recommended)</div>
<div className="text-sm text-muted-foreground">
Automatically choose the best available provider
</div>
</Label>
</div>
{/* Auto-detect */} {/* PyTorch CUDA */}
<div className="flex items-center space-x-2"> <div className="flex items-center justify-between">
<RadioGroupItem value="auto" id="auto" /> <div className="flex items-center space-x-2">
<Label htmlFor="auto"> <RadioGroupItem
<div className="font-medium">Auto-detect (Recommended)</div> value="pytorch-cuda"
<div className="text-sm text-muted-foreground"> id="cuda"
Automatically choose the best available provider disabled={!gpuAvailable}
</div> />
</Label> <Label htmlFor="cuda">
</div> <div className="font-medium">PyTorch CUDA (NVIDIA GPU)</div>
<div className="text-sm text-muted-foreground">
4-5x faster inference on NVIDIA GPUs
</div>
</Label>
</div>
{!installedProviders?.includes("pytorch-cuda") && gpuAvailable && (
<Button
onClick={() => downloadProvider("pytorch-cuda")}
size="sm"
>
Download (2.4GB)
</Button>
)}
</div>
{/* PyTorch CUDA */} {/* PyTorch CPU */}
<div className="flex items-center justify-between"> <div className="flex items-center justify-between">
<div className="flex items-center space-x-2"> <div className="flex items-center space-x-2">
<RadioGroupItem value="pytorch-cuda" id="cuda" disabled={!gpuAvailable} /> <RadioGroupItem value="pytorch-cpu" id="cpu" />
<Label htmlFor="cuda"> <Label htmlFor="cpu">
<div className="font-medium">PyTorch CUDA (NVIDIA GPU)</div> <div className="font-medium">PyTorch CPU</div>
<div className="text-sm text-muted-foreground"> <div className="text-sm text-muted-foreground">
4-5x faster inference on NVIDIA GPUs Works on any system, slower inference
</div> </div>
</Label> </Label>
</div> </div>
{!installedProviders?.includes('pytorch-cuda') && gpuAvailable && ( {!installedProviders?.includes("pytorch-cpu") && (
<Button onClick={() => downloadProvider('pytorch-cuda')} size="sm"> <Button onClick={() => downloadProvider("pytorch-cpu")} size="sm">
Download (2.4GB) Download (300MB)
</Button> </Button>
)} )}
</div> </div>
{/* PyTorch CPU */} {/* MLX (macOS only) */}
<div className="flex items-center justify-between"> {isMacOS && (
<div className="flex items-center space-x-2"> <div className="flex items-center justify-between">
<RadioGroupItem value="pytorch-cpu" id="cpu" /> <div className="flex items-center space-x-2">
<Label htmlFor="cpu"> <RadioGroupItem value="mlx" id="mlx" />
<div className="font-medium">PyTorch CPU</div> <Label htmlFor="mlx">
<div className="text-sm text-muted-foreground"> <div className="font-medium">MLX (Apple Silicon)</div>
Works on any system, slower inference <div className="text-sm text-muted-foreground">
</div> Optimized for M1/M2/M3 chips
</Label> </div>
</div> </Label>
{!installedProviders?.includes('pytorch-cpu') && ( </div>
<Button onClick={() => downloadProvider('pytorch-cpu')} size="sm"> {!installedProviders?.includes("mlx") && (
Download (300MB) <Button onClick={() => downloadProvider("mlx")} size="sm">
</Button> Download (800MB)
)} </Button>
</div> )}
</div>
)}
{/* MLX (macOS only) */} {/* Remote */}
{isMacOS && ( <div className="space-y-2">
<div className="flex items-center justify-between"> <div className="flex items-center space-x-2">
<div className="flex items-center space-x-2"> <RadioGroupItem value="remote" id="remote" />
<RadioGroupItem value="mlx" id="mlx" /> <Label htmlFor="remote">
<Label htmlFor="mlx"> <div className="font-medium">Remote Server</div>
<div className="font-medium">MLX (Apple Silicon)</div> <div className="text-sm text-muted-foreground">
<div className="text-sm text-muted-foreground"> Connect to your own TTS server
Optimized for M1/M2/M3 chips </div>
</div> </Label>
</Label> </div>
</div> {selectedProvider === "remote" && (
{!installedProviders?.includes('mlx') && ( <Input placeholder="http://your-server:8000" className="ml-6" />
<Button onClick={() => downloadProvider('mlx')} size="sm"> )}
Download (800MB) </div>
</Button>
)}
</div>
)}
{/* Remote */} {/* OpenAI */}
<div className="space-y-2"> <div className="space-y-2">
<div className="flex items-center space-x-2"> <div className="flex items-center space-x-2">
<RadioGroupItem value="remote" id="remote" /> <RadioGroupItem value="openai" id="openai" />
<Label htmlFor="remote"> <Label htmlFor="openai">
<div className="font-medium">Remote Server</div> <div className="font-medium">OpenAI API</div>
<div className="text-sm text-muted-foreground"> <div className="text-sm text-muted-foreground">
Connect to your own TTS server Use OpenAI's TTS API (requires API key)
</div> </div>
</Label> </Label>
</div> </div>
{selectedProvider === 'remote' && ( {selectedProvider === "openai" && (
<Input <Input type="password" placeholder="sk-..." className="ml-6" />
placeholder="http://your-server:8000" )}
className="ml-6" </div>
/> </RadioGroup>
)} </CardContent>
</div> </Card>
);
{/* OpenAI */}
<div className="space-y-2">
<div className="flex items-center space-x-2">
<RadioGroupItem value="openai" id="openai" />
<Label htmlFor="openai">
<div className="font-medium">OpenAI API</div>
<div className="text-sm text-muted-foreground">
Use OpenAI's TTS API (requires API key)
</div>
</Label>
</div>
{selectedProvider === 'openai' && (
<Input
type="password"
placeholder="sk-..."
className="ml-6"
/>
)}
</div>
</RadioGroup>
</CardContent>
</Card>
);
} }
``` ```
@@ -718,11 +741,11 @@ Providers have their own version numbers, independent of the main app:
**Example:** **Example:**
| App Version | Min Provider Version | Max Provider Version | | App Version | Min Provider Version | Max Provider Version |
|-------------|---------------------|---------------------| | ----------- | -------------------- | -------------------- |
| v0.2.0 | v1.0.0 | v1.x.x | | v0.2.0 | v1.0.0 | v1.x.x |
| v0.3.0 | v1.0.0 | v1.x.x | | v0.3.0 | v1.0.0 | v1.x.x |
| v0.4.0 | v1.2.0 | v1.x.x | | v0.4.0 | v1.2.0 | v1.x.x |
| v1.0.0 | v2.0.0 | v2.x.x | | v1.0.0 | v2.0.0 | v2.x.x |
**Backend checks compatibility:** **Backend checks compatibility:**
@@ -735,6 +758,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
``` ```
**UI shows warning if incompatible:** **UI shows warning if incompatible:**
``` ```
⚠️ Provider version 0.9.0 is outdated. Update to v1.0.0+ ⚠️ Provider version 0.9.0 is outdated. Update to v1.0.0+
``` ```
@@ -748,6 +772,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
1. User downloads and installs Voicebox (~150MB) 1. User downloads and installs Voicebox (~150MB)
2. App launches → detects no TTS provider installed 2. App launches → detects no TTS provider installed
3. Shows setup wizard: 3. Shows setup wizard:
``` ```
Choose your TTS provider: Choose your TTS provider:
@@ -769,6 +794,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
[ ] OpenAI API [ ] OpenAI API
API Key: ________________ API Key: ________________
``` ```
4. User selects provider → downloads with progress bar 4. User selects provider → downloads with progress bar
5. Provider installs to AppData/Application Support 5. Provider installs to AppData/Application Support
6. App starts provider → ready to use 6. App starts provider → ready to use
@@ -818,16 +844,16 @@ async def check_provider_compatibility(provider_version: str) -> bool:
## Benefits ## Benefits
| Benefit | Details | | Benefit | Details |
|---------|---------| | ----------------------------- | --------------------------------------------------------- |
| **GitHub Releases Work** | Main app ~150MB << 2GB limit | | **GitHub Releases Work** | Main app ~150MB << 2GB limit |
| **Fast Updates** | UI/feature updates don't require re-downloading providers | | **Fast Updates** | UI/feature updates don't require re-downloading providers |
| **User Choice** | CPU, CUDA, MLX, OpenAI, remote server | | **User Choice** | CPU, CUDA, MLX, OpenAI, remote server |
| **External Provider Support** | Users can run their own TTS servers | | **External Provider Support** | Users can run their own TTS servers |
| **Bandwidth Savings** | Only download provider once, app updates are small | | **Bandwidth Savings** | Only download provider once, app updates are small |
| **Future-Proof** | Easy to add new providers (ElevenLabs, custom models) | | **Future-Proof** | Easy to add new providers (ElevenLabs, custom models) |
| **Team Deployments** | Multiple users share one remote provider | | **Team Deployments** | Multiple users share one remote provider |
| **Cloud-Ready** | Works with Modal, Replicate, RunPod, etc. | | **Cloud-Ready** | Works with Modal, Replicate, RunPod, etc. |
--- ---
@@ -838,6 +864,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** Should providers have independent versions or match app version? **Question:** Should providers have independent versions or match app version?
**Options:** **Options:**
- A. Independent (providers: v1.x, app: v0.2.x) - A. Independent (providers: v1.x, app: v0.2.x)
- B. Matched (both use v0.2.x) - B. Matched (both use v0.2.x)
@@ -850,6 +877,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** Should providers auto-update separately from app? **Question:** Should providers auto-update separately from app?
**Options:** **Options:**
- A. Manual updates only (user clicks "Update Provider") - A. Manual updates only (user clicks "Update Provider")
- B. Optional auto-update (user can enable) - B. Optional auto-update (user can enable)
- C. Always auto-update - C. Always auto-update
@@ -863,6 +891,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** How does app find installed providers? **Question:** How does app find installed providers?
**Options:** **Options:**
- A. Check standard paths in AppData/Application Support - A. Check standard paths in AppData/Application Support
- B. Registry (Windows) / plist (macOS) - B. Registry (Windows) / plist (macOS)
- C. Config file with provider locations - C. Config file with provider locations
@@ -876,6 +905,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** What if no provider is installed? **Question:** What if no provider is installed?
**Options:** **Options:**
- A. Show setup wizard on first launch - A. Show setup wizard on first launch
- B. Block app until provider installed - B. Block app until provider installed
- C. Allow app to run in "demo mode" (transcription only) - C. Allow app to run in "demo mode" (transcription only)
@@ -889,6 +919,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** Should provider start automatically with app? **Question:** Should provider start automatically with app?
**Options:** **Options:**
- A. Always start selected provider on app launch - A. Always start selected provider on app launch
- B. Start on-demand (when user generates speech) - B. Start on-demand (when user generates speech)
- C. User preference - C. User preference
@@ -928,5 +959,6 @@ If you want to build a custom TTS provider:
4. Share in GitHub Discussions 4. Share in GitHub Discussions
**Questions?** **Questions?**
- GitHub Issues: [voicebox/issues](https://github.com/jamiepine/voicebox/issues) - GitHub Issues: [voicebox/issues](https://github.com/jamiepine/voicebox/issues)
- Discord: Coming soon - Discord: Coming soon