Update TTS Provider Architecture status to v0.1.13

This commit is contained in:
Jamie Pine
2026-01-31 02:11:41 -08:00
parent 2bc243f93e
commit cb541521d2
+65 -33
View File
@@ -1,6 +1,6 @@
# TTS Provider Architecture # TTS Provider Architecture
**Status:** Planned for v0.2.0 **Status:** Planned for v0.1.13
**Created:** 2025-01-31 **Created:** 2025-01-31
**Problem:** GitHub 2GB release limit + poor UX for frequent updates requiring 2.4GB re-downloads **Problem:** GitHub 2GB release limit + poor UX for frequent updates requiring 2.4GB re-downloads
@@ -14,6 +14,7 @@ Split the monolithic backend into modular components:
2. **TTS Providers** (downloadable plugins): Separate executables for model inference 2. **TTS Providers** (downloadable plugins): Separate executables for model inference
This architecture solves: This architecture solves:
- ✅ GitHub 2GB release artifact limit - ✅ GitHub 2GB release artifact limit
- ✅ Frequent app updates without re-downloading large ML models - ✅ Frequent app updates without re-downloading large ML models
- ✅ User choice of compute backend (CPU/GPU/Cloud) - ✅ User choice of compute backend (CPU/GPU/Cloud)
@@ -68,6 +69,7 @@ This architecture solves:
### Current Architecture Issues ### Current Architecture Issues
**Monolithic Binary:** **Monolithic Binary:**
- CPU version: ~295MB - CPU version: ~295MB
- CUDA version: ~2.37GB - CUDA version: ~2.37GB
- GitHub releases: 2GB file size limit (BLOCKED) - GitHub releases: 2GB file size limit (BLOCKED)
@@ -75,6 +77,7 @@ This architecture solves:
- Poor UX: update app → restart → download CUDA update → restart again - Poor UX: update app → restart → download CUDA update → restart again
**User Pain Points:** **User Pain Points:**
1. Cannot release CUDA version on GitHub (over 2GB) 1. Cannot release CUDA version on GitHub (over 2GB)
2. Every app update forces 2.4GB re-download for GPU users 2. Every app update forces 2.4GB re-download for GPU users
3. No flexibility (can't use OpenAI, remote servers, etc.) 3. No flexibility (can't use OpenAI, remote servers, etc.)
@@ -91,6 +94,7 @@ This architecture solves:
**Size:** ~150-200MB **Size:** ~150-200MB
**Includes:** **Includes:**
- Tauri runtime + React UI - Tauri runtime + React UI
- FastAPI backend (pure Python, no PyTorch) - FastAPI backend (pure Python, no PyTorch)
- Whisper model (tiny, ~50MB) - Whisper model (tiny, ~50MB)
@@ -99,6 +103,7 @@ This architecture solves:
- Provider management system - Provider management system
**Does NOT include:** **Does NOT include:**
- PyTorch (CPU or CUDA) - PyTorch (CPU or CUDA)
- TTS models (Qwen3-TTS) - TTS models (Qwen3-TTS)
- Heavy ML dependencies - Heavy ML dependencies
@@ -113,6 +118,7 @@ This architecture solves:
**Size:** ~300MB **Size:** ~300MB
**Includes:** **Includes:**
- PyTorch CPU build - PyTorch CPU build
- Qwen3-TTS package - Qwen3-TTS package
- Transformers - Transformers
@@ -129,6 +135,7 @@ This architecture solves:
**Size:** ~2.4GB **Size:** ~2.4GB
**Includes:** **Includes:**
- PyTorch CUDA build (cu121) - PyTorch CUDA build (cu121)
- Qwen3-TTS package - Qwen3-TTS package
- CUDA runtime, cuDNN, cuBLAS - CUDA runtime, cuDNN, cuBLAS
@@ -146,6 +153,7 @@ This architecture solves:
**Size:** ~800MB **Size:** ~800MB
**Includes:** **Includes:**
- MLX framework - MLX framework
- MLX-optimized Qwen3-TTS - MLX-optimized Qwen3-TTS
- Metal acceleration - Metal acceleration
@@ -161,11 +169,13 @@ This architecture solves:
**Size:** 0MB **Size:** 0MB
**How it works:** **How it works:**
- User provides URL to their own TTS server - User provides URL to their own TTS server
- Backend proxies requests to that server - Backend proxies requests to that server
- Implements API spec from `EXTERNAL_PROVIDERS.md` - Implements API spec from `EXTERNAL_PROVIDERS.md`
**Use cases:** **Use cases:**
- AMD GPU users running their own server - AMD GPU users running their own server
- Team deployments with shared GPU server - Team deployments with shared GPU server
- Cloud hosting (Modal, RunPod, Replicate) - Cloud hosting (Modal, RunPod, Replicate)
@@ -178,11 +188,13 @@ This architecture solves:
**Size:** 0MB **Size:** 0MB
**How it works:** **How it works:**
- User provides OpenAI API key - User provides OpenAI API key
- Backend wraps OpenAI Audio API - Backend wraps OpenAI Audio API
- Voice profiles map to OpenAI voices - Voice profiles map to OpenAI voices
**Benefits:** **Benefits:**
- Zero local compute - Zero local compute
- Pay-per-use - Pay-per-use
- Instant setup - Instant setup
@@ -200,10 +212,13 @@ All TTS providers must implement these endpoints:
Generate speech from text. Generate speech from text.
**Request:** **Request:**
```json ```json
{ {
"text": "Hello world!", "text": "Hello world!",
"voice_prompt": { /* voice prompt object */ }, "voice_prompt": {
/* voice prompt object */
},
"language": "en", "language": "en",
"seed": 12345, "seed": 12345,
"model_size": "1.7B" "model_size": "1.7B"
@@ -211,6 +226,7 @@ Generate speech from text.
``` ```
**Response:** **Response:**
```json ```json
{ {
"audio": "base64-encoded-audio", "audio": "base64-encoded-audio",
@@ -224,13 +240,17 @@ Generate speech from text.
Create voice prompt from reference audio. Create voice prompt from reference audio.
**Request:** (multipart/form-data) **Request:** (multipart/form-data)
- `audio`: Audio file - `audio`: Audio file
- `reference_text`: Transcript - `reference_text`: Transcript
**Response:** **Response:**
```json ```json
{ {
"voice_prompt": { /* serialized prompt */ } "voice_prompt": {
/* serialized prompt */
}
} }
``` ```
@@ -239,6 +259,7 @@ Create voice prompt from reference audio.
Health check. Health check.
**Response:** **Response:**
```json ```json
{ {
"status": "healthy", "status": "healthy",
@@ -254,6 +275,7 @@ Health check.
Model status. Model status.
**Response:** **Response:**
```json ```json
{ {
"model_loaded": true, "model_loaded": true,
@@ -434,6 +456,7 @@ class ProviderInstaller:
``` ```
**Provider Storage Location:** **Provider Storage Location:**
- Windows: `%APPDATA%/voicebox/providers/` - Windows: `%APPDATA%/voicebox/providers/`
- macOS: `~/Library/Application Support/voicebox/providers/` - macOS: `~/Library/Application Support/voicebox/providers/`
- Linux: `~/.local/share/voicebox/providers/` - Linux: `~/.local/share/voicebox/providers/`
@@ -448,23 +471,24 @@ class ProviderInstaller:
```tsx ```tsx
export function ProviderSettings() { export function ProviderSettings() {
const [selectedProvider, setSelectedProvider] = useState<ProviderType>('auto'); const [selectedProvider, setSelectedProvider] =
const { data: installedProviders } = useQuery({ useState<ProviderType>("auto");
queryKey: ['providers', 'installed'], const {data: installedProviders} = useQuery({
queryFn: () => apiClient.getInstalledProviders() queryKey: ["providers", "installed"],
queryFn: () => apiClient.getInstalledProviders(),
}); });
return ( return (
<Card> <Card>
<CardHeader> <CardHeader>
<CardTitle>TTS Provider</CardTitle> <CardTitle>TTS Provider</CardTitle>
<CardDescription> <CardDescription>Choose how Voicebox generates speech</CardDescription>
Choose how Voicebox generates speech
</CardDescription>
</CardHeader> </CardHeader>
<CardContent> <CardContent>
<RadioGroup value={selectedProvider} onValueChange={setSelectedProvider}> <RadioGroup
value={selectedProvider}
onValueChange={setSelectedProvider}
>
{/* Auto-detect */} {/* Auto-detect */}
<div className="flex items-center space-x-2"> <div className="flex items-center space-x-2">
<RadioGroupItem value="auto" id="auto" /> <RadioGroupItem value="auto" id="auto" />
@@ -479,7 +503,11 @@ export function ProviderSettings() {
{/* PyTorch CUDA */} {/* PyTorch CUDA */}
<div className="flex items-center justify-between"> <div className="flex items-center justify-between">
<div className="flex items-center space-x-2"> <div className="flex items-center space-x-2">
<RadioGroupItem value="pytorch-cuda" id="cuda" disabled={!gpuAvailable} /> <RadioGroupItem
value="pytorch-cuda"
id="cuda"
disabled={!gpuAvailable}
/>
<Label htmlFor="cuda"> <Label htmlFor="cuda">
<div className="font-medium">PyTorch CUDA (NVIDIA GPU)</div> <div className="font-medium">PyTorch CUDA (NVIDIA GPU)</div>
<div className="text-sm text-muted-foreground"> <div className="text-sm text-muted-foreground">
@@ -487,8 +515,11 @@ export function ProviderSettings() {
</div> </div>
</Label> </Label>
</div> </div>
{!installedProviders?.includes('pytorch-cuda') && gpuAvailable && ( {!installedProviders?.includes("pytorch-cuda") && gpuAvailable && (
<Button onClick={() => downloadProvider('pytorch-cuda')} size="sm"> <Button
onClick={() => downloadProvider("pytorch-cuda")}
size="sm"
>
Download (2.4GB) Download (2.4GB)
</Button> </Button>
)} )}
@@ -505,8 +536,8 @@ export function ProviderSettings() {
</div> </div>
</Label> </Label>
</div> </div>
{!installedProviders?.includes('pytorch-cpu') && ( {!installedProviders?.includes("pytorch-cpu") && (
<Button onClick={() => downloadProvider('pytorch-cpu')} size="sm"> <Button onClick={() => downloadProvider("pytorch-cpu")} size="sm">
Download (300MB) Download (300MB)
</Button> </Button>
)} )}
@@ -524,8 +555,8 @@ export function ProviderSettings() {
</div> </div>
</Label> </Label>
</div> </div>
{!installedProviders?.includes('mlx') && ( {!installedProviders?.includes("mlx") && (
<Button onClick={() => downloadProvider('mlx')} size="sm"> <Button onClick={() => downloadProvider("mlx")} size="sm">
Download (800MB) Download (800MB)
</Button> </Button>
)} )}
@@ -543,11 +574,8 @@ export function ProviderSettings() {
</div> </div>
</Label> </Label>
</div> </div>
{selectedProvider === 'remote' && ( {selectedProvider === "remote" && (
<Input <Input placeholder="http://your-server:8000" className="ml-6" />
placeholder="http://your-server:8000"
className="ml-6"
/>
)} )}
</div> </div>
@@ -562,15 +590,10 @@ export function ProviderSettings() {
</div> </div>
</Label> </Label>
</div> </div>
{selectedProvider === 'openai' && ( {selectedProvider === "openai" && (
<Input <Input type="password" placeholder="sk-..." className="ml-6" />
type="password"
placeholder="sk-..."
className="ml-6"
/>
)} )}
</div> </div>
</RadioGroup> </RadioGroup>
</CardContent> </CardContent>
</Card> </Card>
@@ -718,7 +741,7 @@ Providers have their own version numbers, independent of the main app:
**Example:** **Example:**
| App Version | Min Provider Version | Max Provider Version | | App Version | Min Provider Version | Max Provider Version |
|-------------|---------------------|---------------------| | ----------- | -------------------- | -------------------- |
| v0.2.0 | v1.0.0 | v1.x.x | | v0.2.0 | v1.0.0 | v1.x.x |
| v0.3.0 | v1.0.0 | v1.x.x | | v0.3.0 | v1.0.0 | v1.x.x |
| v0.4.0 | v1.2.0 | v1.x.x | | v0.4.0 | v1.2.0 | v1.x.x |
@@ -735,6 +758,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
``` ```
**UI shows warning if incompatible:** **UI shows warning if incompatible:**
``` ```
⚠️ Provider version 0.9.0 is outdated. Update to v1.0.0+ ⚠️ Provider version 0.9.0 is outdated. Update to v1.0.0+
``` ```
@@ -748,6 +772,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
1. User downloads and installs Voicebox (~150MB) 1. User downloads and installs Voicebox (~150MB)
2. App launches → detects no TTS provider installed 2. App launches → detects no TTS provider installed
3. Shows setup wizard: 3. Shows setup wizard:
``` ```
Choose your TTS provider: Choose your TTS provider:
@@ -769,6 +794,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
[ ] OpenAI API [ ] OpenAI API
API Key: ________________ API Key: ________________
``` ```
4. User selects provider → downloads with progress bar 4. User selects provider → downloads with progress bar
5. Provider installs to AppData/Application Support 5. Provider installs to AppData/Application Support
6. App starts provider → ready to use 6. App starts provider → ready to use
@@ -819,7 +845,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
## Benefits ## Benefits
| Benefit | Details | | Benefit | Details |
|---------|---------| | ----------------------------- | --------------------------------------------------------- |
| **GitHub Releases Work** | Main app ~150MB << 2GB limit | | **GitHub Releases Work** | Main app ~150MB << 2GB limit |
| **Fast Updates** | UI/feature updates don't require re-downloading providers | | **Fast Updates** | UI/feature updates don't require re-downloading providers |
| **User Choice** | CPU, CUDA, MLX, OpenAI, remote server | | **User Choice** | CPU, CUDA, MLX, OpenAI, remote server |
@@ -838,6 +864,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** Should providers have independent versions or match app version? **Question:** Should providers have independent versions or match app version?
**Options:** **Options:**
- A. Independent (providers: v1.x, app: v0.2.x) - A. Independent (providers: v1.x, app: v0.2.x)
- B. Matched (both use v0.2.x) - B. Matched (both use v0.2.x)
@@ -850,6 +877,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** Should providers auto-update separately from app? **Question:** Should providers auto-update separately from app?
**Options:** **Options:**
- A. Manual updates only (user clicks "Update Provider") - A. Manual updates only (user clicks "Update Provider")
- B. Optional auto-update (user can enable) - B. Optional auto-update (user can enable)
- C. Always auto-update - C. Always auto-update
@@ -863,6 +891,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** How does app find installed providers? **Question:** How does app find installed providers?
**Options:** **Options:**
- A. Check standard paths in AppData/Application Support - A. Check standard paths in AppData/Application Support
- B. Registry (Windows) / plist (macOS) - B. Registry (Windows) / plist (macOS)
- C. Config file with provider locations - C. Config file with provider locations
@@ -876,6 +905,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** What if no provider is installed? **Question:** What if no provider is installed?
**Options:** **Options:**
- A. Show setup wizard on first launch - A. Show setup wizard on first launch
- B. Block app until provider installed - B. Block app until provider installed
- C. Allow app to run in "demo mode" (transcription only) - C. Allow app to run in "demo mode" (transcription only)
@@ -889,6 +919,7 @@ async def check_provider_compatibility(provider_version: str) -> bool:
**Question:** Should provider start automatically with app? **Question:** Should provider start automatically with app?
**Options:** **Options:**
- A. Always start selected provider on app launch - A. Always start selected provider on app launch
- B. Start on-demand (when user generates speech) - B. Start on-demand (when user generates speech)
- C. User preference - C. User preference
@@ -928,5 +959,6 @@ If you want to build a custom TTS provider:
4. Share in GitHub Discussions 4. Share in GitHub Discussions
**Questions?** **Questions?**
- GitHub Issues: [voicebox/issues](https://github.com/jamiepine/voicebox/issues) - GitHub Issues: [voicebox/issues](https://github.com/jamiepine/voicebox/issues)
- Discord: Coming soon - Discord: Coming soon