mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-09-15 12:50:42 -07:00
- Created a new .npmrc file to enforce bun usage. - Bumped version numbers for multiple packages to 0.1.9 in bun.lock. - Added react-sound-visualizer dependency to enhance audio visualization features. - Introduced convert:assets script in package.json for asset optimization. - Updated CONTRIBUTING.md with instructions for converting assets to web formats. - Added documentation files for API endpoints and developer guidelines in the docs directory.
284 lines
6.7 KiB
Plaintext
284 lines
6.7 KiB
Plaintext
---
|
|
title: "TTS Generation"
|
|
description: "How text-to-speech generation works in Voicebox"
|
|
---
|
|
|
|
## Overview
|
|
|
|
Voicebox uses Qwen3-TTS for voice cloning and text-to-speech generation. The TTS module handles model loading, voice prompt creation, and audio synthesis.
|
|
|
|
## Architecture
|
|
|
|
The TTS system is built around the `TTSModel` class which manages:
|
|
|
|
**Model Loading:** Lazy loading with automatic HuggingFace Hub download.
|
|
|
|
**Voice Prompts:** Converting reference audio into embeddings.
|
|
|
|
**Generation:** Synthesizing speech from text using voice prompts.
|
|
|
|
## TTSModel Class
|
|
|
|
```python
|
|
class TTSModel:
|
|
def __init__(self, model_size: str = "1.7B"):
|
|
self.model = None
|
|
self.model_size = model_size
|
|
self.device = self._get_device() # cuda, mps, or cpu
|
|
```
|
|
|
|
### Device Selection
|
|
|
|
The model automatically selects the best available device:
|
|
|
|
```python
|
|
def _get_device(self) -> str:
|
|
if torch.cuda.is_available():
|
|
return "cuda"
|
|
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
|
return "cpu" # MPS can have issues, use CPU for stability
|
|
return "cpu"
|
|
```
|
|
|
|
## Model Loading
|
|
|
|
Models are downloaded from HuggingFace Hub on first use:
|
|
|
|
```python
|
|
def load_model(self, model_size: Optional[str] = None):
|
|
# Model IDs on HuggingFace Hub
|
|
hf_model_map = {
|
|
"1.7B": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
|
|
"0.6B": "Qwen/Qwen3-TTS-12Hz-0.6B-Base",
|
|
}
|
|
|
|
# Load with progress tracking
|
|
with tracker.patch_download():
|
|
self.model = Qwen3TTSModel.from_pretrained(
|
|
model_path,
|
|
device_map=self.device,
|
|
torch_dtype=torch.bfloat16, # float32 on CPU
|
|
)
|
|
```
|
|
|
|
### Async Loading
|
|
|
|
Loading runs in a thread pool to avoid blocking the event loop:
|
|
|
|
```python
|
|
async def load_model_async(self, model_size: Optional[str] = None):
|
|
if self.model is not None and self._current_model_size == model_size:
|
|
return
|
|
await asyncio.to_thread(self.load_model, model_size)
|
|
```
|
|
|
|
## Voice Prompt Creation
|
|
|
|
Voice prompts are created from reference audio and cached for reuse:
|
|
|
|
```python
|
|
async def create_voice_prompt(
|
|
self,
|
|
audio_path: str,
|
|
reference_text: str,
|
|
use_cache: bool = True,
|
|
) -> Tuple[dict, bool]:
|
|
await self.load_model_async()
|
|
|
|
# Check cache
|
|
if use_cache:
|
|
cache_key = get_cache_key(audio_path, reference_text)
|
|
cached = get_cached_voice_prompt(cache_key)
|
|
if cached:
|
|
return cached, True
|
|
|
|
# Create prompt (blocking, run in thread pool)
|
|
voice_prompt = await asyncio.to_thread(
|
|
self.model.create_voice_clone_prompt,
|
|
ref_audio=audio_path,
|
|
ref_text=reference_text,
|
|
)
|
|
|
|
# Cache the result
|
|
cache_voice_prompt(cache_key, voice_prompt)
|
|
return voice_prompt, False
|
|
```
|
|
|
|
### Combining Multiple Samples
|
|
|
|
When a profile has multiple samples, they're combined:
|
|
|
|
```python
|
|
async def combine_voice_prompts(
|
|
self,
|
|
audio_paths: List[str],
|
|
reference_texts: List[str],
|
|
) -> Tuple[np.ndarray, str]:
|
|
combined_audio = []
|
|
|
|
for audio_path in audio_paths:
|
|
audio, sr = load_audio(audio_path)
|
|
audio = normalize_audio(audio)
|
|
combined_audio.append(audio)
|
|
|
|
# Concatenate and normalize
|
|
mixed = np.concatenate(combined_audio)
|
|
mixed = normalize_audio(mixed)
|
|
|
|
# Combine texts
|
|
combined_text = " ".join(reference_texts)
|
|
|
|
return mixed, combined_text
|
|
```
|
|
|
|
## Speech Generation
|
|
|
|
The core generation function:
|
|
|
|
```python
|
|
async def generate(
|
|
self,
|
|
text: str,
|
|
voice_prompt: dict,
|
|
language: str = "en",
|
|
seed: Optional[int] = None,
|
|
instruct: Optional[str] = None,
|
|
) -> Tuple[np.ndarray, int]:
|
|
await self.load_model_async()
|
|
|
|
def _generate_sync():
|
|
# Set seed for reproducibility
|
|
if seed is not None:
|
|
torch.manual_seed(seed)
|
|
|
|
# Generate audio
|
|
wavs, sample_rate = self.model.generate_voice_clone(
|
|
text=text,
|
|
voice_clone_prompt=voice_prompt,
|
|
instruct=instruct, # Natural language delivery control
|
|
)
|
|
return wavs[0], sample_rate
|
|
|
|
# Run in thread pool
|
|
return await asyncio.to_thread(_generate_sync)
|
|
```
|
|
|
|
### Instruct Feature
|
|
|
|
The `instruct` parameter allows natural language control over speech delivery:
|
|
|
|
```python
|
|
# Examples:
|
|
instruct = "Speak slowly and clearly"
|
|
instruct = "Sound excited and enthusiastic"
|
|
instruct = "Whisper softly"
|
|
```
|
|
|
|
## Caching Strategy
|
|
|
|
Voice prompts are cached to avoid recomputation:
|
|
|
|
```python
|
|
def get_cache_key(audio_path: str, reference_text: str) -> str:
|
|
"""Generate cache key from audio hash and text."""
|
|
audio_hash = hashlib.md5(Path(audio_path).read_bytes()).hexdigest()
|
|
text_hash = hashlib.md5(reference_text.encode()).hexdigest()
|
|
return f"{audio_hash}_{text_hash}"
|
|
```
|
|
|
|
Cache is stored in `data/cache/voice_prompts/`.
|
|
|
|
## Memory Management
|
|
|
|
### Unloading Models
|
|
|
|
Free VRAM/RAM when not needed:
|
|
|
|
```python
|
|
def unload_model(self):
|
|
if self.model is not None:
|
|
del self.model
|
|
self.model = None
|
|
|
|
if torch.cuda.is_available():
|
|
torch.cuda.empty_cache()
|
|
```
|
|
|
|
### Model Switching
|
|
|
|
When switching between model sizes (1.7B ↔ 0.6B):
|
|
|
|
```python
|
|
# Unload existing model first
|
|
if self.model is not None and self._current_model_size != model_size:
|
|
self.unload_model()
|
|
```
|
|
|
|
## Generation Flow
|
|
|
|
1. **Request** → Validate text and profile ID
|
|
2. **Profile** → Load profile samples from database
|
|
3. **Voice Prompt** → Create or retrieve cached prompt
|
|
4. **Generate** → Run TTS inference
|
|
5. **Save** → Write audio to generations directory
|
|
6. **Record** → Create history entry in database
|
|
7. **Response** → Return audio path and metadata
|
|
|
|
## API Endpoints
|
|
|
|
| Method | Endpoint | Description |
|
|
|--------|----------|-------------|
|
|
| POST | `/generate` | Generate speech from text |
|
|
| GET | `/audio/{id}` | Serve generated audio file |
|
|
|
|
### Request Schema
|
|
|
|
```json
|
|
{
|
|
"profile_id": "uuid",
|
|
"text": "Text to synthesize",
|
|
"language": "en",
|
|
"seed": 42,
|
|
"model_size": "1.7B",
|
|
"instruct": "Speak clearly"
|
|
}
|
|
```
|
|
|
|
### Response Schema
|
|
|
|
```json
|
|
{
|
|
"id": "generation_uuid",
|
|
"profile_id": "profile_uuid",
|
|
"text": "Text to synthesize",
|
|
"language": "en",
|
|
"audio_path": "/path/to/audio.wav",
|
|
"duration": 3.5,
|
|
"seed": 42,
|
|
"instruct": "Speak clearly",
|
|
"created_at": "2024-01-15T10:30:00Z"
|
|
}
|
|
```
|
|
|
|
## Performance Considerations
|
|
|
|
### GPU Acceleration
|
|
|
|
- CUDA provides fastest inference
|
|
- MPS (Apple Silicon) has stability issues, uses CPU fallback
|
|
- CPU inference is slower but always works
|
|
|
|
### Batch Size
|
|
|
|
Currently generates one utterance at a time. For long texts, consider:
|
|
- Splitting into sentences
|
|
- Sequential generation
|
|
- Concatenating results
|
|
|
|
### Memory Usage
|
|
|
|
| Model | VRAM/RAM Required |
|
|
|-------|-------------------|
|
|
| 0.6B | ~2GB |
|
|
| 1.7B | ~6GB |
|