mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-09-16 05:10:42 -07:00
Add .npmrc for bun usage and update dependencies
- Created a new .npmrc file to enforce bun usage. - Bumped version numbers for multiple packages to 0.1.9 in bun.lock. - Added react-sound-visualizer dependency to enhance audio visualization features. - Introduced convert:assets script in package.json for asset optimization. - Updated CONTRIBUTING.md with instructions for converting assets to web formats. - Added documentation files for API endpoints and developer guidelines in the docs directory.
This commit is contained in:
@@ -0,0 +1,202 @@
|
||||
---
|
||||
title: "Architecture"
|
||||
description: "Understanding Voicebox's technical architecture"
|
||||
---
|
||||
|
||||
## System Overview
|
||||
|
||||
Voicebox uses a client-server architecture with a React frontend and Python backend. The desktop app is built with Tauri and contains two main layers:
|
||||
|
||||
**Frontend Layer:** A React application that handles the UI components, state management with Zustand, and data fetching with React Query (TanStack Query).
|
||||
|
||||
**Backend Layer:** A Python FastAPI server that provides the REST API, runs the TTS engine (Qwen3-TTS), manages the SQLite database, and handles audio processing.
|
||||
|
||||
These two layers communicate via HTTP, with the frontend making API requests to the backend.
|
||||
|
||||
## Frontend Architecture
|
||||
|
||||
### Tech Stack
|
||||
|
||||
- **Framework**: React 18 with TypeScript
|
||||
- **State Management**: Zustand stores
|
||||
- **Data Fetching**: React Query (TanStack Query)
|
||||
- **Styling**: Tailwind CSS
|
||||
- **Audio**: WaveSurfer.js
|
||||
- **Desktop**: Tauri (Rust)
|
||||
|
||||
### Component Structure
|
||||
|
||||
```
|
||||
app/src/
|
||||
├── components/ # React components
|
||||
│ ├── profiles/ # Voice profile UI
|
||||
│ ├── generation/ # Speech generation UI
|
||||
│ ├── stories/ # Timeline editor
|
||||
│ └── shared/ # Reusable components
|
||||
├── lib/ # Utilities
|
||||
│ ├── api/ # Generated API client
|
||||
│ └── utils/ # Helper functions
|
||||
├── hooks/ # React hooks
|
||||
└── stores/ # Zustand state stores
|
||||
```
|
||||
|
||||
### State Management
|
||||
|
||||
```typescript
|
||||
// Example: Profile store
|
||||
const useProfileStore = create((set) => ({
|
||||
profiles: [],
|
||||
selectedProfile: null,
|
||||
setProfiles: (profiles) => set({ profiles }),
|
||||
selectProfile: (id) => set({ selectedProfile: id })
|
||||
}))
|
||||
```
|
||||
|
||||
## Backend Architecture
|
||||
|
||||
### Tech Stack
|
||||
|
||||
- **Framework**: FastAPI (Python 3.11+)
|
||||
- **TTS Model**: Qwen3-TTS
|
||||
- **Transcription**: Whisper
|
||||
- **Database**: SQLite
|
||||
- **Audio**: librosa, soundfile
|
||||
|
||||
### API Structure
|
||||
|
||||
```python
|
||||
# main.py - API routes
|
||||
@app.post("/generate")
|
||||
async def generate_speech(request: GenerateRequest):
|
||||
# 1. Validate request
|
||||
# 2. Load voice profile
|
||||
# 3. Generate audio with TTS
|
||||
# 4. Save to database
|
||||
# 5. Return response
|
||||
```
|
||||
|
||||
### Data Model
|
||||
|
||||
The database uses three main tables:
|
||||
|
||||
**Profile Table:** Stores voice profiles with fields for id, name, and language.
|
||||
|
||||
**Sample Table:** Stores audio samples linked to profiles via profile_id, with fields for audio_path and duration.
|
||||
|
||||
**Generation Table:** Stores generated audio with fields for id, profile_id, text, and audio_path.
|
||||
|
||||
## Desktop App (Tauri)
|
||||
|
||||
### Rust Backend
|
||||
|
||||
```rust
|
||||
// Sidecar process management
|
||||
// File system access
|
||||
// Native integrations
|
||||
```
|
||||
|
||||
### Responsibilities
|
||||
|
||||
- Launch Python backend as sidecar process
|
||||
- Native file dialogs
|
||||
- System tray integration
|
||||
- Auto-updates
|
||||
- OS-specific features
|
||||
|
||||
## Build Process
|
||||
|
||||
### Development
|
||||
|
||||
```bash
|
||||
# Frontend (Vite dev server)
|
||||
cd app && bun run dev
|
||||
|
||||
# Backend (manual start)
|
||||
cd backend && uvicorn main:app --reload
|
||||
|
||||
# Desktop app (connects to manual backend)
|
||||
bun run dev
|
||||
```
|
||||
|
||||
### Production
|
||||
|
||||
```bash
|
||||
# 1. Build server binary (PyInstaller)
|
||||
./scripts/build-server.sh
|
||||
|
||||
# 2. Build Tauri app (includes server)
|
||||
cd tauri && bun run tauri build
|
||||
```
|
||||
|
||||
## Data Flow
|
||||
|
||||
### Generation Flow
|
||||
|
||||
When a user generates speech, the data flows through the following stages:
|
||||
|
||||
1. **User Input** - User enters text in a React component
|
||||
2. **State Update** - Text is stored in Zustand state
|
||||
3. **API Request** - React Query mutation triggers an API call via fetch
|
||||
4. **Backend Processing** - FastAPI endpoint receives the request
|
||||
5. **TTS Generation** - Qwen3-TTS model generates the audio
|
||||
6. **Storage** - Audio file is saved to disk and a database record is created
|
||||
7. **Response** - Backend returns the audio URL
|
||||
8. **Cache Update** - React Query updates its cache with the response
|
||||
9. **UI Update** - Component re-renders with new data
|
||||
10. **Playback** - User can play the generated audio
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
### Frontend
|
||||
|
||||
- **Code splitting** - Lazy load routes
|
||||
- **Memoization** - React.memo for heavy components
|
||||
- **Virtual scrolling** - For large lists
|
||||
- **Debouncing** - Search and input handling
|
||||
|
||||
### Backend
|
||||
|
||||
- **Async operations** - All I/O is async
|
||||
- **Model caching** - Keep TTS model in memory
|
||||
- **Voice prompt caching** - Reuse embeddings
|
||||
- **Connection pooling** - Database connections
|
||||
|
||||
## Security
|
||||
|
||||
### Current
|
||||
|
||||
- Local-only by default
|
||||
- No authentication (localhost trust)
|
||||
- File system sandboxing via Tauri
|
||||
|
||||
### Planned
|
||||
|
||||
- API key authentication
|
||||
- User accounts
|
||||
- Rate limiting
|
||||
- HTTPS support
|
||||
|
||||
## Deployment Modes
|
||||
|
||||
### Local Mode
|
||||
|
||||
- Backend runs as sidecar
|
||||
- All data stays on device
|
||||
- No network required
|
||||
|
||||
### Remote Mode
|
||||
|
||||
- Backend on separate machine
|
||||
- Frontend connects via HTTP
|
||||
- Shared infrastructure possible
|
||||
|
||||
## Next Steps
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Development Setup" icon="code" href="/development/setup">
|
||||
Set up your dev environment
|
||||
</Card>
|
||||
<Card title="Contributing" icon="code-pull-request" href="/development/contributing">
|
||||
Contribute to Voicebox
|
||||
</Card>
|
||||
</CardGroup>
|
||||
@@ -0,0 +1,310 @@
|
||||
---
|
||||
title: "Audio Channels"
|
||||
description: "How audio output routing works in Voicebox"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Audio channels allow routing voice output to different audio devices. This is useful for multi-output setups where different voices should play through different speakers or applications.
|
||||
|
||||
## Architecture
|
||||
|
||||
**Channel:** A named audio bus that can be assigned to output devices.
|
||||
|
||||
**Device Mapping:** Links channels to OS audio device identifiers.
|
||||
|
||||
**Profile Mapping:** Links voice profiles to channels (many-to-many).
|
||||
|
||||
## Data Model
|
||||
|
||||
### AudioChannel Table
|
||||
|
||||
```python
|
||||
class AudioChannel(Base):
|
||||
__tablename__ = "audio_channels"
|
||||
|
||||
id = Column(String, primary_key=True)
|
||||
name = Column(String, nullable=False)
|
||||
is_default = Column(Boolean, default=False)
|
||||
created_at = Column(DateTime)
|
||||
```
|
||||
|
||||
### ChannelDeviceMapping Table
|
||||
|
||||
```python
|
||||
class ChannelDeviceMapping(Base):
|
||||
__tablename__ = "channel_device_mappings"
|
||||
|
||||
id = Column(String, primary_key=True)
|
||||
channel_id = Column(String, ForeignKey("audio_channels.id"))
|
||||
device_id = Column(String) # OS device identifier
|
||||
```
|
||||
|
||||
### ProfileChannelMapping Table
|
||||
|
||||
```python
|
||||
class ProfileChannelMapping(Base):
|
||||
__tablename__ = "profile_channel_mappings"
|
||||
|
||||
profile_id = Column(String, ForeignKey("profiles.id"), primary_key=True)
|
||||
channel_id = Column(String, ForeignKey("audio_channels.id"), primary_key=True)
|
||||
```
|
||||
|
||||
## Default Channel
|
||||
|
||||
A default channel is created on database initialization:
|
||||
|
||||
```python
|
||||
def init_db():
|
||||
# Create default channel if it doesn't exist
|
||||
default_channel = db.query(AudioChannel).filter(
|
||||
AudioChannel.is_default == True
|
||||
).first()
|
||||
|
||||
if not default_channel:
|
||||
default_channel = AudioChannel(
|
||||
id=str(uuid.uuid4()),
|
||||
name="Default",
|
||||
is_default=True
|
||||
)
|
||||
db.add(default_channel)
|
||||
|
||||
# Assign all existing profiles to default channel
|
||||
profiles = db.query(VoiceProfile).all()
|
||||
for profile in profiles:
|
||||
mapping = ProfileChannelMapping(
|
||||
profile_id=profile.id,
|
||||
channel_id=default_channel.id
|
||||
)
|
||||
db.add(mapping)
|
||||
```
|
||||
|
||||
## Core Operations
|
||||
|
||||
### Creating a Channel
|
||||
|
||||
```python
|
||||
async def create_channel(
|
||||
data: AudioChannelCreate,
|
||||
db: Session,
|
||||
) -> AudioChannelResponse:
|
||||
# Check name uniqueness
|
||||
existing = db.query(DBAudioChannel).filter_by(name=data.name).first()
|
||||
if existing:
|
||||
raise ValueError(f"Channel with name '{data.name}' already exists")
|
||||
|
||||
# Create channel
|
||||
channel = DBAudioChannel(
|
||||
id=str(uuid.uuid4()),
|
||||
name=data.name,
|
||||
is_default=False,
|
||||
)
|
||||
db.add(channel)
|
||||
|
||||
# Add device mappings
|
||||
for device_id in data.device_ids:
|
||||
mapping = DBChannelDeviceMapping(
|
||||
id=str(uuid.uuid4()),
|
||||
channel_id=channel.id,
|
||||
device_id=device_id,
|
||||
)
|
||||
db.add(mapping)
|
||||
|
||||
db.commit()
|
||||
```
|
||||
|
||||
### Updating a Channel
|
||||
|
||||
```python
|
||||
async def update_channel(
|
||||
channel_id: str,
|
||||
data: AudioChannelUpdate,
|
||||
db: Session,
|
||||
) -> AudioChannelResponse:
|
||||
channel = db.query(DBAudioChannel).filter_by(id=channel_id).first()
|
||||
|
||||
# Cannot modify default channel
|
||||
if channel.is_default:
|
||||
raise ValueError("Cannot modify the default channel")
|
||||
|
||||
# Update name
|
||||
if data.name is not None:
|
||||
channel.name = data.name
|
||||
|
||||
# Update device mappings
|
||||
if data.device_ids is not None:
|
||||
# Delete existing
|
||||
db.query(DBChannelDeviceMapping).filter_by(channel_id=channel_id).delete()
|
||||
|
||||
# Add new
|
||||
for device_id in data.device_ids:
|
||||
mapping = DBChannelDeviceMapping(
|
||||
channel_id=channel.id,
|
||||
device_id=device_id,
|
||||
)
|
||||
db.add(mapping)
|
||||
|
||||
db.commit()
|
||||
```
|
||||
|
||||
### Deleting a Channel
|
||||
|
||||
```python
|
||||
async def delete_channel(channel_id: str, db: Session) -> bool:
|
||||
channel = db.query(DBAudioChannel).filter_by(id=channel_id).first()
|
||||
|
||||
# Cannot delete default channel
|
||||
if channel.is_default:
|
||||
raise ValueError("Cannot delete the default channel")
|
||||
|
||||
# Delete device mappings
|
||||
db.query(DBChannelDeviceMapping).filter_by(channel_id=channel_id).delete()
|
||||
|
||||
# Delete profile-channel mappings
|
||||
db.query(DBProfileChannelMapping).filter_by(channel_id=channel_id).delete()
|
||||
|
||||
# Delete channel
|
||||
db.delete(channel)
|
||||
db.commit()
|
||||
```
|
||||
|
||||
## Voice Assignment
|
||||
|
||||
### Assigning Voices to Channel
|
||||
|
||||
```python
|
||||
async def set_channel_voices(
|
||||
channel_id: str,
|
||||
data: ChannelVoiceAssignment,
|
||||
db: Session,
|
||||
) -> None:
|
||||
# Verify channel exists
|
||||
channel = db.query(DBAudioChannel).filter_by(id=channel_id).first()
|
||||
if not channel:
|
||||
raise ValueError(f"Channel {channel_id} not found")
|
||||
|
||||
# Verify all profiles exist
|
||||
for profile_id in data.profile_ids:
|
||||
profile = db.query(DBVoiceProfile).filter_by(id=profile_id).first()
|
||||
if not profile:
|
||||
raise ValueError(f"Profile {profile_id} not found")
|
||||
|
||||
# Delete existing mappings
|
||||
db.query(DBProfileChannelMapping).filter_by(channel_id=channel_id).delete()
|
||||
|
||||
# Add new mappings
|
||||
for profile_id in data.profile_ids:
|
||||
mapping = DBProfileChannelMapping(
|
||||
profile_id=profile_id,
|
||||
channel_id=channel_id,
|
||||
)
|
||||
db.add(mapping)
|
||||
|
||||
db.commit()
|
||||
```
|
||||
|
||||
### Assigning Channels to Voice
|
||||
|
||||
```python
|
||||
async def set_profile_channels(
|
||||
profile_id: str,
|
||||
data: ProfileChannelAssignment,
|
||||
db: Session,
|
||||
) -> None:
|
||||
# Verify profile exists
|
||||
profile = db.query(DBVoiceProfile).filter_by(id=profile_id).first()
|
||||
if not profile:
|
||||
raise ValueError(f"Profile {profile_id} not found")
|
||||
|
||||
# Delete existing mappings
|
||||
db.query(DBProfileChannelMapping).filter_by(profile_id=profile_id).delete()
|
||||
|
||||
# Add new mappings
|
||||
for channel_id in data.channel_ids:
|
||||
mapping = DBProfileChannelMapping(
|
||||
profile_id=profile_id,
|
||||
channel_id=channel_id,
|
||||
)
|
||||
db.add(mapping)
|
||||
|
||||
db.commit()
|
||||
```
|
||||
|
||||
## API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| GET | `/channels` | List all channels |
|
||||
| POST | `/channels` | Create a channel |
|
||||
| GET | `/channels/{id}` | Get channel by ID |
|
||||
| PUT | `/channels/{id}` | Update channel |
|
||||
| DELETE | `/channels/{id}` | Delete channel |
|
||||
| GET | `/channels/{id}/voices` | Get assigned voices |
|
||||
| PUT | `/channels/{id}/voices` | Set assigned voices |
|
||||
| GET | `/profiles/{id}/channels` | Get profile's channels |
|
||||
| PUT | `/profiles/{id}/channels` | Set profile's channels |
|
||||
|
||||
## Request/Response Schemas
|
||||
|
||||
### AudioChannelCreate
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "Speakers",
|
||||
"device_ids": ["device_uuid_1", "device_uuid_2"]
|
||||
}
|
||||
```
|
||||
|
||||
### AudioChannelResponse
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "channel_uuid",
|
||||
"name": "Speakers",
|
||||
"is_default": false,
|
||||
"device_ids": ["device_uuid_1", "device_uuid_2"],
|
||||
"created_at": "2024-01-15T10:30:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
### ChannelVoiceAssignment
|
||||
|
||||
```json
|
||||
{
|
||||
"profile_ids": ["profile_1", "profile_2"]
|
||||
}
|
||||
```
|
||||
|
||||
## Use Cases
|
||||
|
||||
### Multi-Output Setup
|
||||
|
||||
**Scenario:** Stream with different voice characters
|
||||
|
||||
1. Create "Stream" channel → OBS virtual audio
|
||||
2. Create "Monitor" channel → Headphones
|
||||
3. Assign "Narrator" profile → Both channels
|
||||
4. Assign "Character 1" profile → Stream only
|
||||
|
||||
### Virtual Audio Cables
|
||||
|
||||
Common device IDs for virtual audio:
|
||||
- VB-Audio Virtual Cable
|
||||
- BlackHole (macOS)
|
||||
- Soundflower (macOS)
|
||||
|
||||
## Frontend Integration
|
||||
|
||||
The frontend needs to:
|
||||
|
||||
1. **Enumerate devices** using Web Audio API or Tauri
|
||||
2. **Display channel list** with device assignments
|
||||
3. **Allow profile assignment** via drag/drop or dropdown
|
||||
4. **Route playback** to correct device based on profile's channel
|
||||
|
||||
## Limitations
|
||||
|
||||
- Device IDs are OS-specific
|
||||
- Hot-plugging may invalidate device IDs
|
||||
- Default channel cannot be modified/deleted
|
||||
- Frontend handles actual audio routing (backend just stores config)
|
||||
@@ -0,0 +1,84 @@
|
||||
---
|
||||
title: "Auto-Updater"
|
||||
description: "Configure and use the Tauri auto-updater"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Voicebox uses Tauri's built-in auto-updater to deliver updates to users automatically.
|
||||
|
||||
## Quick Reference
|
||||
|
||||
For detailed setup instructions, see the existing documentation:
|
||||
|
||||
- [AUTOUPDATER_QUICKSTART.md](https://github.com/jamiepine/voicebox/blob/main/docs/AUTOUPDATER_QUICKSTART.md)
|
||||
- [AUTOUPDATER.md](https://github.com/jamiepine/voicebox/blob/main/docs/AUTOUPDATER.md)
|
||||
|
||||
## How It Works
|
||||
|
||||
The auto-updater follows a secure update process:
|
||||
|
||||
1. **Check for Updates** - The Voicebox app periodically checks GitHub Releases for new versions
|
||||
2. **Download Update** - If a new version is found, the update package is downloaded
|
||||
3. **Verify Signature** - The downloaded package is cryptographically verified using the public key
|
||||
4. **Install** - After verification, the update is installed
|
||||
5. **Restart** - The app restarts with the new version
|
||||
|
||||
## Configuration
|
||||
|
||||
Updates are configured in `tauri/src-tauri/tauri.conf.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"updater": {
|
||||
"active": true,
|
||||
"endpoints": [
|
||||
"https://github.com/jamiepine/voicebox/releases/latest/download/latest.json"
|
||||
],
|
||||
"dialog": true,
|
||||
"pubkey": "YOUR_PUBLIC_KEY"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Generating Keys
|
||||
|
||||
```bash
|
||||
# Generate signing keys
|
||||
bun run generate:keys
|
||||
|
||||
# Keys saved to ~/.tauri/voicebox.key
|
||||
```
|
||||
|
||||
<Warning>
|
||||
Keep your private key secure! Never commit it to the repository.
|
||||
</Warning>
|
||||
|
||||
## Release Process
|
||||
|
||||
1. **Bump version** using bumpversion
|
||||
2. **Push tag** to trigger CI/CD
|
||||
3. **GitHub Actions** builds and signs releases
|
||||
4. **Users** receive update notification
|
||||
|
||||
## User Experience
|
||||
|
||||
When an update is available:
|
||||
|
||||
1. User sees a notification dialog
|
||||
2. User clicks "Update"
|
||||
3. Update downloads in background
|
||||
4. App restarts with new version
|
||||
|
||||
## For Developers
|
||||
|
||||
See the full documentation files for:
|
||||
|
||||
- Setting up signing keys
|
||||
- Configuring GitHub releases
|
||||
- Testing updates locally
|
||||
- Troubleshooting update failures
|
||||
|
||||
<Card title="View Full Docs" href="https://github.com/jamiepine/voicebox/tree/main/docs">
|
||||
Access AUTOUPDATER.md and AUTOUPDATER_QUICKSTART.md in the repository
|
||||
</Card>
|
||||
@@ -0,0 +1,263 @@
|
||||
---
|
||||
title: "Building"
|
||||
description: "Build Voicebox for production"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Voicebox uses a multi-step build process to create platform-specific installers.
|
||||
|
||||
## Quick Build
|
||||
|
||||
```bash
|
||||
# Build for your current platform
|
||||
make build
|
||||
|
||||
# Or manually
|
||||
cd tauri && bun run tauri build
|
||||
```
|
||||
|
||||
## Build Steps
|
||||
|
||||
### 1. Build Server Binary
|
||||
|
||||
The Python backend must be compiled into a standalone executable first:
|
||||
|
||||
```bash
|
||||
./scripts/build-server.sh
|
||||
```
|
||||
|
||||
This uses PyInstaller to create a binary in `tauri/src-tauri/binaries/`.
|
||||
|
||||
**Platform-specific binaries:**
|
||||
- macOS: `voicebox-server-aarch64-apple-darwin` or `voicebox-server-x86_64-apple-darwin`
|
||||
- Windows: `voicebox-server-x86_64-pc-windows-msvc.exe`
|
||||
- Linux: `voicebox-server-x86_64-unknown-linux-gnu`
|
||||
|
||||
<Note>
|
||||
The build script automatically detects your platform and creates the appropriate binary.
|
||||
</Note>
|
||||
|
||||
### 2. Build Tauri App
|
||||
|
||||
```bash
|
||||
cd tauri
|
||||
bun run tauri build
|
||||
```
|
||||
|
||||
This will:
|
||||
1. Build the React frontend (Vite)
|
||||
2. Compile the Rust backend
|
||||
3. Bundle the server binary as a sidecar
|
||||
4. Create platform-specific installers
|
||||
|
||||
### 3. Output
|
||||
|
||||
Installers are created in `tauri/src-tauri/target/release/bundle/`:
|
||||
|
||||
**macOS:**
|
||||
- `dmg/` - Disk image installer
|
||||
- `macos/` - App bundle
|
||||
|
||||
**Windows:**
|
||||
- `msi/` - MSI installer
|
||||
- `nsis/` - NSIS installer
|
||||
|
||||
**Linux:**
|
||||
- `deb/` - Debian package
|
||||
- `appimage/` - AppImage
|
||||
|
||||
## Advanced Options
|
||||
|
||||
### Building for Specific Platform
|
||||
|
||||
```bash
|
||||
# Build for macOS (Apple Silicon)
|
||||
bun run tauri build -- --target aarch64-apple-darwin
|
||||
|
||||
# Build for macOS (Intel)
|
||||
bun run tauri build -- --target x86_64-apple-darwin
|
||||
|
||||
# Build for Windows
|
||||
bun run tauri build -- --target x86_64-pc-windows-msvc
|
||||
|
||||
# Build for Linux
|
||||
bun run tauri build -- --target x86_64-unknown-linux-gnu
|
||||
```
|
||||
|
||||
### Using Local Qwen3-TTS
|
||||
|
||||
If you're developing Qwen3-TTS locally:
|
||||
|
||||
```bash
|
||||
export QWEN_TTS_PATH=~/path/to/Qwen3-TTS
|
||||
./scripts/build-server.sh
|
||||
```
|
||||
|
||||
This makes PyInstaller use your local version instead of the pip package.
|
||||
|
||||
### Debug Build
|
||||
|
||||
```bash
|
||||
cd tauri
|
||||
bun run tauri build --debug
|
||||
```
|
||||
|
||||
Creates a debug build with symbols and logging.
|
||||
|
||||
## Build Configuration
|
||||
|
||||
### Tauri Config
|
||||
|
||||
Edit `tauri/src-tauri/tauri.conf.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"bundle": {
|
||||
"identifier": "com.voicebox.app",
|
||||
"icon": [
|
||||
"icons/32x32.png",
|
||||
"icons/128x128.png",
|
||||
"icons/icon.icns",
|
||||
"icons/icon.ico"
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Sidecar Configuration
|
||||
|
||||
The Python server is bundled as a sidecar:
|
||||
|
||||
```json
|
||||
{
|
||||
"tauri": {
|
||||
"bundle": {
|
||||
"externalBin": [
|
||||
"binaries/voicebox-server"
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Code Signing
|
||||
|
||||
### macOS
|
||||
|
||||
To sign the app for distribution:
|
||||
|
||||
```bash
|
||||
# Set signing identity
|
||||
export APPLE_SIGNING_IDENTITY="Developer ID Application: Your Name"
|
||||
|
||||
# Build with signing
|
||||
bun run tauri build
|
||||
```
|
||||
|
||||
For notarization:
|
||||
|
||||
```bash
|
||||
# Set credentials
|
||||
export APPLE_ID="[email protected]"
|
||||
export APPLE_PASSWORD="app-specific-password"
|
||||
|
||||
# Build and notarize
|
||||
bun run tauri build
|
||||
```
|
||||
|
||||
### Windows
|
||||
|
||||
For Windows code signing:
|
||||
|
||||
```bash
|
||||
# Set certificate
|
||||
export WINDOWS_CERTIFICATE_PATH="/path/to/cert.pfx"
|
||||
export WINDOWS_CERTIFICATE_PASSWORD="password"
|
||||
|
||||
# Build with signing
|
||||
bun run tauri build
|
||||
```
|
||||
|
||||
## Release Process
|
||||
|
||||
The full release process is automated:
|
||||
|
||||
```bash
|
||||
# 1. Bump version
|
||||
bumpversion patch # or minor/major
|
||||
|
||||
# 2. Build all platforms (CI/CD handles this)
|
||||
git push --tags
|
||||
|
||||
# 3. GitHub Actions creates releases
|
||||
```
|
||||
|
||||
See [CONTRIBUTING.md](/development/contributing) for the full release workflow.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Server Binary Build Fails">
|
||||
**Common issues:**
|
||||
- Missing Python dependencies: `pip install -r requirements.txt`
|
||||
- PyInstaller not found: `pip install pyinstaller`
|
||||
- Qwen3-TTS not installed: `pip install git+https://github.com/QwenLM/Qwen3-TTS.git`
|
||||
|
||||
**Solution:**
|
||||
```bash
|
||||
cd backend
|
||||
source venv/bin/activate
|
||||
pip install -r requirements.txt
|
||||
pip install pyinstaller
|
||||
```
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tauri Build Fails">
|
||||
**Common issues:**
|
||||
- Rust not installed: `curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh`
|
||||
- Server binary missing: Run `./scripts/build-server.sh` first
|
||||
- Node modules outdated: `bun install`
|
||||
|
||||
**Solution:**
|
||||
```bash
|
||||
# Clean and rebuild
|
||||
cd tauri/src-tauri
|
||||
cargo clean
|
||||
cd ../..
|
||||
./scripts/build-server.sh
|
||||
bun run tauri build
|
||||
```
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="App Won't Launch After Build">
|
||||
**Check:**
|
||||
- Server binary has execute permissions
|
||||
- All dependencies are bundled
|
||||
- Check logs in the app's data directory
|
||||
|
||||
**macOS:**
|
||||
```bash
|
||||
tail -f ~/Library/Application\ Support/com.voicebox.app/logs/server.log
|
||||
```
|
||||
|
||||
**Windows:**
|
||||
```bash
|
||||
type %APPDATA%\com.voicebox.app\logs\server.log
|
||||
```
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
## CI/CD
|
||||
|
||||
GitHub Actions automatically builds releases when tags are pushed:
|
||||
|
||||
```yaml
|
||||
# .github/workflows/release.yml
|
||||
on:
|
||||
push:
|
||||
tags:
|
||||
- 'v*'
|
||||
```
|
||||
|
||||
See the [repository](https://github.com/jamiepine/voicebox) for the full CI/CD configuration.
|
||||
@@ -0,0 +1,326 @@
|
||||
---
|
||||
title: "Contributing"
|
||||
description: "How to contribute to Voicebox"
|
||||
---
|
||||
|
||||
Thank you for your interest in contributing to Voicebox! This guide will help you get started.
|
||||
|
||||
## Code of Conduct
|
||||
|
||||
- Be respectful and inclusive
|
||||
- Welcome newcomers and help them learn
|
||||
- Focus on constructive feedback
|
||||
- Respect different viewpoints and experiences
|
||||
|
||||
## Getting Started
|
||||
|
||||
Before you start contributing, make sure you have:
|
||||
|
||||
1. **Read the documentation** to understand how Voicebox works
|
||||
2. **Set up your development environment** - see [Development Setup](/development/setup)
|
||||
3. **Explored the codebase** to understand the project structure
|
||||
4. **Checked existing issues** to see if someone else is working on something similar
|
||||
|
||||
## Ways to Contribute
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Report Bugs" icon="bug">
|
||||
Found a bug? Open an issue with reproduction steps
|
||||
</Card>
|
||||
<Card title="Request Features" icon="lightbulb">
|
||||
Have an idea? Start a discussion or open an issue
|
||||
</Card>
|
||||
<Card title="Improve Docs" icon="book">
|
||||
Fix typos, add examples, or clarify instructions
|
||||
</Card>
|
||||
<Card title="Write Code" icon="code">
|
||||
Fix bugs, add features, or optimize performance
|
||||
</Card>
|
||||
</CardGroup>
|
||||
|
||||
## Development Workflow
|
||||
|
||||
### 1. Fork & Clone
|
||||
|
||||
```bash
|
||||
# Fork the repository on GitHub
|
||||
# Then clone your fork
|
||||
git clone https://github.com/YOUR_USERNAME/voicebox.git
|
||||
cd voicebox
|
||||
```
|
||||
|
||||
### 2. Create a Branch
|
||||
|
||||
Use descriptive branch names:
|
||||
|
||||
```bash
|
||||
# For features
|
||||
git checkout -b feature/voice-effects
|
||||
|
||||
# For bug fixes
|
||||
git checkout -b fix/audio-playback-issue
|
||||
|
||||
# For documentation
|
||||
git checkout -b docs/api-examples
|
||||
```
|
||||
|
||||
### 3. Make Your Changes
|
||||
|
||||
Follow these guidelines:
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Code Style">
|
||||
**TypeScript/React:**
|
||||
- Use TypeScript strict mode
|
||||
- Prefer functional components with hooks
|
||||
- Use named exports
|
||||
- Format with Biome (runs automatically)
|
||||
|
||||
**Python:**
|
||||
- Follow PEP 8
|
||||
- Use type hints
|
||||
- Use async/await for I/O
|
||||
- Document functions with docstrings
|
||||
|
||||
**Rust:**
|
||||
- Follow Rust conventions
|
||||
- Use meaningful names
|
||||
- Handle errors explicitly
|
||||
- Run `rustfmt`
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Commit Messages">
|
||||
Write clear, descriptive commit messages:
|
||||
|
||||
```bash
|
||||
# Good
|
||||
git commit -m "Add voice profile export feature"
|
||||
git commit -m "Fix audio playback stopping after 30 seconds"
|
||||
|
||||
# Avoid
|
||||
git commit -m "Update code"
|
||||
git commit -m "Fix bug"
|
||||
```
|
||||
|
||||
Format:
|
||||
- Use imperative mood ("Add feature" not "Added feature")
|
||||
- Keep first line under 50 characters
|
||||
- Add detailed description if needed
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Testing">
|
||||
- Test your changes manually in the app
|
||||
- Ensure backend API endpoints work
|
||||
- Check for TypeScript/Python errors
|
||||
- Verify UI components render correctly
|
||||
- Add automated tests when possible
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
### 4. Push & Create PR
|
||||
|
||||
```bash
|
||||
# Push your branch
|
||||
git push origin feature/your-feature-name
|
||||
|
||||
# Then create a pull request on GitHub
|
||||
```
|
||||
|
||||
## Pull Request Guidelines
|
||||
|
||||
When creating a pull request:
|
||||
|
||||
<Steps>
|
||||
<Step title="Use a Clear Title">
|
||||
Examples:
|
||||
- "Add voice profile export functionality"
|
||||
- "Fix audio playback stopping after 30 seconds"
|
||||
- "Improve generation speed with caching"
|
||||
</Step>
|
||||
|
||||
<Step title="Provide Description">
|
||||
Include:
|
||||
- What changes you made
|
||||
- Why you made them
|
||||
- How to test them
|
||||
- Screenshots (for UI changes)
|
||||
- Reference related issues
|
||||
</Step>
|
||||
|
||||
<Step title="Update Documentation">
|
||||
- Update relevant docs if behavior changes
|
||||
- Add API documentation for new endpoints
|
||||
- Update README if needed
|
||||
</Step>
|
||||
|
||||
<Step title="Check the Checklist">
|
||||
- [ ] Code follows style guidelines
|
||||
- [ ] Documentation updated
|
||||
- [ ] Changes tested
|
||||
- [ ] No breaking changes (or documented)
|
||||
- [ ] CHANGELOG.md updated
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Project Structure
|
||||
|
||||
Understanding the codebase:
|
||||
|
||||
```
|
||||
voicebox/
|
||||
├── app/ # Shared React frontend
|
||||
│ ├── src/
|
||||
│ │ ├── components/ # UI components
|
||||
│ │ ├── lib/ # Utilities and API client
|
||||
│ │ ├── hooks/ # React hooks
|
||||
│ │ └── stores/ # Zustand state stores
|
||||
├── backend/ # Python FastAPI server
|
||||
│ ├── main.py # API routes
|
||||
│ ├── tts.py # Voice synthesis logic
|
||||
│ ├── database.py # SQLite operations
|
||||
│ └── models.py # Pydantic models
|
||||
├── tauri/ # Desktop app wrapper
|
||||
│ └── src-tauri/ # Rust backend
|
||||
├── web/ # Web deployment
|
||||
├── landing/ # Marketing website
|
||||
└── scripts/ # Build & release scripts
|
||||
```
|
||||
|
||||
## Areas for Contribution
|
||||
|
||||
### Bug Fixes
|
||||
|
||||
- Check [existing issues](https://github.com/jamiepine/voicebox/issues) for bugs
|
||||
- Test your fix thoroughly
|
||||
- Add regression tests if possible
|
||||
|
||||
### New Features
|
||||
|
||||
- Check the [roadmap](https://github.com/jamiepine/voicebox#roadmap) for planned features
|
||||
- Discuss major features in an issue first
|
||||
- Keep features focused and well-scoped
|
||||
|
||||
### Documentation
|
||||
|
||||
- Improve clarity and fix typos
|
||||
- Add code examples
|
||||
- Create tutorials or guides
|
||||
- Document API endpoints
|
||||
|
||||
### UI/UX Improvements
|
||||
|
||||
- Improve accessibility
|
||||
- Enhance visual design
|
||||
- Optimize performance
|
||||
- Add animations/transitions
|
||||
|
||||
### Infrastructure
|
||||
|
||||
- Improve build process
|
||||
- Add CI/CD improvements
|
||||
- Optimize bundle size
|
||||
- Add testing infrastructure
|
||||
|
||||
## API Development
|
||||
|
||||
When adding new API endpoints:
|
||||
|
||||
<Steps>
|
||||
<Step title="Add Route">
|
||||
In `backend/main.py`:
|
||||
|
||||
```python
|
||||
@app.post("/api/new-endpoint")
|
||||
async def new_endpoint(data: RequestModel) -> ResponseModel:
|
||||
"""Endpoint description."""
|
||||
# Implementation
|
||||
return response
|
||||
```
|
||||
</Step>
|
||||
|
||||
<Step title="Create Models">
|
||||
In `backend/models.py`:
|
||||
|
||||
```python
|
||||
class RequestModel(BaseModel):
|
||||
field: str
|
||||
|
||||
class ResponseModel(BaseModel):
|
||||
result: str
|
||||
```
|
||||
</Step>
|
||||
|
||||
<Step title="Regenerate Client">
|
||||
```bash
|
||||
bun run generate:api
|
||||
```
|
||||
|
||||
This updates the TypeScript client with type-safe bindings.
|
||||
</Step>
|
||||
|
||||
<Step title="Update Docs">
|
||||
Add documentation in `/docs/api/`
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
## Testing
|
||||
|
||||
Currently testing is primarily manual. When adding tests:
|
||||
|
||||
**Backend:**
|
||||
```bash
|
||||
cd backend
|
||||
pytest
|
||||
```
|
||||
|
||||
**Frontend:**
|
||||
```bash
|
||||
bun run test
|
||||
```
|
||||
|
||||
**E2E (future):**
|
||||
```bash
|
||||
bun run test:e2e
|
||||
```
|
||||
|
||||
## Release Process
|
||||
|
||||
Releases are managed by maintainers using `bumpversion`:
|
||||
|
||||
```bash
|
||||
# Bump version (patch, minor, or major)
|
||||
bumpversion patch
|
||||
|
||||
# Push with tags
|
||||
git push && git push --tags
|
||||
```
|
||||
|
||||
GitHub Actions automatically builds and publishes releases when tags are pushed.
|
||||
|
||||
## Community
|
||||
|
||||
- **GitHub Issues:** Bug reports and feature requests
|
||||
- **GitHub Discussions:** General questions and ideas
|
||||
- **Discord:** Real-time chat (coming soon)
|
||||
|
||||
## Recognition
|
||||
|
||||
Contributors are recognized in:
|
||||
- [CHANGELOG.md](https://github.com/jamiepine/voicebox/blob/main/CHANGELOG.md)
|
||||
- GitHub contributor list
|
||||
- Release notes
|
||||
|
||||
## License
|
||||
|
||||
By contributing, you agree that your contributions will be licensed under the MIT License.
|
||||
|
||||
## Questions?
|
||||
|
||||
If you have questions:
|
||||
|
||||
1. Check the [documentation](/overview/introduction)
|
||||
2. Search [existing issues](https://github.com/jamiepine/voicebox/issues)
|
||||
3. Open a new issue or discussion
|
||||
4. See [CONTRIBUTING.md](https://github.com/jamiepine/voicebox/blob/main/CONTRIBUTING.md) in the repo
|
||||
|
||||
Thank you for contributing to Voicebox! 🎉
|
||||
@@ -0,0 +1,260 @@
|
||||
---
|
||||
title: "Generation History"
|
||||
description: "How generation history tracking works in Voicebox"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
The history module tracks all generated audio, providing a searchable record of past generations. Each generation stores the text, settings, and a reference to the audio file.
|
||||
|
||||
## Data Model
|
||||
|
||||
### Generation Table
|
||||
|
||||
```python
|
||||
class Generation(Base):
|
||||
__tablename__ = "generations"
|
||||
|
||||
id = Column(String, primary_key=True)
|
||||
profile_id = Column(String, ForeignKey("profiles.id"))
|
||||
text = Column(Text, nullable=False)
|
||||
language = Column(String, default="en")
|
||||
audio_path = Column(String, nullable=False)
|
||||
duration = Column(Float, nullable=False)
|
||||
seed = Column(Integer)
|
||||
instruct = Column(Text)
|
||||
created_at = Column(DateTime)
|
||||
```
|
||||
|
||||
## File Storage
|
||||
|
||||
Generated audio is stored in:
|
||||
|
||||
```
|
||||
data/
|
||||
└── generations/
|
||||
└── {generation_id}.wav
|
||||
```
|
||||
|
||||
## Core Functions
|
||||
|
||||
### Creating a Generation Record
|
||||
|
||||
After TTS generates audio, a history entry is created:
|
||||
|
||||
```python
|
||||
async def create_generation(
|
||||
profile_id: str,
|
||||
text: str,
|
||||
language: str,
|
||||
audio_path: str,
|
||||
duration: float,
|
||||
seed: Optional[int],
|
||||
db: Session,
|
||||
instruct: Optional[str] = None,
|
||||
) -> GenerationResponse:
|
||||
db_generation = DBGeneration(
|
||||
id=str(uuid.uuid4()),
|
||||
profile_id=profile_id,
|
||||
text=text,
|
||||
language=language,
|
||||
audio_path=audio_path,
|
||||
duration=duration,
|
||||
seed=seed,
|
||||
instruct=instruct,
|
||||
created_at=datetime.utcnow(),
|
||||
)
|
||||
|
||||
db.add(db_generation)
|
||||
db.commit()
|
||||
|
||||
return GenerationResponse.model_validate(db_generation)
|
||||
```
|
||||
|
||||
### Listing Generations
|
||||
|
||||
Supports filtering and pagination:
|
||||
|
||||
```python
|
||||
async def list_generations(
|
||||
query: HistoryQuery,
|
||||
db: Session,
|
||||
) -> HistoryListResponse:
|
||||
# Build query with profile name join
|
||||
q = db.query(
|
||||
DBGeneration,
|
||||
DBVoiceProfile.name.label('profile_name')
|
||||
).join(
|
||||
DBVoiceProfile,
|
||||
DBGeneration.profile_id == DBVoiceProfile.id
|
||||
)
|
||||
|
||||
# Apply filters
|
||||
if query.profile_id:
|
||||
q = q.filter(DBGeneration.profile_id == query.profile_id)
|
||||
|
||||
if query.search:
|
||||
q = q.filter(DBGeneration.text.like(f"%{query.search}%"))
|
||||
|
||||
# Order and paginate
|
||||
total = q.count()
|
||||
q = q.order_by(DBGeneration.created_at.desc())
|
||||
q = q.offset(query.offset).limit(query.limit)
|
||||
|
||||
return HistoryListResponse(items=results, total=total)
|
||||
```
|
||||
|
||||
### Getting Statistics
|
||||
|
||||
Aggregate statistics for the dashboard:
|
||||
|
||||
```python
|
||||
async def get_generation_stats(db: Session) -> dict:
|
||||
total = db.query(func.count(DBGeneration.id)).scalar()
|
||||
total_duration = db.query(func.sum(DBGeneration.duration)).scalar()
|
||||
|
||||
by_profile = db.query(
|
||||
DBGeneration.profile_id,
|
||||
func.count(DBGeneration.id).label('count')
|
||||
).group_by(DBGeneration.profile_id).all()
|
||||
|
||||
return {
|
||||
"total_generations": total,
|
||||
"total_duration_seconds": total_duration,
|
||||
"generations_by_profile": {
|
||||
profile_id: count for profile_id, count in by_profile
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
## Deletion
|
||||
|
||||
Deleting a generation removes both the database record and audio file:
|
||||
|
||||
```python
|
||||
async def delete_generation(generation_id: str, db: Session) -> bool:
|
||||
generation = db.query(DBGeneration).filter_by(id=generation_id).first()
|
||||
if not generation:
|
||||
return False
|
||||
|
||||
# Delete audio file
|
||||
audio_path = Path(generation.audio_path)
|
||||
if audio_path.exists():
|
||||
audio_path.unlink()
|
||||
|
||||
# Delete database record
|
||||
db.delete(generation)
|
||||
db.commit()
|
||||
|
||||
return True
|
||||
```
|
||||
|
||||
### Cascade Delete
|
||||
|
||||
When deleting a profile, all its generations are also deleted:
|
||||
|
||||
```python
|
||||
async def delete_generations_by_profile(profile_id: str, db: Session) -> int:
|
||||
generations = db.query(DBGeneration).filter_by(profile_id=profile_id).all()
|
||||
|
||||
for generation in generations:
|
||||
Path(generation.audio_path).unlink(missing_ok=True)
|
||||
db.delete(generation)
|
||||
|
||||
db.commit()
|
||||
return len(generations)
|
||||
```
|
||||
|
||||
## Export/Import
|
||||
|
||||
### Exporting a Generation
|
||||
|
||||
Generations can be exported as ZIP archives:
|
||||
|
||||
```
|
||||
generation_export.zip
|
||||
├── generation.json # Metadata
|
||||
└── audio.wav # Audio file
|
||||
```
|
||||
|
||||
### Importing a Generation
|
||||
|
||||
The import process:
|
||||
|
||||
1. Extract ZIP archive
|
||||
2. Validate metadata and audio
|
||||
3. Create new generation ID
|
||||
4. Copy audio to generations directory
|
||||
5. Create database record
|
||||
|
||||
## API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| GET | `/history` | List generations with filters |
|
||||
| GET | `/history/stats` | Get aggregate statistics |
|
||||
| GET | `/history/{id}` | Get generation by ID |
|
||||
| DELETE | `/history/{id}` | Delete generation |
|
||||
| GET | `/history/{id}/export` | Export as ZIP |
|
||||
| GET | `/history/{id}/export-audio` | Export audio only |
|
||||
| POST | `/history/import` | Import from ZIP |
|
||||
|
||||
### Query Parameters
|
||||
|
||||
```
|
||||
GET /history?profile_id=uuid&search=hello&limit=50&offset=0
|
||||
```
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `profile_id` | string | null | Filter by profile |
|
||||
| `search` | string | null | Search in text |
|
||||
| `limit` | int | 50 | Results per page |
|
||||
| `offset` | int | 0 | Pagination offset |
|
||||
|
||||
### Response Schema
|
||||
|
||||
```json
|
||||
{
|
||||
"items": [
|
||||
{
|
||||
"id": "uuid",
|
||||
"profile_id": "uuid",
|
||||
"profile_name": "My Voice",
|
||||
"text": "Hello world",
|
||||
"language": "en",
|
||||
"audio_path": "/path/to/audio.wav",
|
||||
"duration": 1.5,
|
||||
"seed": 42,
|
||||
"instruct": null,
|
||||
"created_at": "2024-01-15T10:30:00Z"
|
||||
}
|
||||
],
|
||||
"total": 150
|
||||
}
|
||||
```
|
||||
|
||||
## Usage in Stories
|
||||
|
||||
Generations can be added to stories for multi-voice narratives. The story system references generations by ID:
|
||||
|
||||
```python
|
||||
class StoryItem(Base):
|
||||
generation_id = Column(String, ForeignKey("generations.id"))
|
||||
```
|
||||
|
||||
This allows the same generation to be reused across multiple stories without duplicating audio files.
|
||||
|
||||
## Storage Considerations
|
||||
|
||||
### Disk Usage
|
||||
|
||||
Each generation creates a WAV file. For a 10-second clip at 24kHz:
|
||||
- ~480KB per file (mono, 16-bit)
|
||||
|
||||
### Cleanup Strategy
|
||||
|
||||
Consider implementing:
|
||||
- Automatic cleanup of old generations
|
||||
- Storage quota per profile
|
||||
- Compression for archival
|
||||
@@ -0,0 +1,341 @@
|
||||
---
|
||||
title: "Model Management"
|
||||
description: "How model downloading, loading, and status tracking works in Voicebox"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Voicebox manages two types of models:
|
||||
|
||||
**TTS Models:** Qwen3-TTS for voice cloning (0.6B and 1.7B variants).
|
||||
|
||||
**ASR Models:** Whisper for transcription (tiny through large).
|
||||
|
||||
Models are downloaded from HuggingFace Hub on first use and cached locally.
|
||||
|
||||
## Available Models
|
||||
|
||||
### TTS Models
|
||||
|
||||
| Model | HuggingFace ID | Size | VRAM |
|
||||
|-------|----------------|------|------|
|
||||
| 0.6B | `Qwen/Qwen3-TTS-12Hz-0.6B-Base` | ~1.2GB | ~2GB |
|
||||
| 1.7B | `Qwen/Qwen3-TTS-12Hz-1.7B-Base` | ~3.4GB | ~6GB |
|
||||
|
||||
### Whisper Models
|
||||
|
||||
| Model | HuggingFace ID | Size | VRAM |
|
||||
|-------|----------------|------|------|
|
||||
| tiny | `openai/whisper-tiny` | ~150MB | ~1GB |
|
||||
| base | `openai/whisper-base` | ~300MB | ~1GB |
|
||||
| small | `openai/whisper-small` | ~500MB | ~2GB |
|
||||
| medium | `openai/whisper-medium` | ~1.5GB | ~5GB |
|
||||
| large | `openai/whisper-large` | ~3GB | ~10GB |
|
||||
|
||||
## Model Storage
|
||||
|
||||
Models are cached in the HuggingFace cache directory:
|
||||
|
||||
```
|
||||
~/.cache/huggingface/hub/
|
||||
├── models--Qwen--Qwen3-TTS-12Hz-1.7B-Base/
|
||||
├── models--Qwen--Qwen3-TTS-12Hz-0.6B-Base/
|
||||
├── models--openai--whisper-base/
|
||||
└── ...
|
||||
```
|
||||
|
||||
## Progress Tracking
|
||||
|
||||
### Progress Manager
|
||||
|
||||
Tracks download progress across all models:
|
||||
|
||||
```python
|
||||
class ProgressManager:
|
||||
def __init__(self):
|
||||
self._progress = {} # model_name -> progress_info
|
||||
|
||||
def update_progress(
|
||||
self,
|
||||
model_name: str,
|
||||
current: int,
|
||||
total: int,
|
||||
filename: str,
|
||||
status: str,
|
||||
):
|
||||
self._progress[model_name] = {
|
||||
"current": current,
|
||||
"total": total,
|
||||
"filename": filename,
|
||||
"status": status, # downloading, complete, error
|
||||
"updated_at": datetime.utcnow(),
|
||||
}
|
||||
|
||||
def get_progress(self, model_name: str) -> Optional[dict]:
|
||||
return self._progress.get(model_name)
|
||||
```
|
||||
|
||||
### HuggingFace Progress Callback
|
||||
|
||||
Hooks into HuggingFace's download system:
|
||||
|
||||
```python
|
||||
class HFProgressTracker:
|
||||
def __init__(self, callback):
|
||||
self.callback = callback
|
||||
|
||||
@contextmanager
|
||||
def patch_download(self):
|
||||
"""Context manager to intercept HF downloads."""
|
||||
original_download = hf_hub_download
|
||||
|
||||
def patched_download(*args, **kwargs):
|
||||
# Intercept progress
|
||||
result = original_download(*args, **kwargs)
|
||||
self.callback(progress_info)
|
||||
return result
|
||||
|
||||
# Apply patch
|
||||
with patch('huggingface_hub.hf_hub_download', patched_download):
|
||||
yield
|
||||
```
|
||||
|
||||
### Server-Sent Events (SSE)
|
||||
|
||||
Progress is streamed to the frontend:
|
||||
|
||||
```python
|
||||
@app.get("/models/progress/{model_name}")
|
||||
async def get_model_progress(model_name: str):
|
||||
async def event_generator():
|
||||
while True:
|
||||
progress = progress_manager.get_progress(model_name)
|
||||
if progress:
|
||||
yield f"data: {json.dumps(progress)}\n\n"
|
||||
|
||||
if progress and progress["status"] in ["complete", "error"]:
|
||||
break
|
||||
|
||||
await asyncio.sleep(0.5)
|
||||
|
||||
return StreamingResponse(
|
||||
event_generator(),
|
||||
media_type="text/event-stream"
|
||||
)
|
||||
```
|
||||
|
||||
## Task Manager
|
||||
|
||||
Tracks active downloads and generations:
|
||||
|
||||
```python
|
||||
class TaskManager:
|
||||
def __init__(self):
|
||||
self._active_downloads = {}
|
||||
self._active_generations = {}
|
||||
|
||||
def start_download(self, model_name: str):
|
||||
self._active_downloads[model_name] = {
|
||||
"status": "downloading",
|
||||
"started_at": datetime.utcnow(),
|
||||
}
|
||||
|
||||
def complete_download(self, model_name: str):
|
||||
if model_name in self._active_downloads:
|
||||
del self._active_downloads[model_name]
|
||||
|
||||
def get_active_tasks(self) -> dict:
|
||||
return {
|
||||
"downloads": list(self._active_downloads.values()),
|
||||
"generations": list(self._active_generations.values()),
|
||||
}
|
||||
```
|
||||
|
||||
## Model Status
|
||||
|
||||
Check which models are downloaded and loaded:
|
||||
|
||||
```python
|
||||
@app.get("/models/status")
|
||||
async def get_model_status() -> ModelStatusListResponse:
|
||||
models = []
|
||||
|
||||
# Check TTS models
|
||||
for size, hf_id in [("1.7B", "Qwen/Qwen3-TTS-12Hz-1.7B-Base"), ...]:
|
||||
downloaded = is_model_downloaded(hf_id)
|
||||
loaded = tts_model._current_model_size == size
|
||||
|
||||
models.append(ModelStatus(
|
||||
model_name=f"qwen-tts-{size}",
|
||||
display_name=f"Qwen3-TTS {size}",
|
||||
downloaded=downloaded,
|
||||
size_mb=get_model_size_mb(hf_id),
|
||||
loaded=loaded,
|
||||
))
|
||||
|
||||
# Check Whisper models
|
||||
for size in ["tiny", "base", "small", "medium", "large"]:
|
||||
hf_id = f"openai/whisper-{size}"
|
||||
downloaded = is_model_downloaded(hf_id)
|
||||
|
||||
models.append(ModelStatus(
|
||||
model_name=f"whisper-{size}",
|
||||
display_name=f"Whisper {size}",
|
||||
downloaded=downloaded,
|
||||
size_mb=get_model_size_mb(hf_id),
|
||||
loaded=False, # Whisper is loaded on-demand
|
||||
))
|
||||
|
||||
return ModelStatusListResponse(models=models)
|
||||
```
|
||||
|
||||
## Manual Model Operations
|
||||
|
||||
### Load Model
|
||||
|
||||
```python
|
||||
@app.post("/models/load")
|
||||
async def load_model(model_size: str = "1.7B"):
|
||||
tts_model = get_tts_model()
|
||||
await tts_model.load_model_async(model_size)
|
||||
return {"status": "loaded", "model_size": model_size}
|
||||
```
|
||||
|
||||
### Unload Model
|
||||
|
||||
```python
|
||||
@app.post("/models/unload")
|
||||
async def unload_model():
|
||||
tts_model = get_tts_model()
|
||||
tts_model.unload_model()
|
||||
return {"status": "unloaded"}
|
||||
```
|
||||
|
||||
### Trigger Download
|
||||
|
||||
```python
|
||||
@app.post("/models/download")
|
||||
async def trigger_model_download(request: ModelDownloadRequest):
|
||||
# This triggers the download in background
|
||||
# Progress is tracked via /models/progress/{model_name}
|
||||
|
||||
if request.model_name.startswith("qwen-tts"):
|
||||
size = request.model_name.split("-")[-1]
|
||||
asyncio.create_task(download_tts_model(size))
|
||||
elif request.model_name.startswith("whisper"):
|
||||
size = request.model_name.split("-")[-1]
|
||||
asyncio.create_task(download_whisper_model(size))
|
||||
|
||||
return {"status": "downloading"}
|
||||
```
|
||||
|
||||
### Delete Model
|
||||
|
||||
```python
|
||||
@app.delete("/models/{model_name}")
|
||||
async def delete_model(model_name: str):
|
||||
# Find and delete from HuggingFace cache
|
||||
cache_dir = Path.home() / ".cache" / "huggingface" / "hub"
|
||||
|
||||
model_dirs = list(cache_dir.glob(f"models--*--{model_name}*"))
|
||||
for model_dir in model_dirs:
|
||||
shutil.rmtree(model_dir)
|
||||
|
||||
return {"status": "deleted"}
|
||||
```
|
||||
|
||||
## API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| GET | `/models/status` | Get status of all models |
|
||||
| POST | `/models/load` | Load TTS model |
|
||||
| POST | `/models/unload` | Unload TTS model |
|
||||
| POST | `/models/download` | Trigger model download |
|
||||
| GET | `/models/progress/{name}` | Stream download progress (SSE) |
|
||||
| DELETE | `/models/{name}` | Delete downloaded model |
|
||||
| GET | `/tasks/active` | Get active downloads/generations |
|
||||
|
||||
## Response Schemas
|
||||
|
||||
### ModelStatus
|
||||
|
||||
```json
|
||||
{
|
||||
"model_name": "qwen-tts-1.7B",
|
||||
"display_name": "Qwen3-TTS 1.7B",
|
||||
"downloaded": true,
|
||||
"size_mb": 3400,
|
||||
"loaded": true
|
||||
}
|
||||
```
|
||||
|
||||
### ActiveTasksResponse
|
||||
|
||||
```json
|
||||
{
|
||||
"downloads": [
|
||||
{
|
||||
"model_name": "whisper-medium",
|
||||
"status": "downloading",
|
||||
"started_at": "2024-01-15T10:30:00Z"
|
||||
}
|
||||
],
|
||||
"generations": [
|
||||
{
|
||||
"task_id": "uuid",
|
||||
"profile_id": "uuid",
|
||||
"text_preview": "Hello world...",
|
||||
"started_at": "2024-01-15T10:30:00Z"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Frontend Integration
|
||||
|
||||
### Progress Display
|
||||
|
||||
```typescript
|
||||
// Subscribe to download progress via SSE
|
||||
const eventSource = new EventSource(`/models/progress/${modelName}`);
|
||||
|
||||
eventSource.onmessage = (event) => {
|
||||
const progress = JSON.parse(event.data);
|
||||
updateProgressBar(progress.current / progress.total);
|
||||
|
||||
if (progress.status === 'complete') {
|
||||
eventSource.close();
|
||||
}
|
||||
};
|
||||
```
|
||||
|
||||
### Model Status UI
|
||||
|
||||
```typescript
|
||||
// Fetch model status
|
||||
const { data: models } = useQuery({
|
||||
queryKey: ['models', 'status'],
|
||||
queryFn: () => api.getModelStatus(),
|
||||
});
|
||||
|
||||
// Display download/load buttons based on status
|
||||
models.map(model => (
|
||||
<ModelCard
|
||||
name={model.display_name}
|
||||
downloaded={model.downloaded}
|
||||
loaded={model.loaded}
|
||||
onDownload={() => triggerDownload(model.model_name)}
|
||||
onLoad={() => loadModel(model.model_name)}
|
||||
/>
|
||||
));
|
||||
```
|
||||
|
||||
## Error Handling
|
||||
|
||||
| Error | Cause | Solution |
|
||||
|-------|-------|----------|
|
||||
| Download failed | Network issue | Retry download |
|
||||
| OOM on load | Model too large | Use smaller model |
|
||||
| Model not found | Cache corrupted | Re-download |
|
||||
| Slow download | HF rate limit | Wait and retry |
|
||||
@@ -0,0 +1,242 @@
|
||||
---
|
||||
title: "Development Setup"
|
||||
description: "Set up your local development environment for Voicebox"
|
||||
---
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Before you begin, ensure you have the following installed:
|
||||
|
||||
<CardGroup cols={3}>
|
||||
<Card title="Bun" icon="package">
|
||||
[Download Bun](https://bun.sh)
|
||||
```bash
|
||||
curl -fsSL https://bun.sh/install | bash
|
||||
```
|
||||
</Card>
|
||||
<Card title="Python 3.11+" icon="python">
|
||||
[Download Python](https://python.org)
|
||||
```bash
|
||||
python --version
|
||||
```
|
||||
</Card>
|
||||
<Card title="Rust" icon="rust">
|
||||
[Install Rust](https://rustup.rs)
|
||||
```bash
|
||||
rustc --version
|
||||
```
|
||||
</Card>
|
||||
</CardGroup>
|
||||
|
||||
## Clone the Repository
|
||||
|
||||
```bash
|
||||
git clone https://github.com/jamiepine/voicebox.git
|
||||
cd voicebox
|
||||
```
|
||||
|
||||
## Quick Setup (Recommended)
|
||||
|
||||
The easiest way to get started is using the Makefile:
|
||||
|
||||
```bash
|
||||
# Setup everything
|
||||
make setup
|
||||
|
||||
# Start development
|
||||
make dev
|
||||
```
|
||||
|
||||
<Note>
|
||||
The Makefile is available on macOS and Linux. Windows users should follow the manual setup below.
|
||||
</Note>
|
||||
|
||||
## Manual Setup
|
||||
|
||||
### 1. Install JavaScript Dependencies
|
||||
|
||||
```bash
|
||||
bun install
|
||||
```
|
||||
|
||||
This installs dependencies for:
|
||||
- `app/` - Shared React frontend
|
||||
- `tauri/` - Tauri desktop wrapper
|
||||
- `web/` - Web deployment wrapper
|
||||
|
||||
### 2. Set Up Python Backend
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
|
||||
# Create virtual environment
|
||||
python -m venv venv
|
||||
|
||||
# Activate virtual environment
|
||||
source venv/bin/activate # macOS/Linux
|
||||
# or
|
||||
venv\Scripts\activate # Windows
|
||||
|
||||
# Install Python dependencies
|
||||
pip install -r requirements.txt
|
||||
|
||||
# Install Qwen3-TTS
|
||||
pip install git+https://github.com/QwenLM/Qwen3-TTS.git
|
||||
```
|
||||
|
||||
### 3. Initialize Database
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
python -c "from database import init_db; init_db()"
|
||||
```
|
||||
|
||||
This creates the SQLite database at `data/voicebox.db`.
|
||||
|
||||
## Running in Development
|
||||
|
||||
Development requires **two terminals**: one for the Python backend, one for the Tauri app.
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Terminal 1: Backend">
|
||||
Start the Python server first:
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
source venv/bin/activate # Activate venv
|
||||
bun run dev:server
|
||||
```
|
||||
|
||||
Or manually:
|
||||
```bash
|
||||
uvicorn main:app --reload --port 17493
|
||||
```
|
||||
|
||||
Backend will be available at `http://localhost:17493`
|
||||
</Tab>
|
||||
|
||||
<Tab title="Terminal 2: Desktop App">
|
||||
Then start the Tauri app:
|
||||
|
||||
```bash
|
||||
bun run dev
|
||||
```
|
||||
|
||||
This will:
|
||||
- Create a placeholder sidecar binary
|
||||
- Start Vite dev server on port 5173
|
||||
- Launch Tauri window
|
||||
- Enable hot reload
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
<Info>
|
||||
In dev mode, the app connects to your manually-started Python server. The bundled server binary is only used in production builds.
|
||||
</Info>
|
||||
|
||||
### Optional: Web App
|
||||
|
||||
```bash
|
||||
bun run dev:web
|
||||
```
|
||||
|
||||
Web app will be available at `http://localhost:5174`
|
||||
|
||||
## Model Downloads
|
||||
|
||||
Models are automatically downloaded from HuggingFace Hub on first use:
|
||||
|
||||
- **Whisper** (transcription): Auto-downloads on first transcription
|
||||
- **Qwen3-TTS** (voice cloning): Auto-downloads on first generation (~2-4GB)
|
||||
|
||||
<Warning>
|
||||
First-time usage will be slower due to model downloads, but subsequent runs will use cached models.
|
||||
</Warning>
|
||||
|
||||
## Project Structure
|
||||
|
||||
```
|
||||
voicebox/
|
||||
├── app/ # Shared React frontend
|
||||
│ └── src/
|
||||
│ ├── components/ # UI components
|
||||
│ ├── lib/ # Utilities and API client
|
||||
│ └── hooks/ # React hooks
|
||||
├── backend/ # Python FastAPI server
|
||||
│ ├── main.py # API routes
|
||||
│ ├── tts.py # Voice synthesis
|
||||
│ └── database.py # SQLite operations
|
||||
├── tauri/ # Desktop app wrapper
|
||||
│ └── src-tauri/ # Rust backend
|
||||
├── web/ # Web deployment
|
||||
├── landing/ # Marketing website
|
||||
└── scripts/ # Build & release scripts
|
||||
```
|
||||
|
||||
## Available Make Commands
|
||||
|
||||
Run `make help` to see all available commands:
|
||||
|
||||
```bash
|
||||
make setup # Install all dependencies
|
||||
make dev # Start development servers
|
||||
make dev-web # Start web development server
|
||||
make build # Build desktop app
|
||||
make build-web # Build web app
|
||||
make clean # Clean build artifacts
|
||||
make test # Run tests
|
||||
```
|
||||
|
||||
## Generate OpenAPI Client
|
||||
|
||||
After starting the backend server, generate the TypeScript API client:
|
||||
|
||||
```bash
|
||||
./scripts/generate-api.sh
|
||||
# or
|
||||
bun run generate:api
|
||||
```
|
||||
|
||||
This downloads the OpenAPI schema and generates the TypeScript client in `app/src/lib/api/`
|
||||
|
||||
## Next Steps
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Architecture" icon="diagram-project" href="/development/architecture">
|
||||
Understand the system architecture
|
||||
</Card>
|
||||
<Card title="Contributing" icon="code-pull-request" href="/development/contributing">
|
||||
Read the contribution guidelines
|
||||
</Card>
|
||||
<Card title="Building" icon="hammer" href="/development/building">
|
||||
Learn how to build production releases
|
||||
</Card>
|
||||
<Card title="API Reference" icon="code" href="/api/overview">
|
||||
Explore the REST API
|
||||
</Card>
|
||||
</CardGroup>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
<AccordionGroup>
|
||||
<Accordion title="Backend won't start">
|
||||
- Check Python version (must be 3.11+)
|
||||
- Ensure virtual environment is activated
|
||||
- Verify all dependencies are installed: `pip install -r requirements.txt`
|
||||
- Check if port 17493 is available
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Tauri build fails">
|
||||
- Ensure Rust is installed: `rustc --version`
|
||||
- Clean the build: `cd tauri/src-tauri && cargo clean`
|
||||
- Try rebuilding: `bun run dev`
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="OpenAPI client generation fails">
|
||||
- Ensure backend is running: `curl http://localhost:17493/openapi.json`
|
||||
- Check network connectivity
|
||||
- Verify the backend is accessible at localhost:17493
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
See the full [Troubleshooting Guide](/guides/troubleshooting) for more issues and solutions.
|
||||
@@ -0,0 +1,320 @@
|
||||
---
|
||||
title: "Stories & Timeline"
|
||||
description: "How the multi-voice timeline editor works in Voicebox"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Stories allow users to arrange multiple voice generations on a timeline to create multi-voice narratives. The system supports tracks, trimming, splitting, and audio mixing.
|
||||
|
||||
## Architecture
|
||||
|
||||
**Story:** A container that holds story items with metadata.
|
||||
|
||||
**Story Item:** Links a generation to a story with timeline position, track, and trim data.
|
||||
|
||||
**Export:** Combines all items into a single mixed audio file.
|
||||
|
||||
## Data Model
|
||||
|
||||
### Story Table
|
||||
|
||||
```python
|
||||
class Story(Base):
|
||||
__tablename__ = "stories"
|
||||
|
||||
id = Column(String, primary_key=True)
|
||||
name = Column(String, nullable=False)
|
||||
description = Column(Text)
|
||||
created_at = Column(DateTime)
|
||||
updated_at = Column(DateTime)
|
||||
```
|
||||
|
||||
### StoryItem Table
|
||||
|
||||
```python
|
||||
class StoryItem(Base):
|
||||
__tablename__ = "story_items"
|
||||
|
||||
id = Column(String, primary_key=True)
|
||||
story_id = Column(String, ForeignKey("stories.id"))
|
||||
generation_id = Column(String, ForeignKey("generations.id"))
|
||||
start_time_ms = Column(Integer, default=0) # Timeline position
|
||||
track = Column(Integer, default=0) # Track number
|
||||
trim_start_ms = Column(Integer, default=0) # Trim from start
|
||||
trim_end_ms = Column(Integer, default=0) # Trim from end
|
||||
created_at = Column(DateTime)
|
||||
```
|
||||
|
||||
## Timeline Concepts
|
||||
|
||||
### Start Time
|
||||
|
||||
`start_time_ms` defines when an item begins on the timeline:
|
||||
|
||||
```
|
||||
Timeline (ms): 0----1000----2000----3000----4000
|
||||
Item 1: [======]
|
||||
Item 2: [==========]
|
||||
Item 3: [====]
|
||||
```
|
||||
|
||||
### Tracks
|
||||
|
||||
Multiple tracks allow overlapping audio:
|
||||
|
||||
```
|
||||
Track 0: [Item 1] [Item 3]
|
||||
Track 1: [Item 2]
|
||||
```
|
||||
|
||||
### Trimming
|
||||
|
||||
Trim values cut audio from the start or end without destroying the original:
|
||||
|
||||
```
|
||||
Original: [=========AUDIO=========]
|
||||
trim_start: ^^
|
||||
trim_end: ^^
|
||||
Result: [=====AUDIO=====]
|
||||
```
|
||||
|
||||
## Core Operations
|
||||
|
||||
### Adding Items
|
||||
|
||||
When adding a generation to a story:
|
||||
|
||||
```python
|
||||
async def add_item_to_story(
|
||||
story_id: str,
|
||||
data: StoryItemCreate,
|
||||
db: Session,
|
||||
) -> StoryItemDetail:
|
||||
# Calculate start time if not provided
|
||||
if data.start_time_ms is None:
|
||||
# Find the end of all existing items
|
||||
existing_items = get_items_with_durations(story_id, db)
|
||||
max_end_time_ms = max(
|
||||
item.start_time_ms + int(gen.duration * 1000)
|
||||
for item, gen in existing_items
|
||||
)
|
||||
start_time_ms = max_end_time_ms + 200 # 200ms gap
|
||||
|
||||
# Create the item
|
||||
item = DBStoryItem(
|
||||
id=str(uuid.uuid4()),
|
||||
story_id=story_id,
|
||||
generation_id=data.generation_id,
|
||||
start_time_ms=start_time_ms,
|
||||
track=data.track or 0,
|
||||
)
|
||||
db.add(item)
|
||||
db.commit()
|
||||
```
|
||||
|
||||
### Moving Items
|
||||
|
||||
Update position and/or track:
|
||||
|
||||
```python
|
||||
async def move_story_item(
|
||||
story_id: str,
|
||||
item_id: str,
|
||||
data: StoryItemMove,
|
||||
db: Session,
|
||||
) -> StoryItemDetail:
|
||||
item = get_item(story_id, item_id, db)
|
||||
|
||||
item.start_time_ms = data.start_time_ms
|
||||
item.track = data.track
|
||||
|
||||
db.commit()
|
||||
```
|
||||
|
||||
### Trimming Items
|
||||
|
||||
Non-destructive trimming:
|
||||
|
||||
```python
|
||||
async def trim_story_item(
|
||||
story_id: str,
|
||||
item_id: str,
|
||||
data: StoryItemTrim,
|
||||
db: Session,
|
||||
) -> StoryItemDetail:
|
||||
item = get_item(story_id, item_id, db)
|
||||
generation = get_generation(item.generation_id, db)
|
||||
|
||||
# Validate trim doesn't exceed duration
|
||||
max_duration_ms = int(generation.duration * 1000)
|
||||
if data.trim_start_ms + data.trim_end_ms >= max_duration_ms:
|
||||
return None # Invalid trim
|
||||
|
||||
item.trim_start_ms = data.trim_start_ms
|
||||
item.trim_end_ms = data.trim_end_ms
|
||||
|
||||
db.commit()
|
||||
```
|
||||
|
||||
### Splitting Items
|
||||
|
||||
Split one item into two at a specific time:
|
||||
|
||||
```python
|
||||
async def split_story_item(
|
||||
story_id: str,
|
||||
item_id: str,
|
||||
data: StoryItemSplit,
|
||||
db: Session,
|
||||
) -> List[StoryItemDetail]:
|
||||
item = get_item(story_id, item_id, db)
|
||||
generation = get_generation(item.generation_id, db)
|
||||
|
||||
# Calculate split point
|
||||
current_trim_start = item.trim_start_ms
|
||||
current_trim_end = item.trim_end_ms
|
||||
original_duration_ms = int(generation.duration * 1000)
|
||||
absolute_split_ms = current_trim_start + data.split_time_ms
|
||||
|
||||
# Update original: trim from end
|
||||
item.trim_end_ms = original_duration_ms - absolute_split_ms
|
||||
|
||||
# Create new item: trim from start
|
||||
new_item = DBStoryItem(
|
||||
generation_id=item.generation_id, # Same generation
|
||||
start_time_ms=item.start_time_ms + data.split_time_ms,
|
||||
track=item.track,
|
||||
trim_start_ms=absolute_split_ms,
|
||||
trim_end_ms=current_trim_end,
|
||||
)
|
||||
|
||||
db.add(new_item)
|
||||
db.commit()
|
||||
|
||||
return [item, new_item]
|
||||
```
|
||||
|
||||
### Duplicating Items
|
||||
|
||||
Create a copy with all properties:
|
||||
|
||||
```python
|
||||
async def duplicate_story_item(
|
||||
story_id: str,
|
||||
item_id: str,
|
||||
db: Session,
|
||||
) -> StoryItemDetail:
|
||||
original = get_item(story_id, item_id, db)
|
||||
generation = get_generation(original.generation_id, db)
|
||||
|
||||
# Calculate effective duration for positioning
|
||||
effective_duration_ms = (
|
||||
int(generation.duration * 1000)
|
||||
- original.trim_start_ms
|
||||
- original.trim_end_ms
|
||||
)
|
||||
|
||||
# Place copy after original with 200ms gap
|
||||
new_item = DBStoryItem(
|
||||
generation_id=original.generation_id,
|
||||
start_time_ms=original.start_time_ms + effective_duration_ms + 200,
|
||||
track=original.track,
|
||||
trim_start_ms=original.trim_start_ms,
|
||||
trim_end_ms=original.trim_end_ms,
|
||||
)
|
||||
|
||||
db.add(new_item)
|
||||
db.commit()
|
||||
```
|
||||
|
||||
## Audio Export
|
||||
|
||||
### Mixing Algorithm
|
||||
|
||||
The export function mixes all items into a single audio file:
|
||||
|
||||
```python
|
||||
async def export_story_audio(story_id: str, db: Session) -> bytes:
|
||||
items = get_all_items_with_generations(story_id, db)
|
||||
|
||||
# Calculate total duration
|
||||
max_end_time_ms = max(
|
||||
data['start_time_ms'] + data['duration_ms']
|
||||
for data in audio_data
|
||||
)
|
||||
|
||||
# Create output buffer
|
||||
total_samples = int((max_end_time_ms / 1000.0) * sample_rate)
|
||||
final_audio = np.zeros(total_samples, dtype=np.float32)
|
||||
|
||||
# Mix each item at its position
|
||||
for data in audio_data:
|
||||
audio = data['audio']
|
||||
start_sample = int((data['start_time_ms'] / 1000.0) * sample_rate)
|
||||
|
||||
# Apply trim
|
||||
trimmed_audio = audio[trim_start_sample:len(audio) - trim_end_sample]
|
||||
|
||||
# Add to buffer (overlapping items sum together)
|
||||
final_audio[start_sample:start_sample + len(trimmed_audio)] += trimmed_audio
|
||||
|
||||
# Normalize to prevent clipping
|
||||
max_val = np.abs(final_audio).max()
|
||||
if max_val > 1.0:
|
||||
final_audio = final_audio / max_val
|
||||
|
||||
return audio_to_bytes(final_audio, sample_rate)
|
||||
```
|
||||
|
||||
## API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| GET | `/stories` | List all stories |
|
||||
| POST | `/stories` | Create a story |
|
||||
| GET | `/stories/{id}` | Get story with items |
|
||||
| PUT | `/stories/{id}` | Update story metadata |
|
||||
| DELETE | `/stories/{id}` | Delete story |
|
||||
| POST | `/stories/{id}/items` | Add item to story |
|
||||
| DELETE | `/stories/{id}/items/{item_id}` | Remove item |
|
||||
| PUT | `/stories/{id}/items/{item_id}/move` | Move item |
|
||||
| PUT | `/stories/{id}/items/{item_id}/trim` | Trim item |
|
||||
| POST | `/stories/{id}/items/{item_id}/split` | Split item |
|
||||
| POST | `/stories/{id}/items/{item_id}/duplicate` | Duplicate item |
|
||||
| PUT | `/stories/{id}/items/times` | Batch update times |
|
||||
| PUT | `/stories/{id}/items/reorder` | Reorder items |
|
||||
| GET | `/stories/{id}/export-audio` | Export mixed audio |
|
||||
|
||||
## Response Schemas
|
||||
|
||||
### StoryItemDetail
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "item_uuid",
|
||||
"story_id": "story_uuid",
|
||||
"generation_id": "generation_uuid",
|
||||
"start_time_ms": 1500,
|
||||
"track": 0,
|
||||
"trim_start_ms": 200,
|
||||
"trim_end_ms": 100,
|
||||
"profile_id": "profile_uuid",
|
||||
"profile_name": "Narrator",
|
||||
"text": "Hello world",
|
||||
"audio_path": "/path/to/audio.wav",
|
||||
"duration": 2.5,
|
||||
"created_at": "2024-01-15T10:30:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
## Frontend Integration
|
||||
|
||||
The timeline UI needs to:
|
||||
|
||||
1. **Fetch story** with all items
|
||||
2. **Render waveforms** for each item
|
||||
3. **Handle drag/drop** to move items
|
||||
4. **Handle edge drag** for trimming
|
||||
5. **Sync playhead** across all tracks
|
||||
6. **Export** when user clicks download
|
||||
@@ -0,0 +1,299 @@
|
||||
---
|
||||
title: "Transcription"
|
||||
description: "How Whisper-based audio transcription works in Voicebox"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Voicebox uses OpenAI's Whisper model for automatic speech recognition (ASR). This powers the transcription feature for creating reference text from audio recordings.
|
||||
|
||||
## Architecture
|
||||
|
||||
The transcription system is built around the `WhisperModel` class:
|
||||
|
||||
**Model Loading:** Lazy loading with HuggingFace Hub download.
|
||||
|
||||
**Audio Processing:** Resampling and preprocessing for Whisper.
|
||||
|
||||
**Inference:** Running transcription with optional language hints.
|
||||
|
||||
## WhisperModel Class
|
||||
|
||||
```python
|
||||
class WhisperModel:
|
||||
def __init__(self, model_size: str = "base"):
|
||||
self.model = None
|
||||
self.processor = None
|
||||
self.model_size = model_size
|
||||
self.device = self._get_device()
|
||||
```
|
||||
|
||||
### Model Sizes
|
||||
|
||||
| Size | Parameters | VRAM | Speed | Quality |
|
||||
|------|------------|------|-------|---------|
|
||||
| tiny | 39M | ~1GB | Fastest | Basic |
|
||||
| base | 74M | ~1GB | Fast | Good |
|
||||
| small | 244M | ~2GB | Medium | Better |
|
||||
| medium | 769M | ~5GB | Slow | High |
|
||||
| large | 1550M | ~10GB | Slowest | Best |
|
||||
|
||||
Default is `base` for balance of speed and quality.
|
||||
|
||||
## Model Loading
|
||||
|
||||
Models are downloaded from HuggingFace Hub:
|
||||
|
||||
```python
|
||||
def load_model(self, model_size: Optional[str] = None):
|
||||
from transformers import WhisperProcessor, WhisperForConditionalGeneration
|
||||
|
||||
model_name = f"openai/whisper-{model_size}"
|
||||
|
||||
# Track download progress
|
||||
progress_manager = get_progress_manager()
|
||||
task_manager = get_task_manager()
|
||||
task_manager.start_download(f"whisper-{model_size}")
|
||||
|
||||
# Load processor and model
|
||||
with tracker.patch_download():
|
||||
self.processor = WhisperProcessor.from_pretrained(model_name)
|
||||
self.model = WhisperForConditionalGeneration.from_pretrained(model_name)
|
||||
|
||||
self.model.to(self.device)
|
||||
|
||||
# Mark complete
|
||||
progress_manager.mark_complete(f"whisper-{model_size}")
|
||||
task_manager.complete_download(f"whisper-{model_size}")
|
||||
```
|
||||
|
||||
### Async Loading
|
||||
|
||||
Like TTS, loading runs in a thread pool:
|
||||
|
||||
```python
|
||||
async def load_model_async(self, model_size: Optional[str] = None):
|
||||
if self.model is not None and self.model_size == model_size:
|
||||
return
|
||||
await asyncio.to_thread(self.load_model, model_size)
|
||||
```
|
||||
|
||||
## Transcription
|
||||
|
||||
### Basic Transcription
|
||||
|
||||
```python
|
||||
async def transcribe(
|
||||
self,
|
||||
audio_path: str,
|
||||
language: Optional[str] = None,
|
||||
) -> str:
|
||||
await self.load_model_async()
|
||||
|
||||
def _transcribe_sync():
|
||||
# Load and resample to 16kHz (Whisper requirement)
|
||||
audio, sr = load_audio(audio_path, sample_rate=16000)
|
||||
|
||||
# Process audio
|
||||
inputs = self.processor(
|
||||
audio,
|
||||
sampling_rate=16000,
|
||||
return_tensors="pt",
|
||||
)
|
||||
inputs = inputs.to(self.device)
|
||||
|
||||
# Set language hint if provided
|
||||
forced_decoder_ids = None
|
||||
if language:
|
||||
forced_decoder_ids = self.processor.get_decoder_prompt_ids(
|
||||
language=language,
|
||||
task="transcribe",
|
||||
)
|
||||
|
||||
# Generate
|
||||
with torch.no_grad():
|
||||
predicted_ids = self.model.generate(
|
||||
inputs["input_features"],
|
||||
forced_decoder_ids=forced_decoder_ids,
|
||||
)
|
||||
|
||||
# Decode
|
||||
transcription = self.processor.batch_decode(
|
||||
predicted_ids,
|
||||
skip_special_tokens=True,
|
||||
)[0]
|
||||
|
||||
return transcription.strip()
|
||||
|
||||
return await asyncio.to_thread(_transcribe_sync)
|
||||
```
|
||||
|
||||
### Supported Languages
|
||||
|
||||
Whisper supports 99+ languages. Common ones in Voicebox:
|
||||
|
||||
| Code | Language |
|
||||
|------|----------|
|
||||
| en | English |
|
||||
| zh | Chinese |
|
||||
| ja | Japanese |
|
||||
| ko | Korean |
|
||||
| de | German |
|
||||
| fr | French |
|
||||
| ru | Russian |
|
||||
| pt | Portuguese |
|
||||
| es | Spanish |
|
||||
| it | Italian |
|
||||
|
||||
### Language Detection
|
||||
|
||||
When no language is specified, Whisper auto-detects:
|
||||
|
||||
```python
|
||||
# Without language hint - auto-detect
|
||||
transcription = await whisper.transcribe(audio_path)
|
||||
|
||||
# With language hint - more accurate for short clips
|
||||
transcription = await whisper.transcribe(audio_path, language="en")
|
||||
```
|
||||
|
||||
## Transcription with Timestamps
|
||||
|
||||
For advanced use cases, word-level timestamps are available:
|
||||
|
||||
```python
|
||||
async def transcribe_with_timestamps(
|
||||
self,
|
||||
audio_path: str,
|
||||
language: Optional[str] = None,
|
||||
) -> List[Dict[str, any]]:
|
||||
await self.load_model_async()
|
||||
|
||||
def _transcribe_timestamps_sync():
|
||||
audio, sr = load_audio(audio_path, sample_rate=16000)
|
||||
inputs = self.processor(audio, sampling_rate=16000, return_tensors="pt")
|
||||
|
||||
with torch.no_grad():
|
||||
predicted_ids = self.model.generate(
|
||||
inputs["input_features"],
|
||||
return_timestamps=True,
|
||||
)
|
||||
|
||||
# Parse timestamps
|
||||
return [
|
||||
{
|
||||
"text": transcription,
|
||||
"start": 0.0,
|
||||
"end": len(audio) / sr,
|
||||
}
|
||||
]
|
||||
|
||||
return await asyncio.to_thread(_transcribe_timestamps_sync)
|
||||
```
|
||||
|
||||
## Memory Management
|
||||
|
||||
### Unloading
|
||||
|
||||
Free memory when not needed:
|
||||
|
||||
```python
|
||||
def unload_model(self):
|
||||
if self.model is not None:
|
||||
del self.model
|
||||
del self.processor
|
||||
self.model = None
|
||||
self.processor = None
|
||||
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.empty_cache()
|
||||
```
|
||||
|
||||
### Global Instance
|
||||
|
||||
A singleton pattern manages the model:
|
||||
|
||||
```python
|
||||
_whisper_model: Optional[WhisperModel] = None
|
||||
|
||||
def get_whisper_model() -> WhisperModel:
|
||||
global _whisper_model
|
||||
if _whisper_model is None:
|
||||
_whisper_model = WhisperModel()
|
||||
return _whisper_model
|
||||
```
|
||||
|
||||
## Audio Preprocessing
|
||||
|
||||
### Resampling
|
||||
|
||||
Whisper requires 16kHz audio:
|
||||
|
||||
```python
|
||||
audio, sr = load_audio(audio_path, sample_rate=16000)
|
||||
```
|
||||
|
||||
### Format Support
|
||||
|
||||
The `load_audio` utility handles:
|
||||
- WAV
|
||||
- MP3
|
||||
- FLAC
|
||||
- OGG
|
||||
- M4A
|
||||
|
||||
All formats are converted to mono 16kHz.
|
||||
|
||||
## API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| POST | `/transcribe` | Transcribe audio file |
|
||||
|
||||
### Request
|
||||
|
||||
Multipart form data:
|
||||
|
||||
```
|
||||
POST /transcribe
|
||||
Content-Type: multipart/form-data
|
||||
|
||||
file: <audio_file>
|
||||
language: en (optional)
|
||||
```
|
||||
|
||||
### Response
|
||||
|
||||
```json
|
||||
{
|
||||
"text": "Hello, this is a test transcription.",
|
||||
"duration": 3.5
|
||||
}
|
||||
```
|
||||
|
||||
## Use Cases
|
||||
|
||||
### Reference Text for Voice Cloning
|
||||
|
||||
1. User records audio sample
|
||||
2. Audio is sent to `/transcribe`
|
||||
3. Transcription becomes `reference_text`
|
||||
4. Both are added to voice profile
|
||||
|
||||
### Quality Tips
|
||||
|
||||
- Provide language hint for short audio
|
||||
- Use clean audio with minimal noise
|
||||
- Longer audio (>5s) improves accuracy
|
||||
- Consider `small` or `medium` model for better quality
|
||||
|
||||
## Error Handling
|
||||
|
||||
Common issues:
|
||||
|
||||
| Error | Cause | Solution |
|
||||
|-------|-------|----------|
|
||||
| Model not found | First run, download failed | Retry with network |
|
||||
| OOM | Model too large | Use smaller model |
|
||||
| Empty result | No speech detected | Check audio has speech |
|
||||
| Wrong language | Auto-detect failed | Provide language hint |
|
||||
@@ -0,0 +1,283 @@
|
||||
---
|
||||
title: "TTS Generation"
|
||||
description: "How text-to-speech generation works in Voicebox"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Voicebox uses Qwen3-TTS for voice cloning and text-to-speech generation. The TTS module handles model loading, voice prompt creation, and audio synthesis.
|
||||
|
||||
## Architecture
|
||||
|
||||
The TTS system is built around the `TTSModel` class which manages:
|
||||
|
||||
**Model Loading:** Lazy loading with automatic HuggingFace Hub download.
|
||||
|
||||
**Voice Prompts:** Converting reference audio into embeddings.
|
||||
|
||||
**Generation:** Synthesizing speech from text using voice prompts.
|
||||
|
||||
## TTSModel Class
|
||||
|
||||
```python
|
||||
class TTSModel:
|
||||
def __init__(self, model_size: str = "1.7B"):
|
||||
self.model = None
|
||||
self.model_size = model_size
|
||||
self.device = self._get_device() # cuda, mps, or cpu
|
||||
```
|
||||
|
||||
### Device Selection
|
||||
|
||||
The model automatically selects the best available device:
|
||||
|
||||
```python
|
||||
def _get_device(self) -> str:
|
||||
if torch.cuda.is_available():
|
||||
return "cuda"
|
||||
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
||||
return "cpu" # MPS can have issues, use CPU for stability
|
||||
return "cpu"
|
||||
```
|
||||
|
||||
## Model Loading
|
||||
|
||||
Models are downloaded from HuggingFace Hub on first use:
|
||||
|
||||
```python
|
||||
def load_model(self, model_size: Optional[str] = None):
|
||||
# Model IDs on HuggingFace Hub
|
||||
hf_model_map = {
|
||||
"1.7B": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
|
||||
"0.6B": "Qwen/Qwen3-TTS-12Hz-0.6B-Base",
|
||||
}
|
||||
|
||||
# Load with progress tracking
|
||||
with tracker.patch_download():
|
||||
self.model = Qwen3TTSModel.from_pretrained(
|
||||
model_path,
|
||||
device_map=self.device,
|
||||
torch_dtype=torch.bfloat16, # float32 on CPU
|
||||
)
|
||||
```
|
||||
|
||||
### Async Loading
|
||||
|
||||
Loading runs in a thread pool to avoid blocking the event loop:
|
||||
|
||||
```python
|
||||
async def load_model_async(self, model_size: Optional[str] = None):
|
||||
if self.model is not None and self._current_model_size == model_size:
|
||||
return
|
||||
await asyncio.to_thread(self.load_model, model_size)
|
||||
```
|
||||
|
||||
## Voice Prompt Creation
|
||||
|
||||
Voice prompts are created from reference audio and cached for reuse:
|
||||
|
||||
```python
|
||||
async def create_voice_prompt(
|
||||
self,
|
||||
audio_path: str,
|
||||
reference_text: str,
|
||||
use_cache: bool = True,
|
||||
) -> Tuple[dict, bool]:
|
||||
await self.load_model_async()
|
||||
|
||||
# Check cache
|
||||
if use_cache:
|
||||
cache_key = get_cache_key(audio_path, reference_text)
|
||||
cached = get_cached_voice_prompt(cache_key)
|
||||
if cached:
|
||||
return cached, True
|
||||
|
||||
# Create prompt (blocking, run in thread pool)
|
||||
voice_prompt = await asyncio.to_thread(
|
||||
self.model.create_voice_clone_prompt,
|
||||
ref_audio=audio_path,
|
||||
ref_text=reference_text,
|
||||
)
|
||||
|
||||
# Cache the result
|
||||
cache_voice_prompt(cache_key, voice_prompt)
|
||||
return voice_prompt, False
|
||||
```
|
||||
|
||||
### Combining Multiple Samples
|
||||
|
||||
When a profile has multiple samples, they're combined:
|
||||
|
||||
```python
|
||||
async def combine_voice_prompts(
|
||||
self,
|
||||
audio_paths: List[str],
|
||||
reference_texts: List[str],
|
||||
) -> Tuple[np.ndarray, str]:
|
||||
combined_audio = []
|
||||
|
||||
for audio_path in audio_paths:
|
||||
audio, sr = load_audio(audio_path)
|
||||
audio = normalize_audio(audio)
|
||||
combined_audio.append(audio)
|
||||
|
||||
# Concatenate and normalize
|
||||
mixed = np.concatenate(combined_audio)
|
||||
mixed = normalize_audio(mixed)
|
||||
|
||||
# Combine texts
|
||||
combined_text = " ".join(reference_texts)
|
||||
|
||||
return mixed, combined_text
|
||||
```
|
||||
|
||||
## Speech Generation
|
||||
|
||||
The core generation function:
|
||||
|
||||
```python
|
||||
async def generate(
|
||||
self,
|
||||
text: str,
|
||||
voice_prompt: dict,
|
||||
language: str = "en",
|
||||
seed: Optional[int] = None,
|
||||
instruct: Optional[str] = None,
|
||||
) -> Tuple[np.ndarray, int]:
|
||||
await self.load_model_async()
|
||||
|
||||
def _generate_sync():
|
||||
# Set seed for reproducibility
|
||||
if seed is not None:
|
||||
torch.manual_seed(seed)
|
||||
|
||||
# Generate audio
|
||||
wavs, sample_rate = self.model.generate_voice_clone(
|
||||
text=text,
|
||||
voice_clone_prompt=voice_prompt,
|
||||
instruct=instruct, # Natural language delivery control
|
||||
)
|
||||
return wavs[0], sample_rate
|
||||
|
||||
# Run in thread pool
|
||||
return await asyncio.to_thread(_generate_sync)
|
||||
```
|
||||
|
||||
### Instruct Feature
|
||||
|
||||
The `instruct` parameter allows natural language control over speech delivery:
|
||||
|
||||
```python
|
||||
# Examples:
|
||||
instruct = "Speak slowly and clearly"
|
||||
instruct = "Sound excited and enthusiastic"
|
||||
instruct = "Whisper softly"
|
||||
```
|
||||
|
||||
## Caching Strategy
|
||||
|
||||
Voice prompts are cached to avoid recomputation:
|
||||
|
||||
```python
|
||||
def get_cache_key(audio_path: str, reference_text: str) -> str:
|
||||
"""Generate cache key from audio hash and text."""
|
||||
audio_hash = hashlib.md5(Path(audio_path).read_bytes()).hexdigest()
|
||||
text_hash = hashlib.md5(reference_text.encode()).hexdigest()
|
||||
return f"{audio_hash}_{text_hash}"
|
||||
```
|
||||
|
||||
Cache is stored in `data/cache/voice_prompts/`.
|
||||
|
||||
## Memory Management
|
||||
|
||||
### Unloading Models
|
||||
|
||||
Free VRAM/RAM when not needed:
|
||||
|
||||
```python
|
||||
def unload_model(self):
|
||||
if self.model is not None:
|
||||
del self.model
|
||||
self.model = None
|
||||
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.empty_cache()
|
||||
```
|
||||
|
||||
### Model Switching
|
||||
|
||||
When switching between model sizes (1.7B ↔ 0.6B):
|
||||
|
||||
```python
|
||||
# Unload existing model first
|
||||
if self.model is not None and self._current_model_size != model_size:
|
||||
self.unload_model()
|
||||
```
|
||||
|
||||
## Generation Flow
|
||||
|
||||
1. **Request** → Validate text and profile ID
|
||||
2. **Profile** → Load profile samples from database
|
||||
3. **Voice Prompt** → Create or retrieve cached prompt
|
||||
4. **Generate** → Run TTS inference
|
||||
5. **Save** → Write audio to generations directory
|
||||
6. **Record** → Create history entry in database
|
||||
7. **Response** → Return audio path and metadata
|
||||
|
||||
## API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| POST | `/generate` | Generate speech from text |
|
||||
| GET | `/audio/{id}` | Serve generated audio file |
|
||||
|
||||
### Request Schema
|
||||
|
||||
```json
|
||||
{
|
||||
"profile_id": "uuid",
|
||||
"text": "Text to synthesize",
|
||||
"language": "en",
|
||||
"seed": 42,
|
||||
"model_size": "1.7B",
|
||||
"instruct": "Speak clearly"
|
||||
}
|
||||
```
|
||||
|
||||
### Response Schema
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "generation_uuid",
|
||||
"profile_id": "profile_uuid",
|
||||
"text": "Text to synthesize",
|
||||
"language": "en",
|
||||
"audio_path": "/path/to/audio.wav",
|
||||
"duration": 3.5,
|
||||
"seed": 42,
|
||||
"instruct": "Speak clearly",
|
||||
"created_at": "2024-01-15T10:30:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
### GPU Acceleration
|
||||
|
||||
- CUDA provides fastest inference
|
||||
- MPS (Apple Silicon) has stability issues, uses CPU fallback
|
||||
- CPU inference is slower but always works
|
||||
|
||||
### Batch Size
|
||||
|
||||
Currently generates one utterance at a time. For long texts, consider:
|
||||
- Splitting into sentences
|
||||
- Sequential generation
|
||||
- Concatenating results
|
||||
|
||||
### Memory Usage
|
||||
|
||||
| Model | VRAM/RAM Required |
|
||||
|-------|-------------------|
|
||||
| 0.6B | ~2GB |
|
||||
| 1.7B | ~6GB |
|
||||
@@ -0,0 +1,202 @@
|
||||
---
|
||||
title: "Voice Profiles"
|
||||
description: "How voice profile management works in Voicebox"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Voice profiles are the foundation of Voicebox's voice cloning capability. Each profile stores reference audio samples and metadata that the TTS model uses to clone a voice.
|
||||
|
||||
## Architecture
|
||||
|
||||
The voice profile system consists of three main components:
|
||||
|
||||
**Database Layer:** SQLite tables store profile metadata and sample references.
|
||||
|
||||
**File Storage:** Audio samples are stored on disk in a structured directory format.
|
||||
|
||||
**Profile Module:** The `profiles.py` module provides the business logic for CRUD operations.
|
||||
|
||||
## Data Model
|
||||
|
||||
### VoiceProfile Table
|
||||
|
||||
```python
|
||||
class VoiceProfile(Base):
|
||||
__tablename__ = "profiles"
|
||||
|
||||
id = Column(String, primary_key=True)
|
||||
name = Column(String, unique=True, nullable=False)
|
||||
description = Column(Text)
|
||||
language = Column(String, default="en")
|
||||
created_at = Column(DateTime)
|
||||
updated_at = Column(DateTime)
|
||||
```
|
||||
|
||||
### ProfileSample Table
|
||||
|
||||
```python
|
||||
class ProfileSample(Base):
|
||||
__tablename__ = "profile_samples"
|
||||
|
||||
id = Column(String, primary_key=True)
|
||||
profile_id = Column(String, ForeignKey("profiles.id"))
|
||||
audio_path = Column(String, nullable=False)
|
||||
reference_text = Column(Text, nullable=False)
|
||||
```
|
||||
|
||||
## File Structure
|
||||
|
||||
Profiles are stored in the data directory:
|
||||
|
||||
```
|
||||
data/
|
||||
└── profiles/
|
||||
└── {profile_id}/
|
||||
├── {sample_id_1}.wav
|
||||
├── {sample_id_2}.wav
|
||||
└── ...
|
||||
```
|
||||
|
||||
## Core Functions
|
||||
|
||||
### Creating a Profile
|
||||
|
||||
```python
|
||||
async def create_profile(data: VoiceProfileCreate, db: Session) -> VoiceProfileResponse:
|
||||
# 1. Create database record
|
||||
db_profile = DBVoiceProfile(
|
||||
id=str(uuid.uuid4()),
|
||||
name=data.name,
|
||||
description=data.description,
|
||||
language=data.language,
|
||||
)
|
||||
db.add(db_profile)
|
||||
db.commit()
|
||||
|
||||
# 2. Create profile directory
|
||||
profile_dir = profiles_dir / db_profile.id
|
||||
profile_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
return VoiceProfileResponse.model_validate(db_profile)
|
||||
```
|
||||
|
||||
### Adding Samples
|
||||
|
||||
When a sample is added, the audio is validated and copied to the profile directory:
|
||||
|
||||
```python
|
||||
async def add_profile_sample(
|
||||
profile_id: str,
|
||||
audio_path: str,
|
||||
reference_text: str,
|
||||
db: Session,
|
||||
) -> ProfileSampleResponse:
|
||||
# 1. Validate audio (duration, format, quality)
|
||||
is_valid, error_msg = validate_reference_audio(audio_path)
|
||||
if not is_valid:
|
||||
raise ValueError(f"Invalid reference audio: {error_msg}")
|
||||
|
||||
# 2. Copy to profile directory
|
||||
sample_id = str(uuid.uuid4())
|
||||
dest_path = profile_dir / f"{sample_id}.wav"
|
||||
audio, sr = load_audio(audio_path)
|
||||
save_audio(audio, str(dest_path), sr)
|
||||
|
||||
# 3. Create database record
|
||||
db_sample = DBProfileSample(
|
||||
id=sample_id,
|
||||
profile_id=profile_id,
|
||||
audio_path=str(dest_path),
|
||||
reference_text=reference_text,
|
||||
)
|
||||
db.add(db_sample)
|
||||
db.commit()
|
||||
```
|
||||
|
||||
### Voice Prompt Creation
|
||||
|
||||
When generating speech, samples are combined into a voice prompt:
|
||||
|
||||
```python
|
||||
async def create_voice_prompt_for_profile(
|
||||
profile_id: str,
|
||||
db: Session,
|
||||
) -> dict:
|
||||
samples = db.query(DBProfileSample).filter_by(profile_id=profile_id).all()
|
||||
|
||||
if len(samples) == 1:
|
||||
# Single sample - use directly
|
||||
voice_prompt, _ = await tts_model.create_voice_prompt(
|
||||
sample.audio_path,
|
||||
sample.reference_text,
|
||||
)
|
||||
else:
|
||||
# Multiple samples - combine them
|
||||
combined_audio, combined_text = await tts_model.combine_voice_prompts(
|
||||
[s.audio_path for s in samples],
|
||||
[s.reference_text for s in samples],
|
||||
)
|
||||
voice_prompt, _ = await tts_model.create_voice_prompt(
|
||||
combined_audio_path,
|
||||
combined_text,
|
||||
)
|
||||
|
||||
return voice_prompt
|
||||
```
|
||||
|
||||
## Audio Validation
|
||||
|
||||
Reference audio is validated before being accepted:
|
||||
|
||||
- **Duration:** 3-30 seconds recommended
|
||||
- **Format:** WAV, MP3, FLAC, OGG supported
|
||||
- **Sample Rate:** Resampled to 24kHz
|
||||
- **Channels:** Converted to mono if stereo
|
||||
|
||||
## Export/Import
|
||||
|
||||
Profiles can be exported as ZIP archives for sharing:
|
||||
|
||||
```
|
||||
profile_export.zip
|
||||
├── profile.json # Metadata
|
||||
├── samples/
|
||||
│ ├── sample_1.wav
|
||||
│ └── sample_1.json # Reference text
|
||||
└── ...
|
||||
```
|
||||
|
||||
## API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| GET | `/profiles` | List all profiles |
|
||||
| POST | `/profiles` | Create a profile |
|
||||
| GET | `/profiles/{id}` | Get profile by ID |
|
||||
| PUT | `/profiles/{id}` | Update profile |
|
||||
| DELETE | `/profiles/{id}` | Delete profile |
|
||||
| GET | `/profiles/{id}/samples` | Get profile samples |
|
||||
| POST | `/profiles/{id}/samples` | Add sample to profile |
|
||||
| PUT | `/profiles/samples/{id}` | Update sample text |
|
||||
| DELETE | `/profiles/samples/{id}` | Delete sample |
|
||||
| GET | `/profiles/{id}/export` | Export as ZIP |
|
||||
| POST | `/profiles/import` | Import from ZIP |
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Sample Quality
|
||||
|
||||
- Use clean audio with minimal background noise
|
||||
- Ensure the reference text exactly matches what is spoken
|
||||
- Multiple samples (3-5) improve voice cloning quality
|
||||
|
||||
### Language Matching
|
||||
|
||||
- Set the profile language to match the reference audio
|
||||
- Supported languages: en, zh, ja, ko, de, fr, ru, pt, es, it
|
||||
|
||||
### Naming Conventions
|
||||
|
||||
- Use descriptive names that identify the voice
|
||||
- Avoid special characters that may cause filesystem issues
|
||||
Reference in New Issue
Block a user