rewrite docs introduction based on README content

This commit is contained in:
James Pine
2026-03-16 04:09:50 -07:00
parent 7c4afbe4df
commit a8968d4081
+47 -41
View File
@@ -1,58 +1,64 @@
--- ---
title: "Introduction" title: "Introduction"
description: "Welcome to Voicebox - the open-source voice synthesis studio" description: "Voicebox is a local-first voice cloning studio -- a free and open-source alternative to ElevenLabs."
--- ---
## What is Voicebox? ## What is Voicebox?
Voicebox is a **local-first voice cloning studio** with DAW-like features for professional voice synthesis. Think of it as the **Ollama for voice** — download models, clone voices, and generate speech entirely on your machine. Voicebox is a **local-first voice cloning studio** -- a free and open-source alternative to ElevenLabs. Clone voices from a few seconds of audio, generate speech in 23 languages across 4 TTS engines, apply post-processing effects, and compose multi-voice projects with a timeline editor.
<Frame> - **Complete privacy** -- models and voice data stay on your machine
<img src="/images/app-screenshot-1.webp" alt="Voicebox App Screenshot" /> - **4 TTS engines** -- Qwen3-TTS, LuxTTS, Chatterbox Multilingual, and Chatterbox Turbo
</Frame> - **23 languages** -- from English to Arabic, Japanese, Hindi, Swahili, and more
- **Post-processing effects** -- pitch shift, reverb, delay, chorus, compression, and filters
- **Expressive speech** -- paralinguistic tags like `[laugh]`, `[sigh]`, `[gasp]` via Chatterbox Turbo
- **Unlimited length** -- auto-chunking with crossfade for scripts, articles, and chapters
- **Stories editor** -- multi-track timeline for conversations, podcasts, and narratives
- **API-first** -- REST API for integrating voice synthesis into your own projects
- **Native performance** -- built with Tauri (Rust), not Electron
- **Runs everywhere** -- macOS (MLX/Metal), Windows (CUDA), Linux, AMD ROCm, Intel Arc, Docker
Unlike cloud services that lock your voice data behind subscriptions, Voicebox gives you: ## TTS Engines
- **Complete privacy** — models and voice data stay on your machine Four engines with different strengths, switchable per-generation:
- **Professional tools** — multi-track timeline editor, audio trimming, conversation mixing
- **Model flexibility** — currently powered by Qwen3-TTS, with support for XTTS, Bark, and other models coming soon
- **API-first** — use the desktop app or integrate voice synthesis into your own projects
- **Native performance** — built with Tauri (Rust), not Electron
Download a voice model, clone any voice from a few seconds of audio, and compose multi-voice projects with studio-grade editing tools. No Python install required, no cloud dependency, no limits. | Engine | Languages | Strengths |
|--------|-----------|-----------|
| **Qwen3-TTS** (0.6B / 1.7B) | 10 | High-quality multilingual cloning, delivery instructions |
| **LuxTTS** | English | Lightweight (~1GB VRAM), 48kHz output, 150x realtime on CPU |
| **Chatterbox Multilingual** | 23 | Broadest language coverage |
| **Chatterbox Turbo** | English | Fast 350M model with paralinguistic emotion/sound tags |
## Key Features ## GPU Support
<CardGroup cols={2}> | Platform | Backend | Notes |
<Card title="Voice Cloning" icon="microphone"> |----------|---------|-------|
Instant cloning from just a few seconds of audio with Qwen3-TTS | macOS (Apple Silicon) | MLX (Metal) | 4-5x faster via Neural Engine |
</Card> | Windows / Linux (NVIDIA) | PyTorch (CUDA) | Auto-downloads CUDA binary from within the app |
<Card title="Stories Editor" icon="film"> | Linux (AMD) | PyTorch (ROCm) | Auto-configures HSA_OVERRIDE_GFX_VERSION |
Multi-track timeline for creating conversations and narratives | Windows (any GPU) | DirectML | Universal Windows GPU support |
</Card> | Intel Arc | IPEX/XPU | Intel discrete GPU acceleration |
<Card title="Full API" icon="code"> | Any | CPU | Works everywhere, just slower |
REST API for integrating voice synthesis into your apps
</Card>
<Card title="Local-First" icon="shield">
Everything runs on your machine - complete privacy
</Card>
</CardGroup>
## Use Cases ## Use Cases
- **Game Development** — Generate dynamic dialogue for characters - **Game development** -- generate dynamic dialogue for characters
- **Content Creation** — Produce podcasts and video voiceovers - **Content creation** -- produce podcasts and video voiceovers
- **Accessibility** — Build text-to-speech tools - **Accessibility** -- build text-to-speech tools for users who need them
- **Voice Assistants** — Create custom voice interfaces - **Voice assistants** -- create custom voice interfaces
- **Production Pipelines** — Automate voiceover workflows - **Production pipelines** -- automate voiceover workflows via the REST API
## Next Steps ## Tech Stack
<CardGroup cols={2}> | Layer | Technology |
<Card title="Installation" icon="download" href="/overview/installation"> |-------|------------|
Download and install Voicebox on your machine | Desktop App | Tauri (Rust) |
</Card> | Frontend | React, TypeScript, Tailwind CSS |
<Card title="Quick Start" icon="rocket" href="/overview/quick-start"> | State | Zustand, React Query |
Get up and running in 5 minutes | Backend | FastAPI (Python) |
</Card> | TTS Engines | Qwen3-TTS, LuxTTS, Chatterbox, Chatterbox Turbo |
</CardGroup> | Effects | Pedalboard (Spotify) |
| Transcription | Whisper / Whisper Turbo (PyTorch or MLX) |
| Inference | MLX (Apple Silicon) / PyTorch (CUDA/ROCm/XPU/CPU) |
| Database | SQLite |
| Audio | WaveSurfer.js, librosa |