* feat(i18n): add i18next foundation with English + zh-CN locales Installs i18next + react-i18next + language detector and wires up a language selector in the General settings page. Extracts strings from the highest-visibility surfaces: all settings tabs, model management, sidebar nav, main editor, and the floating generate box. Remaining strings (profile forms, history, stories/effects/voices/audio tabs) can land in follow-up PRs. Closes #411. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * fix(i18n): ensure language switch actually re-renders the tree - `nonExplicitSupportedLngs: true` was normalizing `zh-CN` → `zh` in some code paths; since we have explicit `zh-CN` resources, swap it for `load: 'currentOnly'` which keeps the code as-is. - `react: { useSuspense: false }` — react-i18next v17 defaults Suspense on, which can silently suspend components mid-switch and look like "nothing happens" to the user. - Use `i18n.language` (the raw current code) instead of `resolvedLanguage` in the selector so the dropdown always mirrors what we just set. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize ProfileCard, HistoryTable, and relative dates - ProfileCard: "No description", "designed" badge, aria-labels, and delete dialog. - HistoryTable: delete / clear-failed / import / effects dialogs. - formatDate: switch date-fns `formatDistance` locale based on `i18n.language` so "5 minutes ago" becomes "5 分钟前" under zh-CN. HistoryTable now subscribes via useTranslation so the table re-renders when language flips. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize Stories tab (list, content, dialogs, toasts) Covers the title + "New Story" button, empty states, story row metadata (item count, updated time), the create/edit/delete dialogs, and all toast notifications. Also handles StoryContent: "Select a story" placeholder, search popover, "Export Audio" button, and the "Generating N audios" pending indicator. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize history item and story item dropdown menus Covers the "..." action menu on both the History table (Play, Export Audio, Export Package, Apply Effects, Regenerate, Delete) and on individual story chat items (Play from here, Remove from Story). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize Effects tab (list, detail, dialogs, toasts) Covers EffectsList (title, "New Preset", section headers, preset cards) and EffectsDetail (header buttons for Save / Save as Custom / Delete, name/description fields, preview section, Save as Custom dialog, all toasts). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize EffectsChainEditor and built-in preset names - Effect type labels (Chorus/Flanger, Reverb, Delay, Compressor, Gain, High-Pass, Low-Pass, Pitch Shift) and every param label (LFO speed, Modulation depth, Threshold, Ratio, etc.) go through `effects.types.<type>.{label,params.<param>}` with the backend string as defaultValue fallback. - Chain-level controls: "Load preset…", "Add effect…", "Clear", Power/Remove button titles. - Built-in preset names + descriptions (Robotic, Radio, Echo Chamber, Deep Voice) are translated client-side; user-created presets keep their original names. Backend keeps returning English — frontend intercepts and translates via key lookup, defaulting to the backend string so unknown effects/params don't break. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize Create/Edit Voice modal and audio sample panels ProfileForm now routes its title, description, voice-source toggle (Clone from audio / Built-in voice), field labels (Name, Description, Language, Engine, Voice, Reference Text, Default Engine, Default Effects), sample tabs (Upload / Record / System Audio), action buttons, and every toast + Zod validation message through i18n. Also covers the three AudioSample panels (Upload/Record/System) — the choose-file / start-recording / start-capture call-to-actions, the "N remaining" countdown, "Recording complete" / "Capture complete" states, and the Play / Transcribe / Remove / Record Again buttons. SampleList too — the "No samples yet" empty state, per-sample edit mode, mini-player aria labels, Delete Sample dialog, and toasts. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize Audio Channels tab (list, dialogs, device picker) Covers the "Audio Channels" title and "New Channel" button, the empty state, per-channel section labels (Output Devices / Assigned Voices), the Available Devices right pane with its three contextual hints, the "No voices assigned" fallback, and both Create/Edit dialogs (titles, descriptions, field labels, Select placeholders, and the "(default)" badge). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize Voices tab (table header, search, inspector) VoicesTab now translates the "Voices" title, search placeholder, "New Voice" button, all six table column headers (Name, Language, Generations, Samples, Effects, Channels), the avatar alt text, and the per-row channel MultiSelect (placeholder + "(Default)" suffix). VoiceInspector routes its form labels through the existing `profileForm.fields.*` keys, has its own "Default Effects" hint and avatar/save toasts, and reuses the ProfileForm Zod validation. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): localize ProfileList unsupported-model note and empty state Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): add Traditional Chinese (zh-TW) locale Adds a zh-TW translation with Taiwan vocabulary conventions (e.g. 預設 / 儲存 / 載入 / 匯入 / 匯出 / 設定 / 檔案 / 伺服器 / 裝置 / 網路). Registers it alongside en and zh-CN; the language dropdown picks it up automatically from SUPPORTED_LANGUAGES. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * feat(i18n): add Japanese (ja) locale Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * fix(i18n): localize relative dates for ja and zh-TW formatDate only mapped zh-CN, so history timestamps stayed in English for ja and zh-TW users even after the rest of the UI translated. Extend the switch to ja and zhTW from date-fns/locale. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * fix(i18n): address PR review feedback Bugs: - GeneralPage: network access toast used the keep-server-running title key (wrong semantic scope). Add networkAccess.updatedTitle and use it. - GeneralPage: fallback "Unknown" version was stored as a translated string in state, so it stayed stale across language switches. Store null, resolve the label at render time. - GeneralPage: memoize the zod resolver on t and retrigger validation when the locale changes so existing error messages retranslate. - GpuPage: adding t to the CUDA progress EventSource effect deps caused the SSE connection to be torn down and reopened on every language change, potentially dropping in-flight download events. Capture t in a ref. - HistoryTable: Effects dialog still rendered English "Source" / "Select source version" / "Cancel" / "Apply" / "Applying..." — localize them. - Locales: zh-CN / zh-TW / ja devSuffix was missing the leading space before "(开发版)"/"(開發版)"/"(開発版)", so dev builds rendered "v0.4.2(开发版)" instead of "v0.4.2 (开发版)". Nits: - ModelManagement: rename .find((t) => ...) callback param to avoid shadowing useTranslation().t. - GenerationPage: rename chunkLimit.value interpolation key from count → chars so i18next doesn't silently activate pluralization if a translator later adds _one/_other forms. - LanguageSelect: narrow onValueChange handler param to LanguageCode. Key count now 559 across en/zh-CN/zh-TW/ja (added 4 effectsDialog keys plus networkAccess.updatedTitle). Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> --------- Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
Voicebox
The open-source voice synthesis studio.
Clone voices. Generate speech. Apply effects. Build voice-powered apps.
All running locally on your machine.
voicebox.sh • Docs • Download • Features • API • Troubleshooting
Click the image above to watch the demo video on voicebox.sh
What is Voicebox?
Voicebox is a local-first voice cloning studio — a free and open-source alternative to ElevenLabs. Clone voices from a few seconds of audio or pick from 50+ preset voices, generate speech in 23 languages across 7 TTS engines, apply post-processing effects, and compose multi-voice projects with a timeline editor.
- Complete privacy — models and voice data stay on your machine
- 7 TTS engines — Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, HumeAI TADA, and Kokoro
- Cloning and preset voices — zero-shot cloning from a reference sample, or curated preset voices via Kokoro (50 voices) and Qwen CustomVoice (9 voices)
- 23 languages — from English to Arabic, Japanese, Hindi, Swahili, and more
- Post-processing effects — pitch shift, reverb, delay, chorus, compression, and filters
- Expressive speech — paralinguistic tags like
[laugh],[sigh],[gasp]via Chatterbox Turbo; natural-language delivery control via Qwen CustomVoice - Unlimited length — auto-chunking with crossfade for scripts, articles, and chapters
- Stories editor — multi-track timeline for conversations, podcasts, and narratives
- API-first — REST API for integrating voice synthesis into your own projects
- Native performance — built with Tauri (Rust), not Electron
- Runs everywhere — macOS (MLX/Metal), Windows (CUDA), Linux, AMD ROCm, Intel Arc, Docker
Download
| Platform | Download |
|---|---|
| macOS (Apple Silicon) | Download DMG |
| macOS (Intel) | Download DMG |
| Windows | Download MSI |
| Docker | docker compose up |
Linux — Pre-built binaries are not yet available. See voicebox.sh/linux-install for build-from-source instructions.
Having trouble? See the Troubleshooting Guide for common install, generation, model-download, and GPU issues.
Features
Multi-Engine Voice Cloning
Seven TTS engines with different strengths, switchable per-generation:
| Engine | Languages | Strengths |
|---|---|---|
| Qwen3-TTS (0.6B / 1.7B) | 10 | High-quality multilingual cloning, delivery instructions ("speak slowly", "whisper") |
| Qwen CustomVoice | 10 | 9 curated preset voices with natural-language delivery control — no reference audio required |
| LuxTTS | English | Lightweight (~1GB VRAM), 48kHz output, 150x realtime on CPU |
| Chatterbox Multilingual | 23 | Broadest language coverage — Arabic, Danish, Finnish, Greek, Hebrew, Hindi, Malay, Norwegian, Polish, Swahili, Swedish, Turkish and more |
| Chatterbox Turbo | English | Fast 350M model with paralinguistic emotion/sound tags |
| TADA (1B / 3B) | 10 | HumeAI speech-language model — 700s+ coherent audio, text-acoustic dual alignment |
| Kokoro | 8 | 50 curated preset voices, tiny 82M model, fast CPU inference |
Emotions & Paralinguistic Tags
Only Chatterbox Turbo interprets paralinguistic tags like [laugh] and
[sigh]. Qwen3-TTS, LuxTTS, Chatterbox Multilingual, and HumeAI TADA read them
literally as text.
With Chatterbox Turbo selected, type / in the text input to open the tag
inserter and add expressive tags inline with speech:
[laugh] [chuckle] [gasp] [cough] [sigh] [groan] [sniff] [shush] [clear throat]
Post-Processing Effects
8 audio effects powered by Spotify's pedalboard library. Apply after generation, preview in real time, build reusable presets.
| Effect | Description |
|---|---|
| Pitch Shift | Up or down by up to 12 semitones |
| Reverb | Configurable room size, damping, wet/dry mix |
| Delay | Echo with adjustable time, feedback, and mix |
| Chorus / Flanger | Modulated delay for metallic or lush textures |
| Compressor | Dynamic range compression |
| Gain | Volume adjustment (-40 to +40 dB) |
| High-Pass Filter | Remove low frequencies |
| Low-Pass Filter | Remove high frequencies |
Ships with 4 built-in presets (Robotic, Radio, Echo Chamber, Deep Voice) and supports custom presets. Effects can be assigned per-profile as defaults.
Unlimited Generation Length
Text is automatically split at sentence boundaries and each chunk is generated independently, then crossfaded together. Works with all engines.
- Configurable auto-chunking limit (100–5,000 chars)
- Crossfade slider (0–200ms) for smooth transitions
- Max text length: 50,000 characters
- Smart splitting respects abbreviations, CJK punctuation, and
[tags]
Generation Versions
Every generation supports multiple versions with provenance tracking:
- Original — clean TTS output, always preserved
- Effects versions — apply different effects chains from any source version
- Takes — regenerate with a new seed for variation
- Source tracking — each version records its lineage
- Favorites — star generations for quick access
Async Generation Queue
Generation is non-blocking. Submit and immediately start typing the next one.
- Serial execution queue prevents GPU contention
- Real-time SSE status streaming
- Failed generations can be retried
- Stale generations from crashes auto-recover on startup
Voice Profile Management
- Create profiles from audio files or record directly in-app
- Import/export profiles to share or back up
- Multi-sample support for higher quality cloning
- Per-profile default effects chains
- Organize with descriptions and language tags
Stories Editor
Multi-voice timeline editor for conversations, podcasts, and narratives.
- Multi-track composition with drag-and-drop
- Inline audio trimming and splitting
- Auto-playback with synchronized playhead
- Version pinning per track clip
Recording & Transcription
- In-app recording with waveform visualization
- System audio capture (macOS and Windows)
- Automatic transcription powered by Whisper (including Whisper Turbo)
- Export recordings in multiple formats
Model Management
- Per-model unload to free GPU memory without deleting downloads
- Custom models directory via
VOICEBOX_MODELS_DIR - Model folder migration with progress tracking
- Download cancel/clear UI
GPU Support
| Platform | Backend | Notes |
|---|---|---|
| macOS (Apple Silicon) | MLX (Metal) | 4-5x faster via Neural Engine |
| Windows / Linux (NVIDIA) | PyTorch (CUDA) | Auto-downloads CUDA binary from within the app |
| Linux (AMD) | PyTorch (ROCm) | Auto-configures HSA_OVERRIDE_GFX_VERSION |
| Windows (any GPU) | DirectML | Universal Windows GPU support |
| Intel Arc | IPEX/XPU | Intel discrete GPU acceleration |
| Any | CPU | Works everywhere, just slower |
API
Voicebox exposes a full REST API for integrating voice synthesis into your own apps.
# Generate speech
curl -X POST http://localhost:17493/generate \
-H "Content-Type: application/json" \
-d '{"text": "Hello world", "profile_id": "abc123", "language": "en"}'
# List voice profiles
curl http://localhost:17493/profiles
# Create a profile
curl -X POST http://localhost:17493/profiles \
-H "Content-Type: application/json" \
-d '{"name": "My Voice", "language": "en"}'
Use cases: game dialogue, podcast production, accessibility tools, voice assistants, content automation.
Full API documentation available at http://localhost:17493/docs.
Tech Stack
| Layer | Technology |
|---|---|
| Desktop App | Tauri (Rust) |
| Frontend | React, TypeScript, Tailwind CSS |
| State | Zustand, React Query |
| Backend | FastAPI (Python) |
| TTS Engines | Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox, Chatterbox Turbo, TADA, Kokoro |
| Effects | Pedalboard (Spotify) |
| Transcription | Whisper / Whisper Turbo (PyTorch or MLX) |
| Inference | MLX (Apple Silicon) / PyTorch (CUDA/ROCm/XPU/CPU) |
| Database | SQLite |
| Audio | WaveSurfer.js, librosa |
Roadmap
| Feature | Description |
|---|---|
| Real-time Streaming | Stream audio as it generates, word by word |
| Voice Design | Create new voices from text descriptions |
| More Models | XTTS, Bark, and other open-source voice models |
| Plugin Architecture | Extend with custom models and effects |
| Mobile Companion | Control Voicebox from your phone |
For the full engineering status, open-issue triage, and prioritized work queue, see docs/PROJECT_STATUS.md — a living document that tracks what's shipped, what's in-flight, candidate TTS engines under evaluation, and why we've accepted or backlogged specific integrations.
Development
See CONTRIBUTING.md for detailed setup and contribution guidelines.
Quick Start
git clone https://github.com/jamiepine/voicebox.git
cd voicebox
just setup # creates Python venv, installs all deps
just dev # starts backend + desktop app
Install just: brew install just or cargo install just. Run just --list to see all commands.
Prerequisites: Bun, Rust, Python 3.11+, Tauri Prerequisites, and Xcode on macOS.
Building Locally
just build # Build CPU server binary + Tauri app
just build-local # (Windows) Build CPU + CUDA server binaries + Tauri app
Adding New Voice Models
The multi-engine architecture makes adding new TTS engines straightforward. A step-by-step guide covers the full process: dependency research, backend protocol implementation, frontend wiring, and PyInstaller bundling.
The guide is optimized for AI coding agents. An agent skill can pick up a model name and handle the entire integration autonomously — you just test the build locally.
Project Structure
voicebox/
├── app/ # Shared React frontend
├── tauri/ # Desktop app (Tauri + Rust)
├── web/ # Web deployment
├── backend/ # Python FastAPI server
├── landing/ # Marketing website
└── scripts/ # Build & release scripts
Contributing
Contributions welcome! See CONTRIBUTING.md for guidelines.
- Fork the repo
- Create a feature branch
- Make your changes
- Submit a PR
Security
Found a security vulnerability? Please report it responsibly. See SECURITY.md for details.
License
MIT License — see LICENSE for details.

