_model: project --- schema_version: 1 --- project_id: talkbox --- title: TalkBox --- summary: Local-first AI voice studio for multi-engine speech generation, zero-shot voice cloning, global dictation, and MCP agent voice I/O. --- status: active --- started: 2026-08-24 --- author: Labyricorn --- repository_url: https://git.labyricorn.com/Labyricorn/TalkBox --- default_branch: main --- tags: tts, voice-cloning, stt, whisper, mcp, tauri, react, fastapi, rust, python --- body: TalkBox is a local-first AI voice studio combining high-fidelity speech synthesis, zero-shot voice cloning, global dictation with synthetic paste, and Model Context Protocol (MCP) voice I/O. ### Highlights - **Multi-Engine Speech Generation**: Supports 7 TTS engines (Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, HumeAI TADA, Kokoro) covering 23 languages. - **Voice Cloning & Profiles**: Zero-shot voice cloning from reference audio samples, custom effects chains, and personality prompt attachments. - **Global Dictation**: Whisper-backed push-to-talk and toggle dictation anywhere on macOS, Windows, and Linux with focus-aware automatic pasting. - **Agent Voice Output (MCP)**: Streamable HTTP and stdio MCP server (`http://127.0.0.1:17494/mcp`) exposing `talkbox.speak`, `talkbox.transcribe`, `talkbox.list_captures`, and `talkbox.list_profiles` for agent integrations. - **Native & Local Performance**: Built on Tauri v2 (Rust) and React/TypeScript frontend with Python FastAPI backend running locally on Apple Silicon (MLX), NVIDIA (CUDA), AMD (ROCm), DirectML, or CPU.