--- title: "Voice Cloning" description: "Clone any voice from just a few seconds of audio" --- ## Overview Voicebox uses **Qwen3-TTS** from Alibaba to achieve near-perfect voice cloning from just a few seconds of audio. The model captures prosody, emotion, and natural cadence. ## How It Works Provide 10-30 seconds of clear speech from the target voice Qwen3-TTS analyzes vocal characteristics, tone, and speaking patterns The model generates a voice embedding for synthesis Use the profile to generate any text in the cloned voice ## Best Practices ### Sample Quality - Use 10-30 seconds of audio - Clear, consistent speaking - Minimal background noise - Natural speaking pace - Very short clips (< 5 seconds) - Heavy background noise - Music or overlapping voices - Heavily processed audio ### Multiple Samples Adding multiple samples from the same speaker can improve quality: - Different speaking styles (casual, formal) - Different emotions (happy, serious) - Different recording conditions The model will learn a more robust representation from diverse samples. ## Supported Languages Currently supported: - English - Chinese (Mandarin) More languages coming soon. ## Limitations Voice cloning should only be used with consent. Ensure you have permission to clone someone's voice. - Quality depends on sample clarity - Works best with consistent speaking tone - May struggle with extreme accents or speech impediments - Background noise reduces quality