--- title: "Creating Voice Profiles" description: "Advanced guide to creating high-quality voice profiles" --- ## Overview Voice profiles are the foundation of voice cloning in Voicebox. This guide covers best practices for creating professional-quality voice profiles. ## Quick Start 10-30 seconds of clear speech **Profiles** → **+ New Profile** Add your audio file Use the profile to generate speech ## Audio Requirements ### Ideal Sample Characteristics **10-30 seconds** Too short: Poor quality Too long: Unnecessary **Clear speech** No background noise No music or overlapping voices **High fidelity** 44.1kHz or 48kHz sample rate Minimal compression **Natural speech** Conversational tone Complete sentences ### File Formats Supported formats: - **WAV** (recommended) - Lossless quality - **MP3** - Acceptable, minimal compression - **M4A** - Acceptable - **FLAC** - Lossless alternative Use WAV for best results. Avoid heavily compressed formats. ## Recording Tips ### Environment - Record in a quiet room - Turn off fans, AC, appliances - Close windows to reduce outside noise - Use soft furnishings to reduce echo - 6-12 inches from mouth - Slight angle to reduce plosives (p, b, t) - Use a pop filter if available - Maintain consistent distance - 44.1kHz or 48kHz sample rate - 16-bit or 24-bit depth - Mono is fine (stereo will be converted) - Avoid automatic gain control ### Speaking - **Natural pace** - Don't rush or speak too slowly - **Clear articulation** - Pronounce words clearly - **Consistent volume** - Maintain steady loudness - **Normal tone** - Speak as you normally would - **Complete sentences** - Avoid fragments or "ums" ## Multiple Samples Adding multiple samples can significantly improve quality: ### Why Multiple Samples? Model learns a more complete representation Handles different speaking styles better Reduces artifacts and improves naturalness More reliable across different texts ### Sample Variety Consider adding samples with: 1. **Different tones** - Casual conversation - Professional/formal - Excited/enthusiastic - Calm/serious 2. **Different content** - Narratives - Questions - Statements - Emotions (happy, sad, neutral) 3. **Different recording conditions** - Studio quality - Phone call quality (if needed) - Room acoustics All samples should be from the **same speaker**. Mixing voices will produce poor results. ## Processing Existing Audio If you have existing audio (podcasts, videos, etc.): ### Extracting Clean Segments Look for segments with: - Just the target speaker - No background music - Minimal noise Tools like Audacity or Adobe Audition: - Cut out clean 10-30s segments - Remove silence at start/end - Normalize volume if needed Save as high-quality WAV file ### Noise Reduction If you have light background noise: ``` 1. Use noise reduction in Audacity: - Select noise-only section - Get Noise Profile - Select full audio - Apply noise reduction (gentle settings) 2. Avoid over-processing: - Can introduce artifacts - May reduce voice quality ``` ## Testing & Iteration ### Test Your Profile After creating a profile: Generate a simple phrase: ``` "Hello, this is a test of my voice profile." ``` Listen for: - Natural tone - Clear pronunciation - Proper prosody - Lack of artifacts If quality is poor: - Add more samples - Try different source audio - Check sample quality ### Common Issues **Cause**: Poor quality samples or too short **Fix**: Use longer, higher quality samples **Cause**: Sample tone doesn't match desired output **Fix**: Record samples in the style you want to generate **Cause**: Background noise or audio issues in samples **Fix**: Clean up samples or re-record in quieter environment ## Advanced Tips ### Celebrity/Character Voices For cloning public figures or characters: 1. **Legal considerations** - Ensure you have rights or it's fair use 2. **Source quality** - Find high-quality interview audio or clean clips 3. **Consistency** - Use clips where they speak similarly 4. **Multiple samples** - Very important for recognizable voices ### Accent & Dialect The model will preserve accent and dialect: - British English will generate British English - Southern accent will produce Southern accent - Regional pronunciations will be maintained ### Emotion Transfer The emotional tone of samples affects generation: - Energetic samples → Energetic output - Calm samples → Calm output - Mix samples for versatile profile ## Managing Profiles ### Organization - **Descriptive names** - "John Smith - Professional Narrator" - **Add descriptions** - Note recording conditions, use cases - **Language tags** - Mark the primary language - **Archive unused** - Keep profile list manageable ### Export/Import - **Export** profiles to share or backup - **Import** from colleagues or teammates - Profiles include voice embeddings, not original audio ## Next Steps Use your profile to generate speech Create multi-voice narratives