---
title: "Creating Voice Profiles"
description: "Advanced guide to creating high-quality voice profiles"
---
## Overview
Voice profiles are the foundation of voice cloning in Voicebox. This guide covers best practices for creating professional-quality voice profiles.
## Quick Start
10-30 seconds of clear speech
**Profiles** → **+ New Profile**
Add your audio file
Use the profile to generate speech
## Audio Requirements
### Ideal Sample Characteristics
**10-30 seconds**
Too short: Poor quality
Too long: Unnecessary
**Clear speech**
No background noise
No music or overlapping voices
**High fidelity**
44.1kHz or 48kHz sample rate
Minimal compression
**Natural speech**
Conversational tone
Complete sentences
### File Formats
Supported formats:
- **WAV** (recommended) - Lossless quality
- **MP3** - Acceptable, minimal compression
- **M4A** - Acceptable
- **FLAC** - Lossless alternative
Use WAV for best results. Avoid heavily compressed formats.
## Recording Tips
### Environment
- Record in a quiet room
- Turn off fans, AC, appliances
- Close windows to reduce outside noise
- Use soft furnishings to reduce echo
- 6-12 inches from mouth
- Slight angle to reduce plosives (p, b, t)
- Use a pop filter if available
- Maintain consistent distance
- 44.1kHz or 48kHz sample rate
- 16-bit or 24-bit depth
- Mono is fine (stereo will be converted)
- Avoid automatic gain control
### Speaking
- **Natural pace** - Don't rush or speak too slowly
- **Clear articulation** - Pronounce words clearly
- **Consistent volume** - Maintain steady loudness
- **Normal tone** - Speak as you normally would
- **Complete sentences** - Avoid fragments or "ums"
## Multiple Samples
Adding multiple samples can significantly improve quality:
### Why Multiple Samples?
Model learns a more complete representation
Handles different speaking styles better
Reduces artifacts and improves naturalness
More reliable across different texts
### Sample Variety
Consider adding samples with:
1. **Different tones**
- Casual conversation
- Professional/formal
- Excited/enthusiastic
- Calm/serious
2. **Different content**
- Narratives
- Questions
- Statements
- Emotions (happy, sad, neutral)
3. **Different recording conditions**
- Studio quality
- Phone call quality (if needed)
- Room acoustics
All samples should be from the **same speaker**. Mixing voices will produce poor results.
## Processing Existing Audio
If you have existing audio (podcasts, videos, etc.):
### Extracting Clean Segments
Look for segments with:
- Just the target speaker
- No background music
- Minimal noise
Tools like Audacity or Adobe Audition:
- Cut out clean 10-30s segments
- Remove silence at start/end
- Normalize volume if needed
Save as high-quality WAV file
### Noise Reduction
If you have light background noise:
```
1. Use noise reduction in Audacity:
- Select noise-only section
- Get Noise Profile
- Select full audio
- Apply noise reduction (gentle settings)
2. Avoid over-processing:
- Can introduce artifacts
- May reduce voice quality
```
## Testing & Iteration
### Test Your Profile
After creating a profile:
Generate a simple phrase:
```
"Hello, this is a test of my voice profile."
```
Listen for:
- Natural tone
- Clear pronunciation
- Proper prosody
- Lack of artifacts
If quality is poor:
- Add more samples
- Try different source audio
- Check sample quality
### Common Issues
**Cause**: Poor quality samples or too short
**Fix**: Use longer, higher quality samples
**Cause**: Sample tone doesn't match desired output
**Fix**: Record samples in the style you want to generate
**Cause**: Background noise or audio issues in samples
**Fix**: Clean up samples or re-record in quieter environment
## Advanced Tips
### Celebrity/Character Voices
For cloning public figures or characters:
1. **Legal considerations** - Ensure you have rights or it's fair use
2. **Source quality** - Find high-quality interview audio or clean clips
3. **Consistency** - Use clips where they speak similarly
4. **Multiple samples** - Very important for recognizable voices
### Accent & Dialect
The model will preserve accent and dialect:
- British English will generate British English
- Southern accent will produce Southern accent
- Regional pronunciations will be maintained
### Emotion Transfer
The emotional tone of samples affects generation:
- Energetic samples → Energetic output
- Calm samples → Calm output
- Mix samples for versatile profile
## Managing Profiles
### Organization
- **Descriptive names** - "John Smith - Professional Narrator"
- **Add descriptions** - Note recording conditions, use cases
- **Language tags** - Mark the primary language
- **Archive unused** - Keep profile list manageable
### Export/Import
- **Export** profiles to share or backup
- **Import** from colleagues or teammates
- Profiles include voice embeddings, not original audio
## Next Steps
Use your profile to generate speech
Create multi-voice narratives