mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-09-19 14:50:38 -07:00
docs
This commit is contained in:
@@ -28,32 +28,32 @@ Voice profiles are the foundation of voice cloning in Voicebox. This guide cover
|
||||
|
||||
### Ideal Sample Characteristics
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Duration" icon="clock">
|
||||
<Cards>
|
||||
<Card title="Duration">
|
||||
**10-30 seconds**
|
||||
|
||||
Too short: Poor quality
|
||||
Too long: Unnecessary
|
||||
</Card>
|
||||
<Card title="Clarity" icon="volume">
|
||||
<Card title="Clarity">
|
||||
**Clear speech**
|
||||
|
||||
No background noise
|
||||
No music or overlapping voices
|
||||
</Card>
|
||||
<Card title="Quality" icon="sparkles">
|
||||
<Card title="Quality">
|
||||
**High fidelity**
|
||||
|
||||
44.1kHz or 48kHz sample rate
|
||||
Minimal compression
|
||||
</Card>
|
||||
<Card title="Content" icon="microphone">
|
||||
<Card title="Content">
|
||||
**Natural speech**
|
||||
|
||||
Conversational tone
|
||||
Complete sentences
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
### File Formats
|
||||
|
||||
@@ -63,9 +63,9 @@ Supported formats:
|
||||
- **M4A** - Acceptable
|
||||
- **FLAC** - Lossless alternative
|
||||
|
||||
<Tip>
|
||||
<Callout type="info">
|
||||
Use WAV for best results. Avoid heavily compressed formats.
|
||||
</Tip>
|
||||
</Callout>
|
||||
|
||||
## Recording Tips
|
||||
|
||||
@@ -108,20 +108,20 @@ Adding multiple samples can significantly improve quality:
|
||||
|
||||
### Why Multiple Samples?
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Robustness" icon="shield">
|
||||
<Cards>
|
||||
<Card title="Robustness">
|
||||
Model learns a more complete representation
|
||||
</Card>
|
||||
<Card title="Versatility" icon="palette">
|
||||
<Card title="Versatility">
|
||||
Handles different speaking styles better
|
||||
</Card>
|
||||
<Card title="Quality" icon="star">
|
||||
<Card title="Quality">
|
||||
Reduces artifacts and improves naturalness
|
||||
</Card>
|
||||
<Card title="Consistency" icon="check">
|
||||
<Card title="Consistency">
|
||||
More reliable across different texts
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
### Sample Variety
|
||||
|
||||
@@ -144,9 +144,9 @@ Consider adding samples with:
|
||||
- Phone call quality (if needed)
|
||||
- Room acoustics
|
||||
|
||||
<Warning>
|
||||
<Callout type="warn">
|
||||
All samples should be from the **same speaker**. Mixing voices will produce poor results.
|
||||
</Warning>
|
||||
</Callout>
|
||||
|
||||
## Processing Existing Audio
|
||||
|
||||
@@ -286,11 +286,11 @@ The emotional tone of samples affects generation:
|
||||
|
||||
## Next Steps
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Generate Speech" icon="waveform" href="/overview/generating-speech">
|
||||
<Cards>
|
||||
<Card title="Generate Speech" href="/overview/generating-speech">
|
||||
Use your profile to generate speech
|
||||
</Card>
|
||||
<Card title="Build Stories" icon="film" href="/overview/building-stories">
|
||||
<Card title="Build Stories" href="/overview/building-stories">
|
||||
Create multi-voice narratives
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
@@ -43,9 +43,9 @@ Use formatting to suggest emphasis:
|
||||
- Bold for strong emphasis: "This is **very** important"
|
||||
```
|
||||
|
||||
<Note>
|
||||
<Callout type="info">
|
||||
The model interprets these hints but results may vary.
|
||||
</Note>
|
||||
</Callout>
|
||||
|
||||
## Advanced Features
|
||||
|
||||
|
||||
@@ -9,20 +9,20 @@ Voicebox keeps a complete history of all generated audio, making it easy to find
|
||||
|
||||
## Features
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Full History" icon="clock">
|
||||
<Cards>
|
||||
<Card title="Full History" icon={<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round"><circle cx="12" cy="12" r="10"/><polyline points="12 6 12 12 16 14"/></svg>}>
|
||||
Every generation is automatically saved
|
||||
</Card>
|
||||
<Card title="Search & Filter" icon="search">
|
||||
<Card title="Search & Filter" icon={<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round"><circle cx="11" cy="11" r="8"/><path d="m21 21-4.3-4.3"/></svg>}>
|
||||
Find by text, voice, or date
|
||||
</Card>
|
||||
<Card title="Re-generate" icon="rotate">
|
||||
<Card title="Re-generate" icon={<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round"><path d="M3 12a9 9 0 0 1 9-9 9.75 9.75 0 0 1 6.74 2.74L21 8"/><path d="M21 3v5h-5"/><path d="M21 12a9 9 0 0 1-9 9 9.75 9.75 0 0 1-6.74-2.74L3 16"/><path d="M8 16H3v5"/></svg>}>
|
||||
Regenerate any past generation with one click
|
||||
</Card>
|
||||
<Card title="Export" icon="download">
|
||||
<Card title="Export" icon={<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round"><path d="M21 15v4a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2v-4"/><polyline points="7 10 12 15 17 10"/><line x1="12" y1="15" x2="12" y2="3"/></svg>}>
|
||||
Download individual or batch exports
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
## Viewing History
|
||||
|
||||
@@ -54,20 +54,20 @@ Drag generations to the Stories Editor timeline.
|
||||
|
||||
## Search & Filter
|
||||
|
||||
<Tabs>
|
||||
<Tab title="By Text">
|
||||
<Tabs items={["By Text", "By Voice", "By Date"]}>
|
||||
<Tab value="By Text">
|
||||
Search for specific text content
|
||||
```
|
||||
"Hello world"
|
||||
```
|
||||
</Tab>
|
||||
<Tab title="By Voice">
|
||||
<Tab value="By Voice">
|
||||
Filter by voice profile
|
||||
```
|
||||
Select from dropdown
|
||||
```
|
||||
</Tab>
|
||||
<Tab title="By Date">
|
||||
<Tab value="By Date">
|
||||
Filter by date range
|
||||
```
|
||||
Last 7 days, Last 30 days, Custom range
|
||||
@@ -83,6 +83,6 @@ History is stored locally:
|
||||
- **Windows**: `%APPDATA%/com.voicebox.app/data/`
|
||||
- **Linux**: `~/.config/com.voicebox.app/data/`
|
||||
|
||||
<Warning>
|
||||
<Callout type="warn">
|
||||
Deleting the data directory will remove all history. Export important files first.
|
||||
</Warning>
|
||||
</Callout>
|
||||
|
||||
@@ -7,19 +7,19 @@ description: "Download and install Voicebox on macOS, Windows, or Linux"
|
||||
|
||||
Voicebox is available for macOS and Windows, with Linux builds coming soon.
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="macOS" icon="apple">
|
||||
<Cards>
|
||||
<Card title="macOS" icon={<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round"><path d="M12 2c-1.5 0-2.8.4-3.9 1.1A5.5 5.5 0 0 0 4 2.5C2.5 2.5 1 4 1 6c0 3.5 2.5 6 5 7.5C5 16 4 18 4 20c0 1.5.5 2.5 1.5 3C6.5 23.5 8 24 9.5 24c2 0 3.5-.5 5-2 1.5 1.5 3 2 5 2 1.5 0 3-.5 4-1 1-.5 1.5-1.5 1.5-3 0-2-1-4-2-6.5 2.5-1.5 5-4 5-7.5 0-2-1.5-3.5-3-3.5-.9 0-2.1.4-3.1 1.1A6.5 6.5 0 0 0 12 2Z"/></svg>}>
|
||||
Download for Apple Silicon or Intel Macs
|
||||
</Card>
|
||||
<Card title="Windows" icon="windows">
|
||||
<Card title="Windows" icon={<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round"><rect x="2" y="3" width="20" height="14" rx="2" ry="2"/><line x1="8" y1="21" x2="16" y2="21"/><line x1="12" y1="17" x2="12" y2="21"/></svg>}>
|
||||
Download MSI installer or Setup executable
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
### macOS
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Apple Silicon">
|
||||
<Tabs items={["Apple Silicon", "Intel"]}>
|
||||
<Tab value="Apple Silicon">
|
||||
Download: [voicebox_aarch64.app.tar.gz](https://github.com/jamiepine/voicebox/releases/latest/download/voicebox_aarch64.app.tar.gz)
|
||||
|
||||
```bash
|
||||
@@ -30,7 +30,7 @@ Voicebox is available for macOS and Windows, with Linux builds coming soon.
|
||||
mv Voicebox.app /Applications/
|
||||
```
|
||||
</Tab>
|
||||
<Tab title="Intel">
|
||||
<Tab value="Intel">
|
||||
Download: [voicebox_x64.app.tar.gz](https://github.com/jamiepine/voicebox/releases/latest/download/voicebox_x64.app.tar.gz)
|
||||
|
||||
```bash
|
||||
@@ -45,13 +45,13 @@ Voicebox is available for macOS and Windows, with Linux builds coming soon.
|
||||
|
||||
### Windows
|
||||
|
||||
<Tabs>
|
||||
<Tab title="MSI Installer">
|
||||
<Tabs items={["MSI Installer", "Setup Executable"]}>
|
||||
<Tab value="MSI Installer">
|
||||
Download: [voicebox_x64_en-US.msi](https://github.com/jamiepine/voicebox/releases/latest/download/voicebox_x64_en-US.msi)
|
||||
|
||||
Double-click the MSI file and follow the installation wizard.
|
||||
</Tab>
|
||||
<Tab title="Setup Executable">
|
||||
<Tab value="Setup Executable">
|
||||
Download: [voicebox_x64-setup.exe](https://github.com/jamiepine/voicebox/releases/latest/download/voicebox_x64-setup.exe)
|
||||
|
||||
Run the executable and follow the installation wizard.
|
||||
@@ -60,9 +60,9 @@ Voicebox is available for macOS and Windows, with Linux builds coming soon.
|
||||
|
||||
### Linux
|
||||
|
||||
<Note>
|
||||
<Callout type="info">
|
||||
Linux builds are coming soon. Currently blocked by GitHub runner disk space limitations.
|
||||
</Note>
|
||||
</Callout>
|
||||
|
||||
## First Launch
|
||||
|
||||
@@ -76,9 +76,9 @@ When you launch Voicebox for the first time:
|
||||
|
||||
3. **Backend Server** — The bundled Python server starts automatically
|
||||
|
||||
<Tip>
|
||||
<Callout type="info">
|
||||
First generation will be slower due to model downloads. Subsequent runs use cached models.
|
||||
</Tip>
|
||||
</Callout>
|
||||
|
||||
## System Requirements
|
||||
|
||||
@@ -95,9 +95,9 @@ When you launch Voicebox for the first time:
|
||||
- **GPU:** CUDA-capable NVIDIA GPU (for faster generation)
|
||||
- **Storage:** 10GB+ free space
|
||||
|
||||
<Note>
|
||||
<Callout type="info">
|
||||
CPU inference is supported but significantly slower than GPU. A CUDA-capable GPU is highly recommended for real-time workflows.
|
||||
</Note>
|
||||
</Callout>
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -108,12 +108,12 @@ After installation, verify everything works:
|
||||
3. Navigate to **Profiles** and create a test profile
|
||||
4. Generate a short audio clip to verify the TTS engine works
|
||||
|
||||
<Check>
|
||||
<Callout type="success">
|
||||
If you see a green status indicator and can generate audio, you're all set!
|
||||
</Check>
|
||||
</Callout>
|
||||
|
||||
## Next Steps
|
||||
|
||||
<Card title="Quick Start Guide" icon="rocket" href="/overview/quick-start">
|
||||
<Card title="Quick Start Guide" href="/overview/quick-start">
|
||||
Create your first voice profile and generate speech
|
||||
</Card>
|
||||
|
||||
@@ -46,9 +46,9 @@ Voice profiles are the foundation of Voicebox. Each profile contains voice sampl
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
<Callout type="info">
|
||||
For best results, use clean audio with minimal background noise and consistent speaking tone.
|
||||
</Tip>
|
||||
</Callout>
|
||||
|
||||
## Step 2: Generate Speech
|
||||
|
||||
@@ -74,9 +74,9 @@ Now let's use your new voice profile to generate speech.
|
||||
<Step title="Generate">
|
||||
Click **Generate** and wait a few seconds
|
||||
|
||||
<Note>
|
||||
<Callout type="info">
|
||||
First generation may take longer due to model initialization. Subsequent generations will be faster.
|
||||
</Note>
|
||||
</Callout>
|
||||
</Step>
|
||||
|
||||
<Step title="Play & Download">
|
||||
@@ -114,20 +114,20 @@ The Stories Editor lets you create multi-voice narratives with a timeline-based
|
||||
|
||||
## What's Next?
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Voice Cloning Guide" icon="microphone" href="/overview/creating-voice-profiles">
|
||||
<Cards>
|
||||
<Card title="Voice Cloning Guide" href="/overview/creating-voice-profiles">
|
||||
Learn advanced techniques for high-quality voice cloning
|
||||
</Card>
|
||||
<Card title="API Integration" icon="code" href="/api-reference">
|
||||
<Card title="API Integration" href="/api-reference">
|
||||
Integrate Voicebox into your own applications
|
||||
</Card>
|
||||
<Card title="Stories Editor" icon="film" href="/overview/stories-editor">
|
||||
<Card title="Stories Editor" href="/overview/stories-editor">
|
||||
Master the multi-track timeline editor
|
||||
</Card>
|
||||
<Card title="Remote Mode" icon="server" href="/overview/remote-mode">
|
||||
<Card title="Remote Mode" href="/overview/remote-mode">
|
||||
Connect to a GPU server for faster generation
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
## Tips for Success
|
||||
|
||||
|
||||
@@ -59,6 +59,6 @@ Automatic speech-to-text powered by OpenAI's Whisper model.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
<Tip>
|
||||
<Callout type="info">
|
||||
Transcription is useful for creating voice samples from existing audio or generating subtitles.
|
||||
</Tip>
|
||||
</Callout>
|
||||
|
||||
@@ -41,9 +41,9 @@ In Remote Mode, the Voicebox desktop app (running on your local machine) communi
|
||||
uvicorn main:app --host 0.0.0.0 --port 17493
|
||||
```
|
||||
|
||||
<Warning>
|
||||
<Callout type="warn">
|
||||
This exposes the server to your network. Use a firewall or VPN for security.
|
||||
</Warning>
|
||||
</Callout>
|
||||
</Step>
|
||||
|
||||
<Step title="Open Firewall">
|
||||
@@ -108,9 +108,9 @@ In Remote Mode, the Voicebox desktop app (running on your local machine) communi
|
||||
|
||||
## Security Considerations
|
||||
|
||||
<Warning>
|
||||
<Callout type="warn">
|
||||
The API currently has no authentication. Only use on trusted networks or with a VPN.
|
||||
</Warning>
|
||||
</Callout>
|
||||
|
||||
**Best Practices:**
|
||||
- Use a VPN (WireGuard, Tailscale) instead of exposing to the internet
|
||||
@@ -129,9 +129,9 @@ Expected performance on various GPUs:
|
||||
| RTX 3060 | ~5-7s per 10 words |
|
||||
| CPU (12-core) | ~20-30s per 10 words |
|
||||
|
||||
<Tip>
|
||||
<Callout type="info">
|
||||
A GPU with 8GB+ VRAM is recommended for best performance.
|
||||
</Tip>
|
||||
</Callout>
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
|
||||
@@ -9,20 +9,20 @@ The Stories Editor is a DAW-like timeline interface for creating multi-voice nar
|
||||
|
||||
## Features
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Multi-Track Timeline" icon="timeline">
|
||||
<Cards>
|
||||
<Card title="Multi-Track Timeline">
|
||||
Arrange multiple voice tracks in parallel
|
||||
</Card>
|
||||
<Card title="Inline Editing" icon="scissors">
|
||||
<Card title="Inline Editing">
|
||||
Trim and split clips directly in the timeline
|
||||
</Card>
|
||||
<Card title="Auto-Playback" icon="play">
|
||||
<Card title="Auto-Playback">
|
||||
Preview with synchronized playhead
|
||||
</Card>
|
||||
<Card title="Voice Mixing" icon="users">
|
||||
<Card title="Voice Mixing">
|
||||
Build conversations with multiple speakers
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
## Creating a Story
|
||||
|
||||
|
||||
@@ -25,9 +25,9 @@ Windows SmartScreen may warn that the app is unrecognized.
|
||||
- Click "More info"
|
||||
- Click "Run anyway"
|
||||
|
||||
<Note>
|
||||
<Callout type="info">
|
||||
This is expected for unsigned applications. We're working on code signing for future releases.
|
||||
</Note>
|
||||
</Callout>
|
||||
|
||||
## Server Issues
|
||||
|
||||
@@ -163,9 +163,9 @@ This is expected behavior. The first generation downloads the Qwen3-TTS model (~
|
||||
|
||||
Settings → Generation → Use CPU instead of GPU
|
||||
|
||||
<Warning>
|
||||
<Callout type="warn">
|
||||
CPU generation is 5-10x slower but uses system RAM instead of VRAM.
|
||||
</Warning>
|
||||
</Callout>
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Reduce Batch Size">
|
||||
@@ -322,9 +322,9 @@ bun run tauri build
|
||||
|
||||
**Solutions:**
|
||||
|
||||
<Warning>
|
||||
<Callout type="warn">
|
||||
This will delete all your voice profiles and generation history. Export important profiles first if possible.
|
||||
</Warning>
|
||||
</Callout>
|
||||
|
||||
```bash
|
||||
# macOS
|
||||
|
||||
@@ -28,20 +28,20 @@ Voicebox uses **Qwen3-TTS** from Alibaba to achieve near-perfect voice cloning f
|
||||
|
||||
### Sample Quality
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Do" icon="check">
|
||||
<Cards>
|
||||
<Card title="Do">
|
||||
- Use 10-30 seconds of audio
|
||||
- Clear, consistent speaking
|
||||
- Minimal background noise
|
||||
- Natural speaking pace
|
||||
</Card>
|
||||
<Card title="Don't" icon="xmark">
|
||||
<Card title="Don't">
|
||||
- Very short clips (< 5 seconds)
|
||||
- Heavy background noise
|
||||
- Music or overlapping voices
|
||||
- Heavily processed audio
|
||||
</Card>
|
||||
</CardGroup>
|
||||
</Cards>
|
||||
|
||||
### Multiple Samples
|
||||
|
||||
@@ -51,9 +51,9 @@ Adding multiple samples from the same speaker can improve quality:
|
||||
- Different emotions (happy, serious)
|
||||
- Different recording conditions
|
||||
|
||||
<Tip>
|
||||
<Callout type="info">
|
||||
The model will learn a more robust representation from diverse samples.
|
||||
</Tip>
|
||||
</Callout>
|
||||
|
||||
## Supported Languages
|
||||
|
||||
@@ -65,9 +65,9 @@ More languages coming soon.
|
||||
|
||||
## Limitations
|
||||
|
||||
<Warning>
|
||||
<Callout type="warn">
|
||||
Voice cloning should only be used with consent. Ensure you have permission to clone someone's voice.
|
||||
</Warning>
|
||||
</Callout>
|
||||
|
||||
- Quality depends on sample clarity
|
||||
- Works best with consistent speaking tone
|
||||
|
||||
Reference in New Issue
Block a user