Compare commits

...
34 Commits
Author SHA1 Message Date
James Pine 34e17bd469 Fix LuxTTS + Chatterbox in prod: bundle espeak/perth data, fix multiprocessing
- collect-all piper_phonemize to bundle espeak-ng-data for LuxTTS phonemization
- Set ESPEAK_DATA_PATH in frozen builds so the C library finds bundled data
- collect-all perth to bundle pretrained watermark model for Chatterbox
- Add multiprocessing.freeze_support() to fix resource_tracker subprocess crash
2026-03-15 16:02:09 -07:00
James Pine aada13a5c9 Collect all inflect files for PyInstaller (fixes typeguard inspect.getsource) 2026-03-15 14:32:10 -07:00
James Pine de8558d197 Fix prod build: download progress, robust stderr, full tracebacks
- Force tqdm disable=False in TrackedTqdm so byte progress works in prod
  (huggingface_hub disables tqdm based on logger level, which prevents
  self.n from updating — our progress tracking needs the counter even
  though we don't render to terminal)
- Harden devnull redirect to test writability, not just None check
- Add full traceback logging to all backend error handlers
- Add chatterbox/luxtts/zipvoice hidden imports and metadata to spec
2026-03-15 14:23:11 -07:00
Jamie Pine 9d79ea367a Only use --noconsole on Windows, macOS/Linux need stdout for Tauri logs 2026-03-15 12:07:35 -07:00
Jamie Pine 04316f7adc Copy metadata for requests/transformers/huggingface-hub to fix PyInstaller metadata lookup 2026-03-15 11:35:05 -07:00
Jamie Pine 4e4361d350 Fix noconsole crash: redirect None stdout/stderr to devnull on Windows 2026-03-15 11:27:35 -07:00
Jamie Pine e9a249587c Collect all linacodec files for PyInstaller (fixes inspect.getsource in Vocos) 2026-03-15 11:06:13 -07:00
Jamie Pine d8a9ed7d15 Enable updater artifacts with v1Compatible for tauri-action sig generation 2026-03-15 10:54:12 -07:00
Jamie Pine 3dbf1c200e Revert "Bump version: 0.2.3 → 0.2.4"
This reverts commit 40fcb8d917.
2026-03-15 10:20:31 -07:00
Jamie Pine 40fcb8d917 Bump version: 0.2.3 → 0.2.4 2026-03-15 10:18:51 -07:00
Jamie Pine ad64d1c3d9 Collect all zipvoice files for PyInstaller (fixes source code error) 2026-03-15 10:18:40 -07:00
Jamie Pine f826e45250 Install chatterbox-tts in CI release workflow 2026-03-15 10:17:23 -07:00
Jamie Pine 3d53c06c5b Bump version: 0.2.2 → 0.2.3 2026-03-15 10:08:56 -07:00
James Pine 9835b9f6d4 fix: prevent stale release data by removing Next.js fetch cache
Replace next: { revalidate: 600 } with cache: 'no-store' on GitHub
API fetches so new releases show up within 5 minutes (in-memory cache
only, no Next.js/Vercel cache layer on top).
2026-03-15 10:07:50 -07:00
Jamie Pine a15dd30b1e Update tauri-action to v0.6 to fix updater JSON and signature generation 2026-03-15 10:05:36 -07:00
Jamie Pine 1d343ac071 Treat missing/draft releases as up-to-date instead of showing error 2026-03-15 09:52:17 -07:00
James Pine ca602de0ae fix: don't reset audio player when unmuting during playback 2026-03-15 09:29:44 -07:00
James Pine cdc0293ca8 feat: add /linux-install page with build-from-source instructions
Linux download card now links to /linux-install instead of a direct
binary download. The page explains the CI situation and gives
clone + setup + build commands.
2026-03-15 09:17:30 -07:00
Jamie Pine e7f749f082 Add luxtts/zipvoice hidden imports to PyInstaller build 2026-03-15 09:13:59 -07:00
Jamie Pine d42e926e5c Bump version: 0.2.1 → 0.2.2 2026-03-15 09:02:10 -07:00
Jamie Pine 32768ea874 Add chatterbox hidden imports to PyInstaller build 2026-03-15 09:00:13 -07:00
James Pine b585e18ccf fix: fade in hero background glow to avoid Safari rendering flash 2026-03-15 08:53:08 -07:00
Jamie Pine 655910457f Auto-update CUDA binary on app update: check version on startup, download if stale 2026-03-15 08:46:17 -07:00
James Pine d6984f1057 fix: remove mix-blend-lighten and drop-shadow causing boxes in Safari 2026-03-15 08:45:40 -07:00
James Pine a637aebe69 feat: show version and total download count on landing page
Fetches download counts across all GitHub releases (paginated) and
displays version, total downloads, and platform list below the CTA.
2026-03-15 08:37:23 -07:00
James Pine a5269d23db Fix keep-server-running on macOS: ignore SIGHUP, watchdog grace period, build script fixes 2026-03-15 08:22:35 -07:00
Jamie Pine fc450e5024 Hide console window for server binary on Windows 2026-03-15 07:57:58 -07:00
Jamie Pine a99c2b572d Show download progress bar for CUDA backend download 2026-03-15 07:50:23 -07:00
Jamie Pine 96289e95f1 Bump version: 0.2.0 → 0.2.1 2026-03-15 06:36:38 -07:00
Jamie PineandGitHub e316b0b4bb Merge pull request #274 from jamiepine/feat/landing-page-redesign
Landing page v0.2.0 redesign
2026-03-15 06:22:04 -07:00
Jamie PineandGitHub 732270b571 Merge pull request #272 from jamiepine/windows-support
Windows support: CUDA detection, cross-platform justfile, clean server shutdown
2026-03-15 06:20:46 -07:00
James Pine 0c6aa15746 Responsive polish: pointer-events-none on animations, sticky header with scroll fade, desktop scroll-to-active fix, iOS audio unlock, player and UI tweaks
- Add pointer-events-none/select-none to feature cards, voice creator, and ControlUI mock
- Sticky header with gradient fade overlay (matching real app 3-layer technique)
- Fix desktop scroll-to-active: separate mobile/desktop card refs to prevent mobile refs overwriting desktop
- Scroll selected card to 2nd row when outside safe zone above generate box
- iOS Safari audio unlock via WaveSurfer's actual media element
- Player: accent fill play/pause button, padding on volume slider, remove close button
- Profile cards: fixed 143px height, mobile edge fades with scroll-aware left fade
- Generate box: accent effect pill when active, white fill sparkle icon, edge-aligned on desktop
- Voice creator: animated waveform background with height-based bars
- 12 profiles (added Attenborough, Zendaya, Obama) for 4-row grid with scroll
2026-03-15 06:17:54 -07:00
James Pine f80782a90a Landing page v0.2.0 updates: multi-engine copy, star count, model cards, voice creator section, responsive ControlUI, iOS audio fix
- Replace Qwen-specific copy with multi-engine messaging across hero, meta, and features
- Add GitHub star count fetched server-side via /api/stars with Spacedrive-style navbar badge
- Replace 'Why Voicebox exists' section with model cards for all 4 TTS engines
- Enable Linux download card (was 'Coming soon')
- Update GPU support copy to include ROCm, Intel Arc, DirectML
- Add Voice Creator section with animated 3-tab UI (upload, mic, system audio) and waveform background
- Make ControlUI responsive: horizontal scroll cards on mobile, stacked layout, scroll-to-active profile
- Fix iOS Safari audio autoplay (unlock AudioContext on user gesture)
- Fix hero logo square background with mix-blend-lighten
- Remove generation length green coloring, use gray with accent highlights
- Comment out grain overlay (visible tile seams)
- Remove player close button, stack waveform above controls on mobile
- Fixed-height profile cards (143px) with space between badges and buttons
2026-03-15 04:53:28 -07:00
Jamie Pine 8377152d86 Redesign landing page with animated ControlUI hero
New Spacedrive-inspired landing page with dark warm color system, glassmorphic navbar, feature cards with animated illustrations, and an interactive ControlUI mockup that cycles through voice generations with real audio playback via WaveSurfer.

The ControlUI demo script is fully data-driven - profiles, generation text, audio samples, and effects are all configurable from a single DEMO_SCRIPT array.

Includes 6 real voice samples (Jarvis, Morgan Freeman, Sam Altman, Samuel L. Jackson, Linus Tech Tips, Fireship) converted to webm opus.
2026-03-14 23:08:29 -07:00
53 changed files with 3934 additions and 520 deletions
+1 -1
View File
@@ -1,5 +1,5 @@
[bumpversion]
current_version = 0.2.0
current_version = 0.2.3
commit = True
tag = True
tag_name = v{new_version}
+3 -1
View File
@@ -61,6 +61,7 @@ jobs:
python -m pip install --upgrade pip
pip install pyinstaller
pip install -r backend/requirements.txt
pip install --no-deps chatterbox-tts
- name: Install MLX dependencies (Apple Silicon only)
if: matrix.backend == 'mlx'
@@ -122,7 +123,7 @@ jobs:
p12-file-base64: ${{ secrets.APPLE_CERTIFICATE }}
p12-password: ${{ secrets.APPLE_CERTIFICATE_PASSWORD }}
- uses: tauri-apps/tauri-action@v0
- uses: tauri-apps/tauri-action@v0.6
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TAURI_SIGNING_PRIVATE_KEY: ${{ secrets.TAURI_SIGNING_PRIVATE_KEY }}
@@ -173,6 +174,7 @@ jobs:
python -m pip install --upgrade pip
pip install pyinstaller
pip install -r backend/requirements.txt
pip install --no-deps chatterbox-tts
- name: Install PyTorch with CUDA 12.1
run: |
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@voicebox/app",
"version": "0.2.0",
"version": "0.2.3",
"private": true,
"type": "module",
"scripts": {
+31 -133
View File
@@ -139,7 +139,11 @@ export function AudioPlayer() {
barRadius: 2,
height: 80,
normalize: true,
backend: 'WebAudio',
// Use MediaElement backend (default). Unlike the WebAudio backend,
// MediaElement uses a standard <audio> element for playback which
// benefits from the browser/webview's built-in audio session recovery.
// This prevents audio loss when another app steals audio output or
// the system audio session is interrupted.
interact: true, // Enable interaction (click to seek)
mediaControls: false, // Don't show native controls
});
@@ -189,15 +193,6 @@ export function AudioPlayer() {
const currentVolume = usePlayerStore.getState().volume;
wavesurfer.setVolume(currentVolume);
// Get the underlying audio element and ensure it's not muted
// (unless we're using native playback, which will be set later)
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement && !isUsingNativePlaybackRef.current) {
mediaElement.volume = currentVolume;
mediaElement.muted = false;
debug.log('Audio element volume:', mediaElement.volume, 'muted:', mediaElement.muted);
}
// Auto-play when ready - check if we should use native playback
// Get current values from the store and queries at runtime (not captured closure values)
const currentAudioUrl = usePlayerStore.getState().audioUrl;
@@ -264,21 +259,8 @@ export function AudioPlayer() {
debug.log('Should use native playback:', shouldUseNative);
if (!shouldUseNative) {
debug.log('No custom devices assigned, falling back to WaveSurfer');
// Reset native playback flag and unmute WaveSurfer
debug.log('No custom devices assigned, using standard playback');
isUsingNativePlaybackRef.current = false;
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement) {
const currentVolume = usePlayerStore.getState().volume;
mediaElement.volume = currentVolume;
mediaElement.muted = false;
debug.log(
'WaveSurfer unmuted for normal playback - volume:',
mediaElement.volume,
'muted:',
mediaElement.muted,
);
}
} else {
const deviceIds = assignedChannels.flatMap((ch: any) => ch.device_ids);
debug.log('Device IDs to play to:', deviceIds);
@@ -299,19 +281,10 @@ export function AudioPlayer() {
// Mark that we're using native playback
isUsingNativePlaybackRef.current = true;
// Mute WaveSurfer's audio element to prevent UI audio output
// Keep WaveSurfer running for visualization
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement) {
mediaElement.volume = 0;
mediaElement.muted = true;
debug.log(
'WaveSurfer muted for native playback - volume:',
mediaElement.volume,
'muted:',
mediaElement.muted,
);
}
// Mute WaveSurfer's audio output — native handles the actual sound
// Keep WaveSurfer running for waveform visualization
wavesurfer.setVolume(0);
wavesurfer.setMuted(true);
// Start WaveSurfer playback for visualization (muted)
wavesurfer.play().catch((error) => {
@@ -334,38 +307,15 @@ export function AudioPlayer() {
'Native playback failed during auto-play, falling back to WaveSurfer:',
error,
);
// Reset native playback flag and unmute WaveSurfer
isUsingNativePlaybackRef.current = false;
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement) {
const currentVolume = usePlayerStore.getState().volume;
mediaElement.volume = currentVolume;
mediaElement.muted = false;
debug.log(
'WaveSurfer unmuted after native playback failure - volume:',
mediaElement.volume,
'muted:',
mediaElement.muted,
);
}
// Fall through to WaveSurfer playback
}
} else {
debug.log('Not using native playback, using WaveSurfer');
// Reset native playback flag and unmute WaveSurfer
isUsingNativePlaybackRef.current = false;
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement) {
const currentVolume = usePlayerStore.getState().volume;
mediaElement.volume = currentVolume;
mediaElement.muted = false;
debug.log(
'WaveSurfer unmuted for normal playback - volume:',
mediaElement.volume,
'muted:',
mediaElement.muted,
);
}
}
// Standard playback path — ensure WaveSurfer is unmuted
if (!isUsingNativePlaybackRef.current) {
wavesurfer.setMuted(false);
wavesurfer.setVolume(usePlayerStore.getState().volume);
}
// Only auto-play if shouldAutoPlay flag is set (user explicitly clicked to play)
@@ -389,28 +339,6 @@ export function AudioPlayer() {
// Handle play/pause
wavesurfer.on('play', () => {
setIsPlaying(true);
// Ensure audio element volume is set correctly
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement) {
// Double-check: if using native playback, keep WaveSurfer muted
// Otherwise, ensure it's unmuted
if (isUsingNativePlaybackRef.current) {
mediaElement.volume = 0;
mediaElement.muted = true;
debug.log('Playing (native mode) - WaveSurfer muted for visualization only');
} else {
// Ensure WaveSurfer is unmuted for normal playback
const currentVolume = usePlayerStore.getState().volume;
mediaElement.volume = currentVolume;
mediaElement.muted = false;
debug.log(
'Playing (normal mode) - volume:',
mediaElement.volume,
'muted:',
mediaElement.muted,
);
}
}
});
wavesurfer.on('pause', () => setIsPlaying(false));
wavesurfer.on('finish', () => {
@@ -492,11 +420,6 @@ export function AudioPlayer() {
if (wavesurferRef.current) {
debug.log('Destroying WaveSurfer instance');
try {
const mediaElement = wavesurferRef.current.getMediaElement();
if (mediaElement) {
mediaElement.pause();
mediaElement.src = '';
}
wavesurferRef.current.destroy();
} catch (error) {
debug.error('Error destroying WaveSurfer:', error);
@@ -537,13 +460,10 @@ export function AudioPlayer() {
}
// Reset native playback flag when loading new audio
// Also unmute WaveSurfer if it was muted
// Unmute WaveSurfer if it was muted for native playback
if (isUsingNativePlaybackRef.current) {
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement) {
mediaElement.muted = false;
mediaElement.volume = usePlayerStore.getState().volume;
}
wavesurfer.setMuted(false);
wavesurfer.setVolume(usePlayerStore.getState().volume);
}
isUsingNativePlaybackRef.current = false;
@@ -559,16 +479,7 @@ export function AudioPlayer() {
wavesurfer.pause();
}
// Stop the media element explicitly
const mediaElement = wavesurfer.getMediaElement();
if (mediaElement) {
debug.log('Stopping media element');
mediaElement.pause();
mediaElement.currentTime = 0;
mediaElement.src = '';
}
// Use empty() to completely destroy the waveform and media element
// Use empty() to completely destroy the waveform and reset media
debug.log('Calling wavesurfer.empty() to destroy audio');
wavesurfer.empty();
} catch (error) {
@@ -623,20 +534,13 @@ export function AudioPlayer() {
// Sync volume
useEffect(() => {
if (wavesurferRef.current) {
wavesurferRef.current.setVolume(volume);
// Also ensure the underlying audio element volume is set
const mediaElement = wavesurferRef.current.getMediaElement();
if (mediaElement) {
// If using native playback, keep WaveSurfer muted regardless of volume setting
if (isUsingNativePlaybackRef.current) {
mediaElement.volume = 0;
mediaElement.muted = true;
debug.log('Volume sync: Using native playback, keeping WaveSurfer muted');
} else {
mediaElement.volume = volume;
mediaElement.muted = volume === 0;
debug.log('Volume synced:', volume, 'muted:', mediaElement.muted);
}
// If using native playback, keep WaveSurfer muted regardless of volume setting
if (isUsingNativePlaybackRef.current) {
wavesurferRef.current.setVolume(0);
debug.log('Volume sync: Using native playback, keeping WaveSurfer muted');
} else {
wavesurferRef.current.setVolume(volume);
debug.log('Volume synced:', volume);
}
}
}, [volume]);
@@ -757,11 +661,8 @@ export function AudioPlayer() {
isUsingNativePlaybackRef.current = true;
// Mute WaveSurfer and start it for visualization
const mediaElement = wavesurferRef.current.getMediaElement();
if (mediaElement) {
mediaElement.volume = 0;
mediaElement.muted = true;
}
wavesurferRef.current.setVolume(0);
wavesurferRef.current.setMuted(true);
// Start WaveSurfer for visualization (muted)
wavesurferRef.current.play().catch((error) => {
@@ -785,11 +686,8 @@ export function AudioPlayer() {
} else {
// Ensure WaveSurfer is not muted if not using native playback
if (!isUsingNativePlaybackRef.current) {
const mediaElement = wavesurferRef.current.getMediaElement();
if (mediaElement) {
mediaElement.muted = false;
mediaElement.volume = volume;
}
wavesurferRef.current.setMuted(false);
wavesurferRef.current.setVolume(volume);
}
wavesurferRef.current.play().catch((error) => {
@@ -5,6 +5,14 @@ import { useEffect, useRef, useState } from 'react';
import { EffectsChainEditor } from '@/components/Effects/EffectsChainEditor';
import { GenerationPicker } from '@/components/Effects/GenerationPicker';
import { Button } from '@/components/ui/button';
import {
Dialog,
DialogContent,
DialogDescription,
DialogFooter,
DialogHeader,
DialogTitle,
} from '@/components/ui/dialog';
import { Input } from '@/components/ui/input';
import { Label } from '@/components/ui/label';
import { Separator } from '@/components/ui/separator';
@@ -29,6 +37,11 @@ export function EffectsDetail() {
const [saving, setSaving] = useState(false);
const [deleting, setDeleting] = useState(false);
// "Save as Custom" dialog state
const [saveAsDialogOpen, setSaveAsDialogOpen] = useState(false);
const [saveAsName, setSaveAsName] = useState('');
const [saveAsDescription, setSaveAsDescription] = useState('');
// Preview state
const [previewGenId, setPreviewGenId] = useState<string | null>(null);
const [previewLoading, setPreviewLoading] = useState(false);
@@ -165,8 +178,38 @@ export function EffectsDetail() {
}
}
async function handleSaveAsNew() {
await handleSaveNew();
function handleSaveAsNew() {
// Open the dialog with a suggested name based on the current preset
setSaveAsName(`${name} (Copy)`);
setSaveAsDescription(description);
setSaveAsDialogOpen(true);
}
async function handleSaveAsConfirm() {
if (!saveAsName.trim()) {
toast({ title: 'Name required', variant: 'destructive' });
return;
}
setSaving(true);
try {
const created = await apiClient.createEffectPreset({
name: saveAsName.trim(),
description: saveAsDescription.trim() || undefined,
effects_chain: workingChain,
});
queryClient.invalidateQueries({ queryKey: ['effect-presets'] });
setSaveAsDialogOpen(false);
setSelectedPresetId(created.id);
toast({ title: 'Preset saved', description: `"${created.name}" has been created.` });
} catch (error) {
toast({
title: 'Failed to save',
description: error instanceof Error ? error.message : 'Unknown error',
variant: 'destructive',
});
} finally {
setSaving(false);
}
}
async function handleDelete() {
@@ -327,6 +370,53 @@ export function EffectsDetail() {
</p>
</div>
</div>
{/* Save as Custom dialog */}
<Dialog open={saveAsDialogOpen} onOpenChange={setSaveAsDialogOpen}>
<DialogContent className="sm:max-w-md">
<DialogHeader>
<DialogTitle>Save as Custom Preset</DialogTitle>
<DialogDescription>
Create a new custom preset based on the current effects chain.
</DialogDescription>
</DialogHeader>
<div className="space-y-3 py-2">
<div className="space-y-1.5">
<Label className="text-xs">Name</Label>
<Input
value={saveAsName}
onChange={(e) => setSaveAsName(e.target.value)}
placeholder="My preset..."
className="h-9"
autoFocus
onKeyDown={(e) => {
if (e.key === 'Enter' && saveAsName.trim()) {
handleSaveAsConfirm();
}
}}
/>
</div>
<div className="space-y-1.5">
<Label className="text-xs">Description</Label>
<Textarea
value={saveAsDescription}
onChange={(e) => setSaveAsDescription(e.target.value)}
placeholder="Describe what this preset does..."
className="min-h-[60px] resize-none"
/>
</div>
</div>
<DialogFooter>
<Button variant="outline" onClick={() => setSaveAsDialogOpen(false)} disabled={saving}>
Cancel
</Button>
<Button onClick={handleSaveAsConfirm} disabled={saving || !saveAsName.trim()}>
<Save className="h-3.5 w-3.5 mr-1.5" />
{saving ? 'Saving...' : 'Save'}
</Button>
</DialogFooter>
</DialogContent>
</Dialog>
</div>
);
}
@@ -246,13 +246,18 @@ export function GpuAcceleration() {
{/* CUDA download section - only show when no GPU is active (native or CUDA) */}
{!hasNativeGpu && !isCurrentlyCuda && (
<>
{/* Download progress */}
{/* Download progress (manual download or auto-update) */}
{cudaDownloading && downloadProgress && (
<div className="space-y-2">
<div className="flex items-center justify-between text-sm">
<div className="flex items-center gap-2">
<Loader2 className="h-4 w-4 animate-spin" />
<span>{downloadProgress.filename || 'Downloading CUDA backend...'}</span>
<span>
{downloadProgress.filename ||
(cudaAvailable
? 'Updating CUDA backend...'
: 'Downloading CUDA backend...')}
</span>
</div>
{downloadProgress.total > 0 && (
<span className="text-muted-foreground">
+5 -2
View File
@@ -45,12 +45,15 @@ export function Sidebar({ isMacOS }: SidebarProps) {
{/* Navigation Buttons */}
<div className="flex flex-col gap-3">
{tabs.map((tab) => {
{tabs.map((tab, index) => {
const Icon = tab.icon;
// For index route, use exact match; for others, use default matching
const isActive =
tab.path === '/' ? matchRoute({ to: '/', exact: true }) : matchRoute({ to: tab.path });
// Accent fades as buttons get further from the logo
const accentOpacity = Math.max(0.08, 0.5 - index * 0.07);
return (
<Link
key={tab.id}
@@ -70,7 +73,7 @@ export function Sidebar({ isMacOS }: SidebarProps) {
style={{
maskImage: 'linear-gradient(to bottom, black, transparent 60%)',
WebkitMaskImage: 'linear-gradient(to bottom, black, transparent 60%)',
border: '1px solid hsl(var(--accent) / 0.5)',
border: `1px solid hsl(var(--accent) / ${accentOpacity})`,
}}
/>
)}
+1 -1
View File
@@ -1,3 +1,3 @@
# Backend package
__version__ = "0.2.0"
__version__ = "0.2.3"
+2 -1
View File
@@ -224,7 +224,8 @@ class ChatterboxTTSBackend:
task_manager.error_download(model_name, str(e))
raise
except Exception as e:
logger.error(f"Failed to load Chatterbox: {e}")
import traceback
logger.error(f"Failed to load Chatterbox: {e}\n{traceback.format_exc()}")
if not is_cached:
progress_manager.mark_error(model_name, str(e))
task_manager.error_download(model_name, str(e))
+2 -1
View File
@@ -228,7 +228,8 @@ class ChatterboxTurboTTSBackend:
task_manager.error_download(model_name, str(e))
raise
except Exception as e:
logger.error(f"Failed to load Chatterbox Turbo: {e}")
import traceback
logger.error(f"Failed to load Chatterbox Turbo: {e}\n{traceback.format_exc()}")
if not is_cached:
progress_manager.mark_error(model_name, str(e))
task_manager.error_download(model_name, str(e))
+2 -1
View File
@@ -149,7 +149,8 @@ class LuxTTSBackend:
logger.info("LuxTTS loaded successfully")
except Exception as e:
logger.error(f"Failed to load LuxTTS: {e}")
import traceback
logger.error(f"Failed to load LuxTTS: {e}\n{traceback.format_exc()}")
if not is_cached:
progress_manager.mark_error(model_name, str(e))
task_manager.error_download(model_name, str(e))
+31
View File
@@ -37,6 +37,11 @@ def build_server(cuda=False):
'--name', binary_name,
]
# Hide console window on Windows only. On macOS/Linux the sidecar needs
# stdout/stderr for Tauri to capture logs.
if platform.system() == "Windows":
args.append('--noconsole')
# Add local qwen_tts path if specified (for editable installs)
qwen_tts_path = os.getenv('QWEN_TTS_PATH')
if qwen_tts_path and Path(qwen_tts_path).exists():
@@ -67,6 +72,16 @@ def build_server(cuda=False):
'--hidden-import', 'backend.utils.effects',
'--hidden-import', 'backend.versions',
'--hidden-import', 'pedalboard',
'--hidden-import', 'chatterbox',
'--hidden-import', 'chatterbox.tts_turbo',
'--hidden-import', 'chatterbox.mtl_tts',
'--hidden-import', 'backend.backends.chatterbox_backend',
'--hidden-import', 'backend.backends.chatterbox_turbo_backend',
'--hidden-import', 'backend.backends.luxtts_backend',
'--hidden-import', 'zipvoice',
'--hidden-import', 'zipvoice.luxvoice',
'--collect-all', 'zipvoice',
'--collect-all', 'linacodec',
'--hidden-import', 'torch',
'--hidden-import', 'transformers',
'--hidden-import', 'fastapi',
@@ -81,11 +96,27 @@ def build_server(cuda=False):
'--hidden-import', 'qwen_tts.core',
'--hidden-import', 'qwen_tts.cli',
'--copy-metadata', 'qwen-tts',
'--copy-metadata', 'requests',
'--copy-metadata', 'transformers',
'--copy-metadata', 'huggingface-hub',
'--copy-metadata', 'tokenizers',
'--copy-metadata', 'safetensors',
'--copy-metadata', 'tqdm',
'--hidden-import', 'requests',
'--collect-submodules', 'qwen_tts',
'--collect-data', 'qwen_tts',
# Fix for pkg_resources and jaraco namespace packages
'--hidden-import', 'pkg_resources.extern',
'--collect-submodules', 'jaraco',
# inflect uses typeguard @typechecked which calls inspect.getsource()
# at import time — needs .py source files, not just .pyc bytecode
'--collect-all', 'inflect',
# perth ships pretrained watermark model files (hparams.yaml, .pth.tar)
# in perth/perth_net/pretrained/ — needed by chatterbox at runtime
'--collect-all', 'perth',
# piper_phonemize ships espeak-ng-data/ (phoneme tables, language dicts)
# needed by LuxTTS for text-to-phoneme conversion
'--collect-all', 'piper_phonemize',
])
# Add CUDA-specific hidden imports
+63 -2
View File
@@ -129,6 +129,17 @@ async def download_cuda_binary(version: Optional[str] = None):
except Exception as e:
logger.warning(f"Could not fetch checksum file — skipping verification: {e}")
# Get total size across all parts by issuing HEAD requests
total_size = 0
for part_name in parts:
try:
head_resp = await client.head(f"{base_url}/{part_name}")
content_length = int(head_resp.headers.get("content-length", 0))
total_size += content_length
except Exception:
pass
logger.info(f"Total download size: {total_size / 1024 / 1024:.1f} MB")
# Download and concatenate parts
total_downloaded = 0
with open(temp_path, "wb") as f:
@@ -142,8 +153,8 @@ async def download_cuda_binary(version: Optional[str] = None):
f.write(chunk)
total_downloaded += len(chunk)
progress.update_progress(
PROGRESS_KEY, current=total_downloaded, total=0,
filename=f"Part {i + 1}/{len(parts)}",
PROGRESS_KEY, current=total_downloaded, total=total_size,
filename=f"Downloading CUDA backend ({i + 1}/{len(parts)})",
status="downloading",
)
@@ -188,6 +199,56 @@ async def download_cuda_binary(version: Optional[str] = None):
raise
def get_cuda_binary_version() -> Optional[str]:
"""Get the version of the installed CUDA binary, or None if not installed."""
import subprocess
cuda_path = get_cuda_binary_path()
if not cuda_path:
return None
try:
result = subprocess.run(
[str(cuda_path), "--version"],
capture_output=True, text=True, timeout=30,
)
# Output format: "voicebox-server 0.2.0"
for line in result.stdout.strip().splitlines():
if "voicebox-server" in line:
return line.split()[-1]
except Exception as e:
logger.warning(f"Could not get CUDA binary version: {e}")
return None
async def check_and_update_cuda_binary():
"""Check if the CUDA binary is outdated and auto-download if so.
Called on server startup. If a CUDA binary exists but its version
doesn't match the current app version, triggers a background download
of the updated CUDA binary. The download progress is visible to the
frontend via the existing SSE progress endpoint.
"""
cuda_path = get_cuda_binary_path()
if not cuda_path:
return # No CUDA binary installed, nothing to update
cuda_version = get_cuda_binary_version()
current_version = __version__
if cuda_version == current_version:
logger.info(f"CUDA binary is up to date (v{current_version})")
return
logger.info(
f"CUDA binary version mismatch: binary=v{cuda_version}, app=v{current_version}. "
f"Auto-downloading updated CUDA backend..."
)
try:
await download_cuda_binary()
except Exception as e:
logger.error(f"Auto-update of CUDA binary failed: {e}")
async def delete_cuda_binary() -> bool:
"""Delete the downloaded CUDA binary. Returns True if deleted."""
path = get_cuda_binary_path()
+4
View File
@@ -3104,6 +3104,10 @@ async def startup_event():
print(f"Backend: {backend_type.upper()}")
print(f"GPU available: {_get_gpu_status()}")
# Auto-update CUDA binary if installed but outdated
from .cuda_download import check_and_update_cuda_binary
_create_background_task(check_and_update_cuda_binary())
# Initialize progress manager with main event loop for thread-safe operations
try:
progress_manager = get_progress_manager()
+49 -1
View File
@@ -6,6 +6,39 @@ absolute imports instead of relative imports.
"""
import sys
import os
# On Windows with --noconsole (PyInstaller), sys.stdout/stderr are None.
# They can also be broken file objects in some edge cases.
# Redirect to devnull to prevent crashes from print()/tqdm/logging.
def _is_writable(stream):
"""Check if a stream is usable for writing."""
if stream is None:
return False
try:
stream.write("")
return True
except Exception:
return False
if not _is_writable(sys.stdout):
sys.stdout = open(os.devnull, 'w')
if not _is_writable(sys.stderr):
sys.stderr = open(os.devnull, 'w')
# PyInstaller + multiprocessing: child processes re-execute the frozen binary
# with internal arguments. freeze_support() handles this and exits early.
import multiprocessing
multiprocessing.freeze_support()
# In frozen builds, piper_phonemize's espeak-ng C library falls back to
# /usr/share/espeak-ng-data/ which doesn't exist. Point it at the bundled
# data directory instead.
if getattr(sys, 'frozen', False):
_meipass = getattr(sys, '_MEIPASS', os.path.dirname(sys.executable))
_espeak_data = os.path.join(_meipass, 'piper_phonemize', 'espeak-ng-data')
if os.path.isdir(_espeak_data):
os.environ.setdefault('ESPEAK_DATA_PATH', _espeak_data)
# Fast path: handle --version before any heavy imports so the Rust
# version check doesn't block for 30+ seconds loading torch etc.
@@ -58,6 +91,12 @@ def disable_watchdog():
"""Disable the parent watchdog so the server keeps running after parent exits."""
global _watchdog_disabled
_watchdog_disabled = True
# Ignore SIGHUP so the server survives when the parent Tauri process exits.
# On Unix, child processes receive SIGHUP when the parent's session leader
# exits, which would kill the server even though we want it to persist.
if sys.platform != "win32":
import signal
signal.signal(signal.SIGHUP, signal.SIG_IGN)
def _start_parent_watchdog(parent_pid, data_dir=None):
@@ -130,7 +169,16 @@ def _start_parent_watchdog(parent_pid, data_dir=None):
watchdog_logger.info("Watchdog disabled (keep server running), stopping monitor")
return
if not _is_pid_alive(parent_pid):
watchdog_logger.info(f"Parent process {parent_pid} gone, shutting down server...")
# Parent is gone. Before shutting down, give the app a moment
# to send /watchdog/disable — there is a race where the Tauri
# RunEvent::Exit handler sends the disable request while we are
# mid-iteration (already past the _watchdog_disabled check above).
watchdog_logger.info(f"Parent process {parent_pid} gone, waiting for possible disable request...")
time.sleep(1)
if _watchdog_disabled:
watchdog_logger.info("Watchdog was disabled during grace period, keeping server alive")
return
watchdog_logger.info("Watchdog still enabled after grace period, shutting down server...")
if sys.platform == "win32":
# sys.exit triggers SystemExit, allowing uvicorn to run
# shutdown handlers. os.kill(SIGTERM) on Windows calls
+6
View File
@@ -64,11 +64,17 @@ class HFProgressTracker:
if key in tqdm_kwargs:
filtered_kwargs[key] = value
# Force-enable the progress bar — we're tracking progress ourselves,
# we don't need tqdm to render to a terminal, but we DO need
# self.n to be updated when update() is called.
filtered_kwargs['disable'] = False
# Try to initialize with filtered kwargs, fall back to all kwargs if that fails
try:
super().__init__(*args, **filtered_kwargs)
except TypeError:
# If filtering failed, try with all kwargs (maybe tqdm version accepts them)
kwargs['disable'] = False
super().__init__(*args, **kwargs)
self._tracker_filename = filename or "unknown"
+19 -10
View File
@@ -1,35 +1,44 @@
# -*- mode: python ; coding: utf-8 -*-
from PyInstaller.utils.hooks import collect_data_files
from PyInstaller.utils.hooks import collect_submodules
from PyInstaller.utils.hooks import collect_all
from PyInstaller.utils.hooks import copy_metadata
datas = []
hiddenimports = ['backend', 'backend.main', 'backend.config', 'backend.database', 'backend.models', 'backend.profiles', 'backend.history', 'backend.tts', 'backend.transcribe', 'backend.platform_detect', 'backend.backends', 'backend.backends.pytorch_backend', 'backend.utils.audio', 'backend.utils.cache', 'backend.utils.progress', 'backend.utils.hf_progress', 'backend.utils.validation', 'torch', 'transformers', 'fastapi', 'uvicorn', 'sqlalchemy', 'librosa', 'soundfile', 'qwen_tts', 'qwen_tts.inference', 'qwen_tts.inference.qwen3_tts_model', 'qwen_tts.inference.qwen3_tts_tokenizer', 'qwen_tts.core', 'qwen_tts.cli', 'pkg_resources.extern', 'backend.backends.mlx_backend', 'mlx', 'mlx.core', 'mlx.nn', 'mlx_audio', 'mlx_audio.tts', 'mlx_audio.stt']
binaries = []
hiddenimports = ['backend', 'backend.main', 'backend.config', 'backend.database', 'backend.models', 'backend.profiles', 'backend.history', 'backend.tts', 'backend.transcribe', 'backend.platform_detect', 'backend.backends', 'backend.backends.pytorch_backend', 'backend.utils.audio', 'backend.utils.cache', 'backend.utils.progress', 'backend.utils.hf_progress', 'backend.utils.validation', 'backend.cuda_download', 'backend.effects', 'backend.utils.effects', 'backend.versions', 'pedalboard', 'chatterbox', 'chatterbox.tts_turbo', 'chatterbox.mtl_tts', 'backend.backends.chatterbox_backend', 'backend.backends.chatterbox_turbo_backend', 'backend.backends.luxtts_backend', 'zipvoice', 'zipvoice.luxvoice', 'torch', 'transformers', 'fastapi', 'uvicorn', 'sqlalchemy', 'librosa', 'soundfile', 'qwen_tts', 'qwen_tts.inference', 'qwen_tts.inference.qwen3_tts_model', 'qwen_tts.inference.qwen3_tts_tokenizer', 'qwen_tts.core', 'qwen_tts.cli', 'requests', 'pkg_resources.extern', 'backend.backends.mlx_backend', 'mlx', 'mlx.core', 'mlx.nn', 'mlx_audio', 'mlx_audio.tts', 'mlx_audio.stt']
datas += collect_data_files('qwen_tts')
# Use collect_all (not collect_data_files) so native .dylib and .metallib
# files are bundled as binaries, not data. Without this, MLX raises OSError
# when loading Metal shaders inside the PyInstaller bundle.
from PyInstaller.utils.hooks import collect_all as _collect_all
_mlx_datas, _mlx_bins, _mlx_hidden = _collect_all('mlx')
_mlxa_datas, _mlxa_bins, _mlxa_hidden = _collect_all('mlx_audio')
datas += _mlx_datas + _mlxa_datas
datas += copy_metadata('qwen-tts')
datas += copy_metadata('requests')
datas += copy_metadata('transformers')
datas += copy_metadata('huggingface-hub')
datas += copy_metadata('tokenizers')
datas += copy_metadata('safetensors')
datas += copy_metadata('tqdm')
hiddenimports += collect_submodules('qwen_tts')
hiddenimports += collect_submodules('jaraco')
hiddenimports += collect_submodules('mlx')
hiddenimports += collect_submodules('mlx_audio')
tmp_ret = collect_all('zipvoice')
datas += tmp_ret[0]; binaries += tmp_ret[1]; hiddenimports += tmp_ret[2]
tmp_ret = collect_all('linacodec')
datas += tmp_ret[0]; binaries += tmp_ret[1]; hiddenimports += tmp_ret[2]
tmp_ret = collect_all('mlx')
datas += tmp_ret[0]; binaries += tmp_ret[1]; hiddenimports += tmp_ret[2]
tmp_ret = collect_all('mlx_audio')
datas += tmp_ret[0]; binaries += tmp_ret[1]; hiddenimports += tmp_ret[2]
a = Analysis(
['server.py'],
pathex=[],
binaries=_mlx_bins + _mlxa_bins,
binaries=binaries,
datas=datas,
hiddenimports=hiddenimports,
hookspath=[],
hooksconfig={},
runtime_hooks=[],
excludes=[],
excludes=['nvidia', 'nvidia.cublas', 'nvidia.cuda_cupti', 'nvidia.cuda_nvrtc', 'nvidia.cuda_runtime', 'nvidia.cudnn', 'nvidia.cufft', 'nvidia.curand', 'nvidia.cusolver', 'nvidia.cusparse', 'nvidia.nccl', 'nvidia.nvjitlink', 'nvidia.nvtx'],
noarchive=False,
optimize=0,
)
+17 -4
View File
@@ -17,7 +17,7 @@
},
"app": {
"name": "@voicebox/app",
"version": "0.1.13",
"version": "0.2.0",
"dependencies": {
"@dnd-kit/core": "^6.3.1",
"@dnd-kit/sortable": "^10.0.0",
@@ -72,13 +72,15 @@
},
"landing": {
"name": "@voicebox/landing",
"version": "0.1.13",
"version": "0.2.0",
"dependencies": {
"@fontsource/space-grotesk": "^5.2.10",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slot": "^1.2.4",
"autoprefixer": "^10.4.17",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"framer-motion": "^12.36.0",
"lucide-react": "^0.316.0",
"next": "^16.1.3",
"postcss": "^8.4.33",
@@ -87,6 +89,7 @@
"tailwind-merge": "^3.4.0",
"tailwindcss": "^3.4.1",
"tailwindcss-animate": "^1.0.7",
"wavesurfer.js": "^7.12.2",
},
"devDependencies": {
"@types/node": "^20.11.5",
@@ -97,7 +100,7 @@
},
"tauri": {
"name": "@voicebox/tauri",
"version": "0.1.13",
"version": "0.2.0",
"dependencies": {
"@tauri-apps/api": "^2.0.0",
"@tauri-apps/plugin-dialog": "^2.0.0",
@@ -120,7 +123,7 @@
},
"web": {
"name": "@voicebox/web",
"version": "0.1.13",
"version": "0.2.0",
"dependencies": {
"@tanstack/react-query": "^5.0.0",
"react": "^18.3.0",
@@ -274,6 +277,8 @@
"@floating-ui/utils": ["@floating-ui/[email protected]", "", {}, "sha512-aGTxbpbg8/b5JfU1HXSrbH3wXZuLPJcNEcZQFMxLs3oSzgtVu6nFPkbbGGUvBcUjKV2YyB9Wxxabo+HEH9tcRQ=="],
"@fontsource/space-grotesk": ["@fontsource/[email protected]", "", {}, "sha512-XNXEbT74OIITPqw2H6HXwPDp85fy43uxfBwFR5PU+9sLnjuLj12KlhVM9nZVN6q6dlKjkuN8JisW/OBxwxgUew=="],
"@hookform/resolvers": ["@hookform/[email protected]", "", { "peerDependencies": { "react-hook-form": "^7.0.0" } }, "sha512-79Dv+3mDF7i+2ajj7SkypSKHhl1cbln1OGavqrsF7p6mbUv11xpqpacPsGDCTRvCSjEEIez2ef1NveSVL3b0Ag=="],
"@humanwhocodes/config-array": ["@humanwhocodes/[email protected]", "", { "dependencies": { "@humanwhocodes/object-schema": "^2.0.3", "debug": "^4.3.1", "minimatch": "^3.0.5" } }, "sha512-DZLEEqFWQFiyK6h5YIeynKx7JlvCYWL0cImfSRXZ9l4Sg2efkFGTuFf6vzXjK1cq6IYkU+Eg/JizXw+TD2vRNw=="],
@@ -1152,12 +1157,16 @@
"@typescript-eslint/typescript-estree/semver": ["[email protected]", "", { "bin": { "semver": "bin/semver.js" } }, "sha512-SdsKMrI9TdgjdweUSR9MweHA4EJ8YxHn8DFaDisvhVlUOe4BF1tLD7GAj0lIqWVl+dPb/rExr0Btby5loQm20Q=="],
"@voicebox/landing/framer-motion": ["[email protected]", "", { "dependencies": { "motion-dom": "^12.36.0", "motion-utils": "^12.36.0", "tslib": "^2.4.0" }, "peerDependencies": { "@emotion/is-prop-valid": "*", "react": "^18.0.0 || ^19.0.0", "react-dom": "^18.0.0 || ^19.0.0" }, "optionalPeers": ["@emotion/is-prop-valid", "react", "react-dom"] }, "sha512-4PqYHAT7gev0ke0wos+PyrcFxI0HScjm3asgU8nSYa8YzJFuwgIvdj3/s3ZaxLq0bUSboIn19A2WS/MHwLCvfw=="],
"@voicebox/landing/lucide-react": ["[email protected]", "", { "peerDependencies": { "react": "^16.5.1 || ^17.0.0 || ^18.0.0" } }, "sha512-dTmYX1H4IXsRfVcj/KUxworV6814ApTl7iXaS21AimK2RUEl4j4AfOmqD3VR8phe5V91m4vEJ8tCK4uT1jE5nA=="],
"@voicebox/landing/tailwind-merge": ["[email protected]", "", {}, "sha512-uSaO4gnW+b3Y2aWoWfFpX62vn2sR3skfhbjsEnaBI81WD1wBLlHZe5sWf0AqjksNdYTbGBEd0UasQMT3SNV15g=="],
"@voicebox/landing/tailwindcss": ["[email protected]", "", { "dependencies": { "@alloc/quick-lru": "^5.2.0", "arg": "^5.0.2", "chokidar": "^3.6.0", "didyoumean": "^1.2.2", "dlv": "^1.1.3", "fast-glob": "^3.3.2", "glob-parent": "^6.0.2", "is-glob": "^4.0.3", "jiti": "^1.21.7", "lilconfig": "^3.1.3", "micromatch": "^4.0.8", "normalize-path": "^3.0.0", "object-hash": "^3.0.0", "picocolors": "^1.1.1", "postcss": "^8.4.47", "postcss-import": "^15.1.0", "postcss-js": "^4.0.1", "postcss-load-config": "^4.0.2 || ^5.0 || ^6.0", "postcss-nested": "^6.2.0", "postcss-selector-parser": "^6.1.2", "resolve": "^1.22.8", "sucrase": "^3.35.0" }, "bin": { "tailwind": "lib/cli.js", "tailwindcss": "lib/cli.js" } }, "sha512-3ofp+LL8E+pK/JuPLPggVAIaEuhvIz4qNcf3nA1Xn2o/7fb7s/TYpHhwGDv1ZU3PkBluUVaF8PyCHcm48cKLWQ=="],
"@voicebox/landing/wavesurfer.js": ["[email protected]", "", {}, "sha512-akVYISAHCw2gNw/7n8Pk/zH1Zz91WJyL/2MaNQCLD1XV3A226gKlWoDHWp9UdWqQ3zXnWttDf9ewZQQ3cxbOmQ=="],
"chokidar/glob-parent": ["[email protected]", "", { "dependencies": { "is-glob": "^4.0.1" } }, "sha512-AOIgSQCepiJYwP3ARnGx+5VnTu2HBYdzbGP45eLw1vr3zB3vZLeyed1sC9hnbcOc9/SrMyM5RPQrkGz4aS9Zow=="],
"fast-glob/glob-parent": ["[email protected]", "", { "dependencies": { "is-glob": "^4.0.1" } }, "sha512-AOIgSQCepiJYwP3ARnGx+5VnTu2HBYdzbGP45eLw1vr3zB3vZLeyed1sC9hnbcOc9/SrMyM5RPQrkGz4aS9Zow=="],
@@ -1169,5 +1178,9 @@
"tinyglobby/picomatch": ["[email protected]", "", {}, "sha512-5gTmgEY/sqK6gFXLIsQNH19lWb4ebPDLA4SdLP7dsWkIXHWlG66oPuVvXSGFPppYZz8ZDZq0dYYrbHfBCVUb1Q=="],
"@typescript-eslint/typescript-estree/minimatch/brace-expansion": ["[email protected]", "", { "dependencies": { "balanced-match": "^1.0.0" } }, "sha512-Jt0vHyM+jmUBqojB7E1NIYadt0vI0Qxjxd2TErW94wDz+E2LAm5vKMXXwg6ZZBTHPuUlDgQHKXvjGBdfcF1ZDQ=="],
"@voicebox/landing/framer-motion/motion-dom": ["[email protected]", "", { "dependencies": { "motion-utils": "^12.36.0" } }, "sha512-Ep1pq8P88rGJ75om8lTCA13zqd7ywPGwCqwuWwin6BKc0hMLkVfcS6qKlRqEo2+t0DwoUcgGJfXwaiFn4AOcQA=="],
"@voicebox/landing/framer-motion/motion-utils": ["[email protected]", "", {}, "sha512-eHWisygbiwVvf6PZ1vhaHCLamvkSbPIeAYxWUuL3a2PD/TROgE7FvfHWTIH4vMl798QLfMw15nRqIaRDXTlYRg=="],
}
}
+163
View File
@@ -0,0 +1,163 @@
# Voicebox v0.2.0 -- Release Notes
## The story
Voicebox v0.1.x shipped as a single-engine voice cloning app built around Qwen3-TTS. It worked, but it was limited: one model family, 10 languages, English-centric emotion, a synchronous generation pipeline that locked the UI, and a hard ceiling on how much text you could generate at once.
v0.2.0 is a ground-up rethink. Voicebox is now a **multi-engine voice cloning platform**. Four TTS engines. 23 languages. Expressive paralinguistic controls. A full post-processing effects pipeline. Unlimited generation length. Asynchronous everything. And it runs on every major GPU vendor -- NVIDIA, AMD, Intel Arc, Apple Silicon -- plus Docker for headless deployment.
This is the release where Voicebox stops being a proof of concept and starts being a real tool.
---
## Major New Features
### Multi-Engine Architecture
Voicebox now supports **four TTS engines**, each with different strengths. Switch between them per-generation from a single unified interface:
| Engine | Languages | Strengths |
|--------|-----------|-----------|
| **Qwen3-TTS** (0.6B / 1.7B) | 10 | High-quality multilingual cloning, delivery instructions ("speak slowly", "whisper") |
| **LuxTTS** | English | Lightweight (~1GB VRAM), 48kHz output, 150x realtime on CPU |
| **Chatterbox Multilingual** | 23 | Broadest language coverage -- Arabic, Danish, Finnish, Greek, Hebrew, Hindi, Malay, Norwegian, Polish, Swahili, Swedish, Turkish and more |
| **Chatterbox Turbo** | English | Fast 350M model with paralinguistic emotion/sound tags |
### Emotions and Paralinguistic Tags (Chatterbox Turbo)
Type `/` in the text input to open an autocomplete for **9 expressive tags** that the model synthesizes inline with speech:
`[laugh]` `[chuckle]` `[gasp]` `[cough]` `[sigh]` `[groan]` `[sniff]` `[shush]` `[clear throat]`
Tags render as inline badges in a rich text editor and serialize cleanly to the API. This makes generated speech sound natural and expressive in a way that plain TTS can't.
### 23 Languages via Chatterbox Multilingual
The Chatterbox Multilingual engine brings zero-shot voice cloning to **23 languages**: Arabic, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, and Turkish. The language dropdown dynamically filters to show only languages supported by the selected engine.
### Unlimited Generation Length (Auto-Chunking)
Previously, long text would hit model context limits and degrade. Now, text is **automatically split at sentence boundaries** and each chunk is generated independently, then crossfaded back together. This is fully engine-agnostic and works with all four engines.
- **Auto-chunking limit slider** (100-5,000 chars, default 800) -- controls when text gets split
- **Crossfade slider** (0-200ms, default 50ms) -- blends chunk boundaries smoothly, or set to 0 for a hard cut
- **Max text length raised to 50,000 characters** -- generate entire scripts, chapters, or articles in one go
- Smart splitting respects abbreviations (Dr., e.g., a.m.), CJK punctuation, and never breaks inside paralinguistic `[tags]`
### Asynchronous Generation Queue
Generation is now fully **non-blocking**. Submit a generation and immediately start typing the next one -- no more frozen UI waiting for inference to complete.
- Serial execution queue prevents GPU contention across all backends
- Real-time SSE status streaming (`generating` -> `completed` / `failed`)
- Failed generations can be retried without re-entering text
- Stale generations from crashes are auto-recovered on startup
- Generating status pill shown inline in the story editor
### Post-Processing Effects Pipeline
A full audio effects system powered by Spotify's `pedalboard` library. Apply effects after generation, preview them in real time, and build reusable presets -- all without leaving the app.
**8 effects available:**
| Effect | What it does |
|--------|-------------|
| **Pitch Shift** | Shift pitch up or down by up to 12 semitones |
| **Reverb** | Room reverb with configurable size, damping, and wet/dry mix |
| **Delay** | Echo with adjustable delay time, feedback, and mix |
| **Chorus / Flanger** | Modulated delay -- short for metallic flanger, longer for lush chorus |
| **Compressor** | Dynamic range compression with threshold, ratio, attack, and release |
| **Gain** | Volume adjustment from -40 to +40 dB |
| **High-Pass Filter** | Remove low frequencies below a configurable cutoff |
| **Low-Pass Filter** | Remove high frequencies above a configurable cutoff |
**Effects presets** -- Four built-in presets ship out of the box (Robotic, Radio, Echo Chamber, Deep Voice), and you can create unlimited custom presets. Presets are drag-and-drop chains of effects with per-parameter sliders.
**Per-profile default effects** -- Assign an effects chain to a voice profile and it applies automatically to every generation with that voice. Override per-generation from the generate box.
**Live preview** -- Audition any effects chain against an existing generation before committing. The preview streams processed audio without saving anything.
### Generation Versions
Every generation now supports **multiple versions** with full provenance tracking:
- **Original** -- the clean, unprocessed TTS output (always preserved)
- **Effects versions** -- apply different effects chains to create new versions from any source version
- **Takes** -- regenerate with the same text and voice but a new seed for variation
- **Source tracking** -- each version records which version it was derived from
- **Version pinning in stories** -- pin a specific version to a track clip in the story editor, independent of the generation's default
- **Favorites** -- star generations to mark them for quick access
---
## New Platform Support
### Linux (Native)
Full Linux support with `.deb` and `.rpm` packages. Includes PulseAudio/PipeWire audio capture for voice sample recording.
### AMD ROCm GPU Acceleration
AMD GPU users now get hardware-accelerated inference via ROCm, with automatic `HSA_OVERRIDE_GFX_VERSION` configuration for GPUs not officially in the ROCm compatibility list (e.g., RX 6600).
### NVIDIA CUDA Backend Swap
The CPU-only release can download and swap in a CUDA-accelerated backend binary from within the app -- no reinstall required. Handles GitHub's 2GB asset limit by downloading split parts and verifying SHA-256 checksums.
### Intel Arc (XPU) and DirectML
PyTorch backend also supports Intel Arc GPUs via IPEX/XPU and Windows any-GPU via DirectML.
### Docker + Web Deployment
Run Voicebox headless as a Docker container with the full web UI:
```bash
docker compose up
```
3-stage build, non-root runtime, health checks, persistent model cache across rebuilds. Binds to localhost only by default.
---
## Model Management
- **Per-model unload** -- free GPU memory without deleting downloaded models
- **Custom models directory** -- set `VOICEBOX_MODELS_DIR` to store models anywhere
- **Model folder migration** -- move all models to a new location with progress tracking
- **Whisper Turbo** -- added `openai/whisper-large-v3-turbo` as a transcription model option
- **Download cancel/clear UI** -- cancel in-progress downloads, VS Code-style problems panel for errors
---
## Security
- **CORS hardening** -- replaced wildcard `*` with an explicit allowlist of local origins; extensible via `VOICEBOX_CORS_ORIGINS` env var
- **Network access toggle** -- fully disable outbound network requests for air-gapped deployments
## Accessibility
- Comprehensive screen reader support (tested with NVDA/Narrator) across all major UI surfaces
- Keyboard navigation for voice cards, history rows, model management, and story editor
- State-aware `aria-label` attributes on all interactive controls
## Reliability
- **Atomic audio saves** -- two-phase write prevents corrupted files on crash/interrupt
- **Filesystem health endpoint** -- proactive disk space and directory writability checks
- **Errno-specific error messages** -- clear feedback for permission denied, disk full, missing directory
## UX Polish
- Responsive layout with horizontal-scroll voice cards on mobile
- App version shown in sidebar
- Voice card heights normalized
- Audio player title hidden at narrow widths to prevent overflow
---
## Installation
| Platform | Download |
|----------|----------|
| **macOS (Apple Silicon)** | `Voicebox_0.2.0_aarch64.dmg` |
| **macOS (Intel)** | `Voicebox_0.2.0_x64.dmg` |
| **Windows** | `Voicebox_0.2.0_x64_en-US.msi` or `x64-setup.exe` |
| **Linux** | `.deb` / `.rpm` packages |
| **Docker** | `docker compose up` |
The app includes automatic updates -- future patches will be installed automatically.
---
## Video Script Beats
For the marketing video, focus on these six beats:
1. **"Four engines, one app"** -- show the engine dropdown switching between Qwen, LuxTTS, Chatterbox, and Turbo
2. **"23 languages"** -- generate the same voice clone in Arabic, Japanese, Hindi, etc.
3. **"Make it expressive"** -- type `/laugh` and `/sigh` with Chatterbox Turbo, play back the result
4. **"Shape your sound"** -- apply the Robotic or Deep Voice preset, preview it live, then build a custom effects chain with drag-and-drop
5. **"No limits"** -- paste a long script, show it auto-chunk and generate seamlessly
6. **"Queue and go"** -- fire off multiple generations back-to-back without waiting
+5 -2
View File
@@ -1,6 +1,6 @@
{
"name": "@voicebox/landing",
"version": "0.2.0",
"version": "0.2.3",
"description": "Landing page for voicebox.sh",
"scripts": {
"dev": "bun --bun next dev --turbo",
@@ -9,11 +9,13 @@
"lint": "next lint"
},
"dependencies": {
"@fontsource/space-grotesk": "^5.2.10",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slot": "^1.2.4",
"autoprefixer": "^10.4.17",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"framer-motion": "^12.36.0",
"lucide-react": "^0.316.0",
"next": "^16.1.3",
"postcss": "^8.4.33",
@@ -21,7 +23,8 @@
"react-dom": "^18.2.0",
"tailwind-merge": "^3.4.0",
"tailwindcss": "^3.4.1",
"tailwindcss-animate": "^1.0.7"
"tailwindcss-animate": "^1.0.7",
"wavesurfer.js": "^7.12.2"
},
"devDependencies": {
"@types/node": "^20.11.5",
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 594 KiB

-1
View File
@@ -2,7 +2,6 @@ import { NextResponse } from 'next/server';
import { getLatestRelease } from '@/lib/releases';
export const dynamic = 'force-dynamic';
export const revalidate = 600; // Revalidate every 10 minutes
export async function GET() {
try {
+15
View File
@@ -0,0 +1,15 @@
import { NextResponse } from 'next/server';
import { getStarCount } from '@/lib/releases';
export const dynamic = 'force-dynamic';
export const revalidate = 600;
export async function GET() {
try {
const count = await getStarCount();
return NextResponse.json({ count });
} catch (error) {
console.error('Error fetching star count:', error);
return NextResponse.json({ error: 'Failed to fetch star count' }, { status: 500 });
}
}
+102 -23
View File
@@ -16,7 +16,7 @@
--secondary-foreground: 0 0% 0%;
--muted: 0 0% 96%;
--muted-foreground: 0 0% 45%;
--accent: 0 0% 96%;
--accent: 43 50% 50%;
--accent-foreground: 0 0% 0%;
--destructive: 0 0% 0%;
--destructive-foreground: 0 0% 100%;
@@ -27,26 +27,52 @@
}
.dark {
--background: 0 0% 3%;
--foreground: 0 0% 98%;
--card: 0 0% 8% / 0.6;
--card-foreground: 0 0% 98%;
--popover: 0 0% 8% / 0.8;
--popover-foreground: 0 0% 98%;
--primary: 0 0% 98%;
--primary-foreground: 0 0% 8%;
--secondary: 0 0% 12% / 0.5;
--secondary-foreground: 0 0% 98%;
--muted: 0 0% 12% / 0.4;
--muted-foreground: 0 0% 65%;
--accent: 0 0% 15% / 0.5;
--accent-foreground: 0 0% 98%;
/* Surfaces -- slightly warm-tinted darks */
--background: 30 4% 4%;
--foreground: 30 10% 94%;
--card: 30 4% 7%;
--card-foreground: 30 10% 94%;
--popover: 30 4% 7%;
--popover-foreground: 30 10% 94%;
--primary: 30 10% 94%;
--primary-foreground: 30 4% 7%;
--secondary: 30 4% 10%;
--secondary-foreground: 30 10% 94%;
--muted: 30 3% 12%;
--muted-foreground: 30 5% 55%;
--accent: 43 50% 45%;
--accent-foreground: 30 10% 94%;
--destructive: 0 62% 50%;
--destructive-foreground: 0 0% 98%;
--border: 0 0% 15% / 0.5;
--input: 0 0% 15% / 0.5;
--ring: 0 0% 98% / 0.2;
--radius: 1rem;
--destructive-foreground: 30 10% 94%;
--border: 30 4% 13%;
--input: 30 4% 13%;
--ring: 30 10% 94% / 0.2;
--radius: 0.75rem;
/* App-specific surface tokens */
--app: 30 4% 4%;
--app-box: 30 4% 7%;
--app-dark-box: 30 4% 5%;
--app-darker-box: 30 4% 3%;
--app-light-box: 30 4% 14%;
--app-line: 30 4% 13%;
--app-button: 30 4% 11%;
--app-hover: 30 4% 15%;
--app-selected: 30 4% 17%;
/* Text hierarchy */
--ink: 30 10% 94%;
--ink-dull: 30 5% 55%;
--ink-faint: 30 3% 38%;
/* Accent shades */
--accent-faint: 43 45% 55%;
--accent-deep: 43 55% 35%;
--accent-glow: 43 60% 50%;
/* Sidebar */
--sidebar: 30 4% 3%;
--sidebar-line: 30 4% 10%;
}
}
@@ -60,9 +86,9 @@
body {
@apply bg-background text-foreground antialiased;
overflow-x: hidden;
background-image:
radial-gradient(at 0% 0%, rgba(255, 255, 255, 0.03) 0px, transparent 50%),
radial-gradient(at 100% 100%, rgba(255, 255, 255, 0.02) 0px, transparent 50%);
font-family:
ui-sans-serif, system-ui, -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto,
"Helvetica Neue", Arial, sans-serif;
}
}
@@ -71,3 +97,56 @@
text-wrap: balance;
}
}
/* Staggered fade-in animation for hero elements */
@keyframes fadeUp {
from {
opacity: 0;
transform: translateY(16px);
}
to {
opacity: 1;
transform: translateY(0);
}
}
.fade-in {
opacity: 0;
animation: fadeUp 0.6s ease-out forwards;
}
.hero-glow-fade {
opacity: 0;
animation: fadeIn 2s ease-out 0.3s forwards;
}
@keyframes fadeIn {
from {
opacity: 0;
}
to {
opacity: 1;
}
}
/* Noise texture overlay for hero glow */
/* .hero-glow::after {
content: "";
position: absolute;
inset: 0;
z-index: 5;
pointer-events: none;
background: url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='2048' height='2048'%3E%3Cfilter id='n'%3E%3CfeTurbulence type='fractalNoise' baseFrequency='1.5' numOctaves='4' stitchTiles='stitch'/%3E%3C/filter%3E%3Crect width='100%25' height='100%25' filter='url(%23n)'/%3E%3C/svg%3E") center / 100% 100% no-repeat;
opacity: 0.35;
mix-blend-mode: overlay;
will-change: transform;
} */
/* Scrollbar hiding */
::-webkit-scrollbar {
display: none;
}
* {
scrollbar-width: none;
}
+23 -20
View File
@@ -1,17 +1,19 @@
import type { Metadata } from 'next';
import { Inter } from 'next/font/google';
import './globals.css';
import { Banner } from '@/components/Banner';
import { Footer } from '@/components/Footer';
import { Header } from '@/components/Header';
const inter = Inter({ subsets: ['latin'], variable: '--font-sans' });
export const metadata: Metadata = {
title: 'Voicebox - Open Source Voice Cloning Desktop App Powered by Qwen3-TTS',
title: 'Voicebox - Open Source Voice Cloning Desktop App',
description:
'Near-perfect voice cloning powered by Qwen3-TTS. Desktop app for Mac, Windows, and Linux. Multi-sample support, smart caching, local or remote inference.',
keywords: ['voice cloning', 'TTS', 'Qwen3', 'desktop app', 'AI voice'],
'Near-perfect voice cloning with multiple TTS engines. Desktop app for Mac, Windows, and Linux. Multi-sample support, smart caching, local or remote inference.',
keywords: [
'voice cloning',
'TTS',
'multi-engine',
'desktop app',
'AI voice',
'open source',
'text to speech',
],
icons: {
icon: [
{ url: '/favicon.png', type: 'image/png' },
@@ -20,8 +22,8 @@ export const metadata: Metadata = {
apple: [{ url: '/apple-touch-icon.png', sizes: '180x180', type: 'image/png' }],
},
openGraph: {
title: 'voicebox',
description: 'Professional voice cloning with Qwen3-TTS',
title: 'Voicebox',
description: 'Open source voice cloning. Local-first. Free forever.',
type: 'website',
url: 'https://voicebox.sh',
},
@@ -30,15 +32,16 @@ export const metadata: Metadata = {
export default function RootLayout({ children }: { children: React.ReactNode }) {
return (
<html lang="en" suppressHydrationWarning className="dark">
<body className={inter.variable}>
<div className="relative min-h-screen bg-background font-sans flex flex-col">
<Banner />
<Header />
<main className="container mx-auto px-4 sm:px-6 md:px-4 flex-1 py-4 sm:py-6 md:py-0">
{children}
</main>
<Footer />
</div>
<head>
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossOrigin="anonymous" />
<link
href="https://fonts.googleapis.com/css2?family=Caveat:wght@400;500&display=swap"
rel="stylesheet"
/>
</head>
<body>
<div className="relative min-h-screen bg-background font-sans">{children}</div>
</body>
</html>
);
+169
View File
@@ -0,0 +1,169 @@
import type { Metadata } from 'next';
import { Footer } from '@/components/Footer';
import { Navbar } from '@/components/Navbar';
import { GITHUB_REPO } from '@/lib/constants';
export const metadata: Metadata = {
title: 'Linux Install - Voicebox',
description: 'Build Voicebox from source on Linux. Clone, setup, and build in three commands.',
};
export default function LinuxInstall() {
return (
<>
<Navbar />
<section className="relative pt-32 pb-24">
<div className="mx-auto max-w-2xl px-6">
<h1 className="text-3xl font-bold tracking-tight text-foreground">Install on Linux</h1>
<p className="mt-4 text-muted-foreground">
We&apos;re currently working through CI issues that prevent us from shipping a reliable
pre-built binary for Linux. In the meantime, building from source is straightforward and
takes just a few minutes.
</p>
<div className="mt-10 space-y-6">
{/* Prerequisites */}
<div>
<h2 className="text-sm font-medium text-muted-foreground uppercase tracking-wider mb-3">
Prerequisites
</h2>
<ul className="list-disc list-inside text-sm text-muted-foreground space-y-1">
<li>
<a
href="https://git-scm.com"
target="_blank"
rel="noopener noreferrer"
className="text-foreground hover:underline"
>
Git
</a>
</li>
<li>
<a
href="https://www.rust-lang.org/tools/install"
target="_blank"
rel="noopener noreferrer"
className="text-foreground hover:underline"
>
Rust
</a>
</li>
<li>
<a
href="https://github.com/casey/just#installation"
target="_blank"
rel="noopener noreferrer"
className="text-foreground hover:underline"
>
just
</a>{' '}
— install via{' '}
<code className="text-xs bg-muted px-1.5 py-0.5 rounded">cargo install just</code>
</li>
<li>
<a
href="https://bun.sh"
target="_blank"
rel="noopener noreferrer"
className="text-foreground hover:underline"
>
Bun
</a>
</li>
<li>
Tauri system deps —{' '}
<a
href="https://v2.tauri.app/start/prerequisites/#linux"
target="_blank"
rel="noopener noreferrer"
className="text-foreground hover:underline"
>
see Tauri docs
</a>
</li>
</ul>
</div>
{/* Steps */}
<div>
<h2 className="text-sm font-medium text-muted-foreground uppercase tracking-wider mb-3">
Build from source
</h2>
<div className="space-y-3">
<div className="rounded-lg border border-border bg-card/60 p-4 font-mono text-sm">
<div className="text-muted-foreground select-none"># Clone the repo</div>
<div>git clone https://github.com/jamiepine/voicebox.git</div>
<div>cd voicebox</div>
</div>
<div className="rounded-lg border border-border bg-card/60 p-4 font-mono text-sm">
<div className="text-muted-foreground select-none">
# Install all dependencies (Python venv, JS deps, etc.)
</div>
<div>just setup</div>
</div>
<div className="rounded-lg border border-border bg-card/60 p-4 font-mono text-sm">
<div className="text-muted-foreground select-none"># Build the app</div>
<div>just build</div>
</div>
</div>
<p className="mt-4 text-sm text-muted-foreground">
The built app will be in{' '}
<code className="text-xs bg-muted px-1.5 py-0.5 rounded">
tauri/src-tauri/target/release/bundle/
</code>
</p>
</div>
{/* Dev mode */}
<div>
<h2 className="text-sm font-medium text-muted-foreground uppercase tracking-wider mb-3">
Or run in dev mode
</h2>
<div className="rounded-lg border border-border bg-card/60 p-4 font-mono text-sm">
<div className="text-muted-foreground select-none">
# Start the dev server with hot reload
</div>
<div>just dev</div>
</div>
</div>
</div>
{/* Links */}
<div className="mt-12 pt-8 border-t border-border flex flex-wrap gap-4 text-sm">
<a
href={GITHUB_REPO}
target="_blank"
rel="noopener noreferrer"
className="text-muted-foreground hover:text-foreground transition-colors"
>
GitHub Repo
</a>
<a
href={`${GITHUB_REPO}/issues`}
target="_blank"
rel="noopener noreferrer"
className="text-muted-foreground hover:text-foreground transition-colors"
>
Report an issue
</a>
<a
href={`${GITHUB_REPO}/blob/main/CONTRIBUTING.md`}
target="_blank"
rel="noopener noreferrer"
className="text-muted-foreground hover:text-foreground transition-colors"
>
Contributing guide
</a>
</div>
</div>
</section>
<Footer />
</>
);
}
+297 -274
View File
@@ -1,304 +1,327 @@
'use client';
import { Cloud, Code, Cpu, Github, Shield, Zap } from 'lucide-react';
import Image from 'next/image';
import { Github, Globe, Languages, MessageSquare, Zap } from 'lucide-react';
import { useEffect, useState } from 'react';
import { ControlUI } from '@/components/ControlUI';
import { Features } from '@/components/Features';
import { Footer } from '@/components/Footer';
import { Navbar } from '@/components/Navbar';
import { AppleIcon, LinuxIcon, WindowsIcon } from '@/components/PlatformIcons';
import { Button } from '@/components/ui/button';
import { Section } from '@/components/ui/section';
import { VoiceCreator } from '@/components/VoiceCreator';
import { DOWNLOAD_LINKS, GITHUB_REPO } from '@/lib/constants';
import type { DownloadLinks } from '@/lib/releases';
import { FeatureCard } from '../components/ui/feature-card';
export default function Home() {
const [downloadLinks, setDownloadLinks] = useState<DownloadLinks>(DOWNLOAD_LINKS);
const [version, setVersion] = useState<string | null>(null);
const [totalDownloads, setTotalDownloads] = useState<number | null>(null);
useEffect(() => {
// Fetch latest release info
fetch('/api/releases')
.then((res) => {
if (!res.ok) {
throw new Error('Failed to fetch releases');
}
if (!res.ok) throw new Error('Failed to fetch releases');
return res.json();
})
.then((data) => {
if (data.downloadLinks) {
setDownloadLinks(data.downloadLinks);
}
if (data.downloadLinks) setDownloadLinks(data.downloadLinks);
if (data.version) setVersion(data.version);
if (data.totalDownloads != null) setTotalDownloads(data.totalDownloads);
})
.catch((error) => {
console.error('Failed to fetch release info:', error);
// Keep fallback links (releases page) on error
});
}, []);
const features = [
{
title: 'Near-Perfect Voice Cloning',
description:
"Powered by Alibaba's Qwen3-TTS model for exceptional voice quality and accuracy.",
icon: <Zap className="h-6 w-6" />,
},
{
title: 'Stories Editor',
description:
'Create multi-voice narratives with a timeline-based editor. Arrange tracks, trim clips, and mix conversations.',
icon: <Code className="h-6 w-6" />,
},
{
title: 'Multi-Sample Support',
description:
'Combine multiple voice samples for higher quality and more natural-sounding results.',
icon: <Code className="h-6 w-6" />,
},
{
title: 'Local or Remote',
description:
'Run GPU inference locally or connect to a remote machine. One-click server setup.',
icon: <Cloud className="h-6 w-6" />,
},
{
title: 'Audio Transcription',
description:
'Powered by Whisper for accurate speech-to-text. Extract reference text from voice samples automatically.',
icon: <Shield className="h-6 w-6" />,
},
{
title: 'Cross-Platform',
description: 'Available for macOS, Windows, and Linux. No Python installation required.',
icon: <Cpu className="h-6 w-6" />,
},
];
return (
<div className="space-y-12 sm:space-y-16 md:space-y-20">
{/* Hero Section */}
<section className="relative py-12 sm:py-16 md:py-20 lg:py-24">
<div className="container mx-auto px-4 sm:px-6 lg:px-8 relative max-w-7xl">
<div className="grid grid-cols-1 lg:grid-cols-2 gap-8 lg:gap-12 items-start">
{/* Left side - Content */}
<div className="space-y-6 lg:pr-8">
<div className="flex lg:justify-start justify-center mb-6">
<Image
src="/voicebox-logo-2.png"
alt="Voicebox Logo"
width={1024}
height={1024}
className="w-32 sm:w-40 md:w-48 h-auto"
priority
/>
</div>
<h1 className="text-5xl sm:text-6xl md:text-7xl lg:text-8xl font-bold leading-tight text-center lg:text-left">
Voicebox
</h1>
<p className="text-lg sm:text-xl md:text-2xl text-foreground/70 max-w-xl text-center lg:text-left mx-auto lg:mx-0">
Open source voice cloning powered by Qwen3-TTS. Create natural-sounding speech from
text with near-perfect voice replication.
</p>
<>
<Navbar />
{/* Mobile: centered screenshot above download buttons */}
<div className="flex justify-center lg:hidden my-8">
<div className="w-full max-w-2xl">
<Image
src="/assets/app-screenshot-1.webp"
alt="Voicebox Application Screenshot"
width={1920}
height={1080}
className="w-full h-auto rounded-lg shadow-lg"
priority
/>
</div>
</div>
{/* Download buttons under left content */}
<div className="space-y-4 pt-4">
<div className="grid grid-cols-1 sm:grid-cols-2 gap-3 w-full max-w-2xl">
<Button asChild size="lg" className="w-full px-0">
<a
href={downloadLinks.macArm}
download
className="flex items-center w-full relative"
>
<div className="flex items-center gap-2 flex-shrink-0 pl-4">
<AppleIcon className="h-5 w-5" />
<div className="h-5 w-px bg-border" />
</div>
<span className="flex-1 text-center px-4">macOS (ARM)</span>
</a>
</Button>
<Button asChild size="lg" className="w-full px-0">
<a
href={downloadLinks.macIntel}
download
className="flex items-center w-full relative"
>
<div className="flex items-center gap-2 flex-shrink-0 pl-4">
<AppleIcon className="h-5 w-5" />
<div className="h-5 w-px bg-border" />
</div>
<span className="flex-1 text-center px-4">macOS (Intel)</span>
</a>
</Button>
<Button asChild size="lg" className="w-full px-0">
<a
href={downloadLinks.windows}
download
className="flex items-center w-full relative"
>
<div className="flex items-center gap-2 flex-shrink-0 pl-4">
<WindowsIcon className="h-5 w-5" />
<div className="h-5 w-px bg-border" />
</div>
<span className="flex-1 text-center px-4">Windows</span>
</a>
</Button>
<Button asChild size="lg" className="w-full px-0" disabled>
<a
href={downloadLinks.linux}
onClick={(e) => e.preventDefault()}
className="flex items-center w-full relative opacity-50 cursor-not-allowed"
title="Linux builds coming soon — Currently blocked by GitHub runner disk space limitations."
aria-label="Linux builds coming soon — Currently blocked by GitHub runner disk space limitations."
>
<div className="flex items-center gap-2 flex-shrink-0 pl-4">
<LinuxIcon className="h-5 w-5" />
<div className="h-5 w-px bg-border" />
</div>
<span className="flex-1 text-center px-4">Linux</span>
</a>
</Button>
</div>
<Button variant="outline" size="lg" asChild className="w-full max-w-2xl">
<a href={GITHUB_REPO} target="_blank" rel="noopener noreferrer">
<Github className="h-4 w-4 mr-2" />
View on GitHub
</a>
</Button>
</div>
</div>
{/* Desktop: Large screenshot positioned off-screen */}
<div className="hidden lg:block relative">
<div className="absolute right-0 top-0 -mt-10 w-[200%] -mr-[100%]">
<Image
src="/assets/app-screenshot-1.webp"
alt="Voicebox Application Screenshot"
width={1920}
height={1080}
className="w-full h-auto"
priority
/>
</div>
</div>
</div>
{/* ── Hero Section ─────────────────────────────────────────────── */}
<section className="relative pt-32 pb-16">
{/* Background glow */}
<div className="hero-glow hero-glow-fade pointer-events-none absolute inset-0 -top-32">
<div className="absolute left-1/2 top-0 -translate-x-1/2 w-[800px] h-[600px] rounded-full bg-accent/15 blur-[150px]" />
<div className="absolute left-1/2 top-12 -translate-x-1/2 w-[500px] h-[400px] rounded-full bg-accent/10 blur-[80px]" />
</div>
</section>
{/* Screenshots Section */}
<section className="py-12 sm:py-16 md:py-20">
<div className="w-full md:w-[150%] md:-ml-[25%]">
<div className="grid grid-cols-1 md:grid-cols-3 gap-6 px-8">
<div className="w-full">
<Image
src="/assets/app-screenshot-2.webp"
alt="Voicebox Screenshot 2"
width={1920}
height={1080}
className="w-full h-auto rounded-lg shadow-lg"
/>
</div>
<div className="w-full">
<Image
src="/assets/app-screenshot-1.webp"
alt="Voicebox Screenshot 1"
width={1920}
height={1080}
className="w-full h-auto rounded-lg shadow-lg"
/>
</div>
<div className="w-full">
<Image
src="/assets/app-screenshot-3.webp"
alt="Voicebox Screenshot 3"
width={1920}
height={1080}
className="w-full h-auto rounded-lg shadow-lg"
/>
</div>
</div>
</div>
</section>
{/* Description Section */}
<section className="py-12 sm:py-16 md:py-20">
<div className="container mx-auto px-4 sm:px-6 lg:px-8 max-w-4xl">
<h2 className="text-3xl sm:text-4xl md:text-5xl font-bold text-center mb-6 sm:mb-8">
What is Voicebox?
</h2>
<div className="space-y-6 text-lg text-foreground/80 text-center">
<p>
Voicebox is a <strong>local-first voice cloning studio</strong> with DAW-like features
for professional voice synthesis. Think of it as a{' '}
<strong>local, free and open-source alternative to ElevenLabs</strong> — download
models, clone voices, and generate speech entirely on your machine.
</p>
<p>
Unlike cloud services that lock your voice data behind subscriptions, Voicebox gives
you complete privacy, professional tools, and native performance. Download a voice
model, clone any voice from a few seconds of audio, and compose multi-voice projects
with studio-grade editing tools.
</p>
<p>
Optimized for performance with <strong>Metal acceleration on Mac</strong> and{' '}
<strong>CUDA acceleration on Windows/Linux</strong> for fast, local inference.
</p>
<p className="text-foreground/60">No Python install required.</p>
</div>
</div>
</section>
{/* Demo Video Section */}
<section className="py-12 sm:py-16 md:py-20">
<div className="container mx-auto px-4 sm:px-6 lg:px-8 max-w-7xl">
<h2 className="text-3xl sm:text-4xl md:text-5xl font-bold text-center mb-8 sm:mb-12">
See it in action...
</h2>
<div className="flex justify-center">
<div className="w-full max-w-5xl">
{/** biome-ignore lint/a11y/useMediaCaption: not generating captions for this, ya damn linter */}
<video
className="w-full h-auto rounded-lg shadow-lg"
controls
playsInline
preload="metadata"
poster="/assets/app-screenshot-1.webp"
>
<source
src="/voicebox-demo.webm"
type="video/webm"
aria-label="Voicebox Demo Video"
/>
Your browser does not support the video tag.
</video>
</div>
</div>
</div>
</section>
{/* Features Section */}
<Section id="features">
<div className="grid grid-cols-1 md:grid-cols-2 lg:grid-cols-3 gap-4 sm:gap-6">
{features.map((feature) => (
<FeatureCard
key={feature.title}
title={feature.title}
description={feature.description}
icon={feature.icon}
<div className="relative mx-auto max-w-7xl px-6 text-center">
{/* Logo */}
<div
className="fade-in mx-auto mb-8 h-[120px] w-[120px] md:h-[160px] md:w-[160px]"
style={{ animationDelay: '0ms' }}
>
{/* eslint-disable-next-line @next/next/no-img-element */}
<img
src="/voicebox-logo-app.webp"
alt="Voicebox"
className="h-full w-full object-contain"
/>
))}
</div>
{/* Headline */}
<div className="fade-in relative" style={{ animationDelay: '100ms' }}>
<h1 className="text-5xl font-bold tracking-tighter leading-[0.9] text-foreground md:text-7xl lg:text-8xl">
Your voice, your machine.
</h1>
</div>
{/* Subtitle */}
<p
className="fade-in mx-auto mt-6 max-w-2xl text-lg text-muted-foreground md:text-xl"
style={{ animationDelay: '200ms' }}
>
Open source voice cloning studio with support for multiple TTS engines. Clone any voice,
generate natural speech, and compose multi-voice projects — all running locally.
</p>
{/* CTAs */}
<div
className="fade-in mt-10 flex flex-col sm:flex-row items-center justify-center gap-4"
style={{ animationDelay: '300ms' }}
>
<a
href="#download"
className="rounded-full bg-accent px-8 py-3.5 text-sm font-semibold uppercase tracking-wider text-white shadow-[0_4px_20px_hsl(43_60%_50%/0.3),inset_0_2px_0_rgba(255,255,255,0.2),inset_0_-2px_0_rgba(0,0,0,0.1)] transition-all hover:bg-accent-faint active:shadow-[0_2px_10px_hsl(43_60%_50%/0.3),inset_0_4px_8px_rgba(0,0,0,0.3)]"
>
Download
</a>
<a
href={GITHUB_REPO}
target="_blank"
rel="noopener noreferrer"
className="flex items-center gap-2 rounded-full border border-border/60 bg-card/40 backdrop-blur-sm px-6 py-3 text-sm font-medium text-muted-foreground transition-colors hover:text-foreground hover:border-border"
>
<Github className="h-4 w-4" />
View on GitHub
</a>
</div>
{/* Version + downloads */}
<p
className="fade-in mt-4 text-xs text-muted-foreground/50"
style={{ animationDelay: '400ms' }}
>
{version ?? ''}
{version && totalDownloads != null ? ' \u00b7 ' : ''}
{totalDownloads != null ? `${totalDownloads.toLocaleString()} downloads` : ''}
{version || totalDownloads != null ? ' \u00b7 ' : ''}
macOS, Windows, Linux
</p>
</div>
</Section>
</div>
{/* ── ControlUI mockup ─────────────────────────────────────── */}
<div className="mt-16">
<ControlUI />
</div>
</section>
{/* ── Features ─────────────────────────────────────────────── */}
<Features />
{/* ── Voice Creator ────────────────────────────────────────── */}
<VoiceCreator />
{/* ── Models ─────────────────────────────────────────────────── */}
<section id="about" className="border-t border-border py-24">
<div className="mx-auto max-w-5xl px-6">
<div className="text-center mb-14">
<h2 className="text-3xl font-semibold tracking-tight text-foreground md:text-4xl mb-4">
Multi-Engine Architecture
</h2>
<p className="text-muted-foreground max-w-2xl mx-auto">
Choose the right model for every job. All models run locally on your hardware —
download once, use forever.
</p>
</div>
<div className="grid grid-cols-1 md:grid-cols-2 gap-4">
{/* Qwen3-TTS */}
<div className="rounded-xl border border-border bg-card/60 backdrop-blur-sm p-6 transition-colors hover:border-accent/30">
<div className="flex items-start justify-between mb-3">
<div>
<h3 className="text-base font-semibold text-foreground">Qwen3-TTS</h3>
<span className="text-xs text-muted-foreground/60">by Alibaba</span>
</div>
<div className="flex gap-1.5">
<span className="text-[10px] px-2 py-0.5 rounded-full border border-border bg-background text-muted-foreground">
1.7B
</span>
<span className="text-[10px] px-2 py-0.5 rounded-full border border-border bg-background text-muted-foreground">
0.6B
</span>
</div>
</div>
<p className="text-sm text-muted-foreground leading-relaxed mb-4">
High-quality multilingual voice cloning with natural prosody. The only engine with
delivery instructions — control tone, pace, and emotion with natural language.
</p>
<div className="flex flex-wrap gap-2">
<span className="flex items-center gap-1 text-[11px] text-muted-foreground/70">
<Globe className="h-3 w-3" />
10 languages
</span>
<span className="flex items-center gap-1 text-[11px] text-muted-foreground/70">
<MessageSquare className="h-3 w-3" />
Delivery instructions
</span>
</div>
</div>
{/* Chatterbox */}
<div className="rounded-xl border border-border bg-card/60 backdrop-blur-sm p-6 transition-colors hover:border-accent/30">
<div className="flex items-start justify-between mb-3">
<div>
<h3 className="text-base font-semibold text-foreground">Chatterbox</h3>
<span className="text-xs text-muted-foreground/60">by Resemble AI</span>
</div>
</div>
<p className="text-sm text-muted-foreground leading-relaxed mb-4">
Production-grade voice cloning with the broadest language support. 23 languages with
zero-shot cloning and emotion exaggeration control.
</p>
<div className="flex flex-wrap gap-2">
<span className="flex items-center gap-1 text-[11px] text-muted-foreground/70">
<Languages className="h-3 w-3" />
23 languages
</span>
</div>
</div>
{/* Chatterbox Turbo */}
<div className="rounded-xl border border-border bg-card/60 backdrop-blur-sm p-6 transition-colors hover:border-accent/30">
<div className="flex items-start justify-between mb-3">
<div>
<h3 className="text-base font-semibold text-foreground">Chatterbox Turbo</h3>
<span className="text-xs text-muted-foreground/60">by Resemble AI</span>
</div>
<span className="text-[10px] px-2 py-0.5 rounded-full border border-border bg-background text-muted-foreground">
350M
</span>
</div>
<p className="text-sm text-muted-foreground leading-relaxed mb-4">
Lightweight and fast. Supports paralinguistic tags — embed [laugh], [sigh], [gasp]
and more directly in your text for expressive, natural speech.
</p>
<div className="flex flex-wrap gap-2">
<span className="flex items-center gap-1 text-[11px] text-muted-foreground/70">
<Zap className="h-3 w-3" />
350M params
</span>
<span className="flex items-center gap-1 text-[11px] text-muted-foreground/70">
<MessageSquare className="h-3 w-3" />
[laugh] [sigh] tags
</span>
</div>
</div>
{/* LuxTTS */}
<div className="rounded-xl border border-border bg-card/60 backdrop-blur-sm p-6 transition-colors hover:border-accent/30">
<div className="flex items-start justify-between mb-3">
<div>
<h3 className="text-base font-semibold text-foreground">LuxTTS</h3>
<span className="text-xs text-muted-foreground/60">by ZipVoice</span>
</div>
</div>
<p className="text-sm text-muted-foreground leading-relaxed mb-4">
Ultra-fast, CPU-friendly voice cloning at 48kHz. Exceeds 150x realtime on CPU with
~1GB VRAM. The fastest engine for quick iterations.
</p>
<div className="flex flex-wrap gap-2">
<span className="flex items-center gap-1 text-[11px] text-muted-foreground/70">
<Zap className="h-3 w-3" />
150x realtime
</span>
<span className="flex items-center gap-1 text-[11px] text-muted-foreground/70">
48kHz output
</span>
</div>
</div>
</div>
</div>
</section>
{/* ── Download Section ─────────────────────────────────────── */}
<section id="download" className="border-t border-border py-24">
<div className="mx-auto max-w-4xl px-6">
<div className="text-center mb-12">
<h2 className="text-3xl font-semibold tracking-tight text-foreground md:text-4xl mb-4">
Download Voicebox
</h2>
<p className="text-muted-foreground">
Available for macOS, Windows, and Linux. No dependencies required.
</p>
</div>
<div className="grid grid-cols-1 sm:grid-cols-2 gap-3 max-w-2xl mx-auto">
{/* macOS ARM */}
<a
href={downloadLinks.macArm}
download
className="flex items-center rounded-xl border border-border bg-card/60 backdrop-blur-sm px-5 py-4 transition-all hover:border-accent/30 hover:bg-card group"
>
<AppleIcon className="h-6 w-6 shrink-0 text-muted-foreground group-hover:text-foreground transition-colors" />
<div className="ml-4">
<div className="text-sm font-medium">macOS</div>
<div className="text-xs text-muted-foreground">Apple Silicon (ARM)</div>
</div>
</a>
{/* macOS Intel */}
<a
href={downloadLinks.macIntel}
download
className="flex items-center rounded-xl border border-border bg-card/60 backdrop-blur-sm px-5 py-4 transition-all hover:border-accent/30 hover:bg-card group"
>
<AppleIcon className="h-6 w-6 shrink-0 text-muted-foreground group-hover:text-foreground transition-colors" />
<div className="ml-4">
<div className="text-sm font-medium">macOS</div>
<div className="text-xs text-muted-foreground">Intel (x64)</div>
</div>
</a>
{/* Windows */}
<a
href={downloadLinks.windows}
download
className="flex items-center rounded-xl border border-border bg-card/60 backdrop-blur-sm px-5 py-4 transition-all hover:border-accent/30 hover:bg-card group"
>
<WindowsIcon className="h-6 w-6 shrink-0 text-muted-foreground group-hover:text-foreground transition-colors" />
<div className="ml-4">
<div className="text-sm font-medium">Windows</div>
<div className="text-xs text-muted-foreground">64-bit (MSI)</div>
</div>
</a>
{/* Linux */}
<a
href="/linux-install"
className="flex items-center rounded-xl border border-border bg-card/60 backdrop-blur-sm px-5 py-4 transition-all hover:border-accent/30 hover:bg-card group"
>
<LinuxIcon className="h-6 w-6 shrink-0 text-muted-foreground group-hover:text-foreground transition-colors" />
<div className="ml-4">
<div className="text-sm font-medium">Linux</div>
<div className="text-xs text-muted-foreground">Build from source</div>
</div>
</a>
</div>
{/* GitHub link */}
<div className="mt-6 text-center">
<a
href={`${GITHUB_REPO}/releases`}
target="_blank"
rel="noopener noreferrer"
className="inline-flex items-center gap-2 text-sm text-muted-foreground hover:text-foreground transition-colors"
>
<Github className="h-4 w-4" />
View all releases on GitHub
</a>
</div>
</div>
</section>
{/* ── Footer ───────────────────────────────────────────────── */}
<Footer />
</>
);
}
+889
View File
@@ -0,0 +1,889 @@
'use client';
import { motion } from 'framer-motion';
import {
AudioLines,
Box,
Download,
Mic,
MoreHorizontal,
Pencil,
Server,
Sparkles,
Speaker,
Star,
Trash2,
Volume2,
Wand2,
} from 'lucide-react';
import { useCallback, useEffect, useRef, useState } from 'react';
import { LandingAudioPlayer, unlockAudioContext } from './LandingAudioPlayer';
// ─── Data ───────────────────────────────────────────────────────────────────
// Edit this section to customise all the content shown in the ControlUI demo.
interface VoiceProfile {
name: string;
description: string;
language: string;
hasEffects: boolean;
}
/** Voice profiles shown in the grid / scroll strip. Index matters — DemoScript references profiles by index. */
const PROFILES: VoiceProfile[] = [
{
name: 'Jarvis',
description: 'Dry wit, composed British AI assistant',
language: 'en',
hasEffects: true,
},
{
name: 'Samuel L. Jackson',
description: 'Commanding intensity with sharp, punchy delivery',
language: 'en',
hasEffects: true,
},
{
name: 'Bob Ross',
description: 'Gentle, soothing voice full of quiet encouragement',
language: 'en',
hasEffects: false,
},
{
name: 'Sam Altman',
description: 'Measured, thoughtful Silicon Valley cadence',
language: 'en',
hasEffects: false,
},
{
name: 'Morgan Freeman',
description: 'Rich, warm baritone with gravitas and calm authority',
language: 'en',
hasEffects: false,
},
{
name: 'Linus Tech Tips',
description: 'Enthusiastic, fast-paced tech explainer energy',
language: 'en',
hasEffects: false,
},
{
name: 'Fireship',
description: 'Rapid-fire, deadpan tech humor with zero filler',
language: 'en',
hasEffects: false,
},
{
name: 'Scarlett Johansson',
description: 'Smooth, low alto with understated warmth',
language: 'en',
hasEffects: false,
},
{
name: 'Dario Amodei',
description: 'Calm, precise articulation with academic depth',
language: 'en',
hasEffects: false,
},
{
name: 'David Attenborough',
description: 'Warm, reverent narration with wonder and precision',
language: 'en',
hasEffects: false,
},
{
name: 'Zendaya',
description: 'Relaxed, modern delivery with effortless cool',
language: 'en',
hasEffects: false,
},
{
name: 'Barack Obama',
description: 'Measured cadence with rhythmic pauses and gravitas',
language: 'en',
hasEffects: false,
},
];
/** Each entry is one cycle of the demo animation: select a profile → type text → generate → play audio. */
interface DemoStep {
profileIndex: number;
text: string;
audioUrl: string;
engine: string;
duration: string;
effect?: string;
}
const DEMO_SCRIPT: DemoStep[] = [
{
profileIndex: 0,
text: 'Sir, I have completed the analysis. Your code has twelve critical vulnerabilities, your coffee is cold, and frankly your commit messages could use some work.',
audioUrl: '/audio/jarvis.webm',
engine: 'Qwen 1.7B',
duration: '0:10',
effect: 'Robot',
},
{
profileIndex: 4,
text: "I've narrated penguins, galaxies, and the entire history of mankind. But nothing prepared me for the moment a computer learned to do my job from a five second audio clip.",
audioUrl: '/audio/morganfreeman.webm',
engine: 'Qwen 1.7B',
duration: '0:11',
effect: 'Radio',
},
{
profileIndex: 3,
text: "Open source? [laugh] What's that?",
audioUrl: '/audio/samaltman.webm',
engine: 'Chatterbox',
duration: '0:03',
},
{
profileIndex: 1,
text: "So let me get this straight. You downloaded an app, pressed a button, and now there's two of me? The world was not ready for one",
audioUrl: '/audio/samjackson.webm',
engine: 'Qwen 1.7B',
duration: '0:10',
},
{
profileIndex: 5,
text: "So we got this voice cloning software and honestly it's kind of terrifying. Like, my wife could not tell the difference. Voicebox dot s h, link in the description!",
audioUrl: '/audio/linus.webm',
engine: 'Qwen 1.7B',
duration: '0:11',
},
{
profileIndex: 6,
text: 'This is Voicebox in one hundred seconds. It clones voices locally, it runs on your GPU, and no, OpenAI cannot hear you. Lets go.',
audioUrl: '/audio/fireship.webm',
engine: 'Qwen 0.6B',
duration: '0:09',
},
];
/** History rows pre-populated on first load. Oldest first visually (array index 0 = top row). */
interface Generation {
id: number;
profileName: string;
text: string;
language: string;
engine: string;
duration: string;
timeAgo: string;
favorited: boolean;
versions: number;
}
const INITIAL_GENERATIONS: Generation[] = [
{
id: 1,
profileName: 'Morgan Freeman',
text: 'The neural pathways of human speech contain more complexity than any language model can fully capture, yet we keep pushing the boundaries of what is possible.',
language: 'en',
engine: 'Qwen 1.7B',
duration: '0:08',
timeAgo: '2 minutes ago',
favorited: true,
versions: 3,
},
{
id: 2,
profileName: 'Samuel L. Jackson',
text: 'In a world increasingly shaped by artificial intelligence, the human voice remains our most powerful tool for connection and storytelling.',
language: 'en',
engine: 'Qwen 1.7B',
duration: '0:07',
timeAgo: '15 minutes ago',
favorited: false,
versions: 1,
},
{
id: 3,
profileName: 'Jarvis',
text: 'The architecture of modern text-to-speech systems reveals an elegant interplay between transformer models and acoustic feature prediction.',
language: 'en',
engine: 'Qwen 0.6B',
duration: '0:09',
timeAgo: '1 hour ago',
favorited: false,
versions: 2,
},
{
id: 4,
profileName: 'Bob Ross',
text: 'Welcome to the next chapter. Every great story begins with a single voice, and today that voice can be yours.',
language: 'en',
engine: 'Chatterbox',
duration: '0:06',
timeAgo: '3 hours ago',
favorited: true,
versions: 1,
},
{
id: 5,
profileName: 'Linus Tech Tips',
text: 'Local inference gives you complete control over your voice data. No cloud, no subscriptions, no compromises.',
language: 'en',
engine: 'Qwen 1.7B',
duration: '0:05',
timeAgo: '5 hours ago',
favorited: false,
versions: 1,
},
];
const SIDEBAR_ITEMS = [
{ icon: Volume2, label: 'Generate' },
{ icon: AudioLines, label: 'Stories' },
{ icon: Mic, label: 'Voices' },
{ icon: Wand2, label: 'Effects' },
{ icon: Speaker, label: 'Audio' },
{ icon: Box, label: 'Models' },
{ icon: Server, label: 'Server' },
];
// ─── Phase system ───────────────────────────────────────────────────────────
type Phase = 'idle' | 'selecting' | 'typing' | 'generating' | 'complete' | 'playing';
const PHASE_DURATIONS: Record<Phase, number> = {
idle: 2500,
selecting: 800,
typing: 6000,
generating: 2800,
complete: 1200,
playing: 4000,
};
// ─── Typewriter ─────────────────────────────────────────────────────────────
function TypewriterText({ text, speed }: { text: string; speed?: number }) {
// Default: fill the typing phase duration, leaving 500ms buffer at the end
const resolvedSpeed =
speed ?? Math.max(20, Math.floor((PHASE_DURATIONS.typing - 500) / text.length));
const [displayed, setDisplayed] = useState('');
const indexRef = useRef(0);
useEffect(() => {
indexRef.current = 0;
setDisplayed('');
const interval = setInterval(() => {
indexRef.current += 1;
if (indexRef.current <= text.length) {
setDisplayed(text.slice(0, indexRef.current));
} else {
clearInterval(interval);
}
}, resolvedSpeed);
return () => clearInterval(interval);
}, [text, resolvedSpeed]);
return (
<>
{displayed}
<span className="inline-block h-3.5 w-[2px] animate-pulse bg-foreground/70 ml-[1px] align-middle" />
</>
);
}
// ─── Loading bars (simplified react-loaders replacement) ────────────────────
function LoadingBars({ mode }: { mode: 'idle' | 'generating' | 'playing' }) {
const barColor = mode !== 'idle' ? 'bg-accent' : 'bg-muted-foreground/40';
return (
<div className="flex items-center gap-[2px] h-5">
{[0, 1, 2, 3, 4].map((i) => (
<motion.div
key={`${i}-${mode}`}
className={`w-[3px] rounded-full ${barColor}`}
animate={
mode === 'generating'
? { height: ['6px', '16px', '6px'] }
: mode === 'playing'
? { height: ['8px', '14px', '4px', '12px', '8px'] }
: { height: '8px' }
}
transition={
mode === 'generating'
? { duration: 0.6, repeat: Infinity, delay: i * 0.08, ease: 'easeInOut' }
: mode === 'playing'
? { duration: 1.2, repeat: Infinity, delay: i * 0.15, ease: 'easeInOut' }
: {}
}
/>
))}
</div>
);
}
// ─── Profile Card ───────────────────────────────────────────────────────────
const ProfileCard = ({
profile,
selected,
selecting,
cardRef,
}: {
profile: VoiceProfile;
selected: boolean;
selecting: boolean;
cardRef?: React.Ref<HTMLDivElement>;
}) => {
return (
<motion.div
ref={cardRef}
className={`rounded-xl border-2 bg-card p-3.5 flex flex-col h-[143px] transition-all duration-200 ${
selected ? 'border-accent shadow-md' : 'border-border/50 hover:shadow-sm'
} ${selecting && !selected ? 'opacity-60' : ''}`}
animate={selecting && selected ? { scale: [1, 1.02, 1] } : {}}
transition={{ duration: 0.3 }}
>
<div className="text-[15px] font-bold leading-tight line-clamp-2">{profile.name}</div>
<div className="text-[10px] text-muted-foreground line-clamp-2 leading-relaxed mt-1">
{profile.description}
</div>
<div className="flex items-center gap-1.5 mt-2">
<span className="text-[10px] px-1.5 py-0.5 rounded-md border border-border text-muted-foreground">
{profile.language}
</span>
{profile.hasEffects && <Sparkles className="h-3 w-3 text-accent fill-accent" />}
</div>
<div className="flex items-center gap-1 mt-auto justify-end">
<Download className="h-3.5 w-3.5 text-muted-foreground/40" />
<Pencil className="h-3.5 w-3.5 text-muted-foreground/40" />
<Trash2 className="h-3.5 w-3.5 text-muted-foreground/40" />
</div>
</motion.div>
);
};
// ─── History Row ────────────────────────────────────────────────────────────
function HistoryRow({
gen,
mode,
isNew,
}: {
gen: Generation;
mode: 'idle' | 'generating' | 'playing';
isNew: boolean;
}) {
return (
<motion.div
className={`border rounded-md transition-colors text-left w-full ${
mode === 'playing' ? 'bg-muted/70' : 'bg-card'
}`}
initial={isNew ? { opacity: 0, y: -8 } : false}
animate={{ opacity: 1, y: 0 }}
transition={{ duration: 0.3, ease: 'easeOut' }}
>
<div className="flex items-stretch gap-3 h-[80px] p-2.5">
{/* Status icon */}
<div className="w-8 flex items-center justify-center shrink-0">
<LoadingBars mode={mode} />
</div>
{/* Meta info */}
<div className="flex flex-col gap-1 w-36 shrink-0 justify-center">
<div className="text-[12px] font-medium truncate">{gen.profileName}</div>
<div className="flex items-center gap-2 text-[10px] text-muted-foreground">
<span>{gen.language}</span>
<span>{gen.engine}</span>
{mode !== 'generating' && <span>{gen.duration}</span>}
</div>
<div className="text-[10px] text-muted-foreground">
{mode === 'generating' ? (
<span className="text-accent">Generating...</span>
) : (
gen.timeAgo
)}
</div>
</div>
{/* Transcript */}
<div className="flex-1 min-w-0 flex items-center">
<div className="text-[11px] text-muted-foreground line-clamp-3 leading-relaxed">
{gen.text}
</div>
</div>
{/* Action buttons */}
<div className="flex flex-col justify-center items-center gap-0.5 shrink-0">
<button className="h-5 w-5 flex items-center justify-center rounded-sm hover:bg-muted">
<Star
className={`h-2.5 w-2.5 ${
gen.favorited ? 'text-accent fill-accent' : 'text-muted-foreground/50'
}`}
/>
</button>
{gen.versions > 1 && (
<button className="h-5 w-5 flex items-center justify-center rounded-sm hover:bg-muted">
<AudioLines className="h-2.5 w-2.5 text-muted-foreground/50" />
</button>
)}
<button className="h-5 w-5 flex items-center justify-center rounded-sm hover:bg-muted">
<MoreHorizontal className="h-2.5 w-2.5 text-muted-foreground/50" />
</button>
</div>
</div>
</motion.div>
);
}
// ─── Floating Generate Box ──────────────────────────────────────────────────
function FloatingGenerateBox({
phase,
typingText,
selectedProfile,
engine,
effect,
}: {
phase: Phase;
typingText: string;
selectedProfile: VoiceProfile | null;
engine: string;
effect?: string;
}) {
const isFocused = phase === 'typing' || phase === 'generating';
const isGenerating = phase === 'generating';
return (
<motion.div
className="bg-background/30 backdrop-blur-2xl border border-accent/20 rounded-[1.5rem] shadow-2xl p-2.5"
animate={{
borderColor: isGenerating
? 'hsl(43 50% 45% / 0.35)'
: isFocused
? 'hsl(43 50% 45% / 0.25)'
: 'hsl(43 50% 45% / 0.15)',
}}
transition={{ duration: 0.3 }}
>
{/* Text area + generate button */}
<div className="flex items-start gap-2">
<div className="flex-1 min-w-0">
<motion.div
className="overflow-hidden"
animate={{ height: isFocused ? 100 : 32 }}
transition={{ duration: 0.25, ease: 'easeOut' }}
>
<div
className="text-[12.5px] text-muted-foreground/60 px-2 py-1 leading-relaxed"
style={{ minHeight: isFocused ? 100 : 32 }}
>
{phase === 'typing' ? (
<span className="text-foreground">
<TypewriterText text={typingText} />
</span>
) : phase === 'generating' ? (
<span className="text-muted-foreground/40">{typingText}</span>
) : (
<span>
{selectedProfile
? `Generate speech using ${selectedProfile.name}...`
: 'Select a voice profile above...'}
</span>
)}
</div>
</motion.div>
</div>
{/* Generate button */}
<button className="h-8 w-8 rounded-full bg-accent flex items-center justify-center shrink-0 shadow-lg">
<Sparkles className="h-3.5 w-3.5 text-white fill-white" />
</button>
</div>
{/* Bottom selectors */}
<div className="flex items-center gap-1.5 mt-2">
<span className="text-[10px] px-2 py-1 rounded-full border border-border bg-card text-muted-foreground">
English
</span>
<span className="text-[10px] px-2 py-1 rounded-full border border-border bg-card text-muted-foreground">
{engine}
</span>
<span
className={`text-[10px] px-2 py-1 rounded-full border flex items-center gap-1 ${
effect
? 'border-accent/30 bg-accent/10 text-accent'
: 'border-border bg-card text-muted-foreground'
}`}
>
<Sparkles className={`h-2.5 w-2.5 ${effect ? 'fill-accent' : ''}`} />
{effect || 'Effect'}
</span>
</div>
</motion.div>
);
}
// ─── Main ControlUI ─────────────────────────────────────────────────────────
export function ControlUI() {
const [phase, setPhase] = useState<Phase>('idle');
const [selectedIndex, setSelectedIndex] = useState(DEMO_SCRIPT[0].profileIndex);
const [cycle, setCycle] = useState(0);
const [newGenId, setNewGenId] = useState<number | null>(null);
const [generations, setGenerations] = useState<Generation[]>([...INITIAL_GENERATIONS]);
const [isMuted, setIsMuted] = useState(true);
const [isVisible, setIsVisible] = useState(true);
const [pageHidden, setPageHidden] = useState(false);
const containerRef = useRef<HTMLDivElement>(null);
const phaseRef = useRef(phase);
const mobileCardRefs = useRef<Map<number, HTMLDivElement>>(new Map());
const desktopCardRefs = useRef<Map<number, HTMLDivElement>>(new Map());
const profileGridRef = useRef<HTMLDivElement>(null);
const [scrollLeft, setScrollLeft] = useState(0);
phaseRef.current = phase;
const step = DEMO_SCRIPT[cycle % DEMO_SCRIPT.length];
const selectedProfile = PROFILES[selectedIndex];
// Scroll to selected profile card — accounts for generate box overlay on desktop
useEffect(() => {
const isMobile = window.innerWidth < 768;
if (isMobile) {
const el = mobileCardRefs.current.get(selectedIndex);
if (el) el.scrollIntoView({ behavior: 'smooth', block: 'nearest', inline: 'center' });
return;
}
// Desktop
const el = desktopCardRefs.current.get(selectedIndex);
const scrollContainer = profileGridRef.current;
if (!el || !scrollContainer) return;
const containerTop = scrollContainer.getBoundingClientRect().top;
const elTop = el.getBoundingClientRect().top;
const elRelTop = elTop - containerTop + scrollContainer.scrollTop;
const rowHeight = 145;
const generateBoxHeight = 200;
const visibleTop = scrollContainer.scrollTop;
const visibleBottom = visibleTop + scrollContainer.clientHeight - generateBoxHeight;
const elRelBottom = elRelTop + el.offsetHeight;
if (elRelTop >= visibleTop && elRelBottom <= visibleBottom) {
return;
}
const target = elRelTop - rowHeight;
scrollContainer.scrollTo({ top: Math.max(0, target), behavior: 'smooth' });
}, [selectedIndex]);
// Visibility detection
useEffect(() => {
const observer = new IntersectionObserver(([entry]) => setIsVisible(entry.isIntersecting), {
threshold: 0,
});
if (containerRef.current) observer.observe(containerRef.current);
const handleVisibility = () => setPageHidden(document.visibilityState !== 'visible');
document.addEventListener('visibilitychange', handleVisibility);
return () => {
observer.disconnect();
document.removeEventListener('visibilitychange', handleVisibility);
};
}, []);
const paused = !isVisible || pageHidden;
// Phase cycling — `playing` phase is driven by audio finish, not a timeout
useEffect(() => {
if (paused || phase === 'playing') return;
const duration = PHASE_DURATIONS[phase];
const timer = setTimeout(() => {
console.log(
'[ControlUI] phase transition',
phase,
'→ next, cycle:',
cycle,
'step profile:',
PROFILES[step.profileIndex].name,
);
switch (phase) {
case 'idle': {
setSelectedIndex(step.profileIndex);
setPhase('selecting');
break;
}
case 'selecting':
setPhase('typing');
break;
case 'typing': {
const profile = PROFILES[step.profileIndex];
const newGen: Generation = {
id: Date.now(),
profileName: profile.name,
text: step.text,
language: profile.language,
engine: step.engine,
duration: step.duration,
timeAgo: 'just now',
favorited: false,
versions: 1,
};
setGenerations((prev) => [newGen, ...prev.slice(0, 5)]);
setNewGenId(newGen.id);
setPhase('generating');
break;
}
case 'generating':
setPhase('playing');
break;
}
}, duration);
return () => clearTimeout(timer);
}, [phase, paused, step, cycle]);
const handleAudioFinish = useCallback(() => {
if (phaseRef.current !== 'playing') return;
setPhase('idle');
setCycle((c) => c + 1);
setNewGenId(null);
}, []);
const isGenerating = phase === 'generating';
return (
<div ref={containerRef} className="relative z-20 mx-auto w-full max-w-6xl px-6">
{/* Unmute button with handwritten hint */}
<div className="flex justify-end mb-3">
<div className="relative">
{/* Handwritten hint — absolutely positioned above the button */}
{isMuted && (
<motion.div
className="absolute select-none pointer-events-none"
style={{ top: -30, right: 100 }}
initial={{ opacity: 0, y: 6 }}
animate={{ opacity: 1, y: 0 }}
transition={{ delay: 2, duration: 0.6, ease: 'easeOut' }}
>
<span
className="text-xl text-accent/80 whitespace-nowrap"
style={{
fontFamily: "'Caveat', 'Segoe Script', 'Comic Sans MS', cursive",
letterSpacing: '0.02em',
}}
>
try me!
</span>
{/* Curved arrow from text down-right toward the button */}
<svg
width="22"
height="11"
viewBox="0 0 80 40"
fill="none"
className="text-accent/70 absolute"
style={{ top: 14, left: 60 }}
aria-hidden="true"
>
<title>Arrow</title>
<path
d="M4 4 C20 4, 40 8, 55 20 C62 26, 66 32, 70 36"
stroke="currentColor"
strokeWidth="3"
strokeLinecap="round"
fill="none"
/>
<path
d="M58 42 L70 36 L64 22"
stroke="currentColor"
strokeWidth="3"
strokeLinecap="round"
strokeLinejoin="round"
transform="rotate(35, 70, 36)"
fill="none"
/>
</svg>
</motion.div>
)}
<button
onClick={() => {
unlockAudioContext();
setIsMuted(!isMuted);
}}
className="flex items-center gap-2 px-3 py-1.5 rounded-full border border-border bg-card/50 backdrop-blur text-xs text-muted-foreground hover:text-foreground transition-colors"
>
{isMuted ? (
<>
<Volume2 className="h-3.5 w-3.5" />
<span>Unmute</span>
</>
) : (
<>
<Volume2 className="h-3.5 w-3.5 text-accent" />
<span>Mute</span>
</>
)}
</button>
</div>
</div>
<div className="overflow-hidden rounded-2xl border border-app-line bg-app-box shadow-[0_25px_60px_rgba(0,0,0,0.5),0_8px_20px_rgba(0,0,0,0.3)] md:h-[640px] pointer-events-none select-none">
<div className="flex flex-col md:flex-row h-full">
{/* ── Sidebar (hidden on mobile) ─────────────────────────── */}
<div className="hidden md:flex w-16 shrink-0 border-r border-app-line bg-sidebar flex-col items-center py-4 gap-4">
{/* Logo */}
<div className="mb-1">
<div
className="w-9 h-9 rounded-lg overflow-hidden"
style={{
filter:
'drop-shadow(0 0 6px hsl(43 50% 45% / 0.5)) drop-shadow(0 0 14px hsl(43 50% 45% / 0.35))',
}}
>
<img
src="/voicebox-logo-app.webp"
alt=""
className="w-full h-full object-contain"
/>
</div>
</div>
{/* Nav items */}
<div className="flex flex-col gap-2">
{SIDEBAR_ITEMS.map((item, i) => {
const Icon = item.icon;
const active = i === 0;
return (
<div
key={item.label}
className={`w-9 h-9 rounded-full flex items-center justify-center transition-all duration-200 ${
active
? 'bg-white/[0.07] text-foreground shadow-lg backdrop-blur-sm border border-white/[0.08]'
: 'text-muted-foreground/60'
}`}
>
<Icon className="h-4 w-4" />
</div>
);
})}
</div>
{/* Version */}
<div className="mt-auto text-[8px] text-muted-foreground/40">v0.2.0</div>
</div>
{/* ── Main content ──────────────────────────────────────── */}
<div className="flex-1 flex flex-col md:flex-row min-w-0 relative">
{/* Left: Profiles + Generate box */}
<div className="flex flex-col min-w-0 relative md:flex-1 md:overflow-hidden">
{/* Gradient fade overlay — sits between header and scroll content */}
<div className="hidden md:block absolute top-0 left-0 right-0 h-16 bg-gradient-to-b from-app-box to-transparent z-[1] pointer-events-none" />
{/* Header — floats above everything */}
<div className="absolute top-0 left-0 right-0 z-10 px-4 pt-4 md:pt-6 pb-2 flex items-center justify-between">
<h2 className="text-base font-bold">Voicebox</h2>
<div className="flex items-center gap-1.5">
<button className="h-6 text-[10px] px-2.5 rounded-full border border-border bg-card text-muted-foreground flex items-center gap-1">
Import Voice
</button>
<button className="h-6 text-[10px] px-2.5 rounded-full bg-accent text-accent-foreground flex items-center">
Create Voice
</button>
</div>
</div>
{/* Scrollable profile cards — scrolls behind header + gradient */}
<div
ref={profileGridRef}
className="flex-1 min-h-0 md:overflow-y-auto md:pt-14 pt-12"
>
<div className="px-4">
{/* Mobile: horizontal scroll strip with edge fade */}
<div className="relative md:hidden">
{scrollLeft > 0 && (
<div className="absolute left-0 top-0 bottom-0 w-6 bg-gradient-to-r from-app-box to-transparent z-10" />
)}
<div className="absolute right-0 top-0 bottom-0 w-6 bg-gradient-to-l from-app-box to-transparent z-10" />
<div
className="flex gap-2 overflow-x-auto pb-2"
onScroll={(e) => setScrollLeft(e.currentTarget.scrollLeft)}
>
{PROFILES.map((profile, i) => (
<div
key={profile.name}
className="shrink-0 w-[140px]"
ref={(el) => {
if (el) mobileCardRefs.current.set(i, el);
}}
>
<ProfileCard
profile={profile}
selected={i === selectedIndex}
selecting={phase === 'selecting'}
/>
</div>
))}
</div>
</div>
{/* Desktop: 3-col grid */}
<div className="hidden md:grid grid-cols-3 gap-2 mt-1 pb-44">
{PROFILES.map((profile, i) => (
<ProfileCard
key={profile.name}
profile={profile}
selected={i === selectedIndex}
selecting={phase === 'selecting'}
cardRef={(el: HTMLDivElement | null) => {
if (el) desktopCardRefs.current.set(i, el);
}}
/>
))}
</div>
</div>
</div>
{/* Floating generate box — desktop: absolute overlay, mobile: inline */}
<div className="px-3 pt-2 pb-3 md:pt-0 md:absolute md:left-4 md:right-4 md:bottom-[117px] md:z-20 md:pb-0 md:px-0">
<FloatingGenerateBox
phase={phase}
typingText={step.text}
selectedProfile={selectedProfile}
engine={step.engine}
effect={step.effect}
/>
</div>
</div>
{/* Right/Below: History */}
<div className="md:w-[48%] shrink-0 flex flex-col min-w-0 border-t md:border-t-0 border-app-line">
<div className="max-h-[360px] md:max-h-none flex-1 overflow-hidden px-3 pt-3 md:pt-6 pb-3">
<div className="flex flex-col gap-2">
{generations.map((gen) => {
const isThisNew = gen.id === newGenId;
const rowMode: 'idle' | 'generating' | 'playing' =
isThisNew && isGenerating
? 'generating'
: isThisNew && phase === 'playing'
? 'playing'
: 'idle';
return <HistoryRow key={gen.id} gen={gen} mode={rowMode} isNew={isThisNew} />;
})}
</div>
</div>
</div>
{/* Audio player */}
<LandingAudioPlayer
audioUrl={step.audioUrl}
title={selectedProfile.name}
playing={phase === 'playing'}
muted={isMuted}
onFinish={handleAudioFinish}
onClose={() => {}}
/>
</div>
</div>
</div>
</div>
);
}
+855
View File
@@ -0,0 +1,855 @@
'use client';
import { motion } from 'framer-motion';
import { AudioLines, Cloud, MessageSquareText, Mic, Sparkles, TextCursorInput } from 'lucide-react';
import { useEffect, useMemo, useRef, useState } from 'react';
// ─── Lazy load wrapper ──────────────────────────────────────────────────────
function LazyLoad({
children,
className,
rootMargin = '200px',
}: {
children: React.ReactNode;
className?: string;
rootMargin?: string;
}) {
const ref = useRef<HTMLDivElement>(null);
const [visible, setVisible] = useState(false);
useEffect(() => {
const el = ref.current;
if (!el) return;
const observer = new IntersectionObserver(
([entry]) => {
if (entry.isIntersecting) {
setVisible(true);
observer.disconnect();
}
},
{ rootMargin },
);
observer.observe(el);
return () => observer.disconnect();
}, [rootMargin]);
return (
<div ref={ref} className={className}>
{visible ? children : null}
</div>
);
}
// ─── Animation: Voice Cloning ───────────────────────────────────────────────
function VoiceCloningAnimation() {
const [phase, setPhase] = useState(0);
useEffect(() => {
const interval = setInterval(() => {
setPhase((p) => (p + 1) % 3);
}, 2400);
return () => clearInterval(interval);
}, []);
const samples = ['Sample 1', 'Sample 2', 'Sample 3'];
const bars = [0.4, 0.7, 0.5, 0.9, 0.3, 0.6, 0.8, 0.4, 0.7, 0.5, 0.3, 0.6];
return (
<div className="h-40 w-full flex items-center justify-center overflow-hidden rounded-md bg-app-darkerBox/50 p-4">
<div className="flex flex-col items-center gap-3 w-full max-w-[200px]">
{/* Sample pills */}
<div className="flex gap-1.5">
{samples.map((s, i) => (
<motion.div
key={s}
className="text-[9px] px-2 py-1 rounded-full border font-medium"
animate={{
borderColor: i === phase ? 'hsl(43 50% 45% / 0.5)' : 'rgba(255,255,255,0.06)',
backgroundColor: i === phase ? 'hsl(43 50% 45% / 0.08)' : 'rgba(255,255,255,0.02)',
color: i === phase ? 'hsl(43 50% 45%)' : 'rgba(255,255,255,0.4)',
}}
transition={{ duration: 0.3 }}
>
{s}
</motion.div>
))}
</div>
{/* Waveform visualization */}
<div className="flex items-center gap-[2px] h-10 w-full justify-center">
{bars.map((h, i) => (
<motion.div
key={i}
className="w-[4px] rounded-full"
animate={{
height: `${h * 100}%`,
backgroundColor: phase === 2 ? 'hsl(43 50% 45%)' : 'rgba(255,255,255,0.15)',
}}
transition={{
height: { duration: 0.6, delay: i * 0.04, ease: 'easeInOut' },
backgroundColor: { duration: 0.3 },
}}
/>
))}
</div>
{/* Result label */}
<motion.div
className="text-[9px] font-mono"
animate={{
opacity: phase === 2 ? 1 : 0.3,
color: phase === 2 ? 'hsl(43 50% 45%)' : 'rgba(255,255,255,0.3)',
}}
transition={{ duration: 0.3 }}
>
voice profile ready
</motion.div>
</div>
</div>
);
}
// ─── Mini waveform for clips ────────────────────────────────────────────────
// Fixed-width dense waveform that overflows — the clip container clips it.
// This way resizing a clip just reveals/hides bars instead of re-rendering.
const WAVEFORM_BAR_COUNT = 60;
function MiniWaveform({ seed, color }: { seed: number; color: string }) {
// Deterministic pseudo-random waveform that looks like real speech audio.
// Uses layered noise at different frequencies for natural envelope + detail.
const bars = useMemo(() => {
// Seeded pseudo-random number generator (deterministic per seed)
let s = seed * 9301 + 49297;
const rand = () => {
s = (s * 16807 + 0) % 2147483647;
return s / 2147483647;
};
// Pre-generate random values
const r = Array.from({ length: WAVEFORM_BAR_COUNT }, () => rand());
return Array.from({ length: WAVEFORM_BAR_COUNT }, (_, i) => {
const t = i / WAVEFORM_BAR_COUNT;
// Slow envelope — broad amplitude shape (words / phrases)
const envelope =
0.3 +
0.35 *
Math.sin(t * Math.PI * (2 + (seed % 3))) *
Math.sin(t * Math.PI * (1.3 + seed * 0.7)) +
0.2 * Math.sin(t * Math.PI * (4.7 + seed * 1.3));
// Medium variation — syllable-level bumps
const mid = 0.15 * Math.sin(i * 0.8 + seed * 3.1) * Math.cos(i * 1.3 + seed);
// High-frequency noise — individual sample jitter
const noise = (r[i] - 0.5) * 0.25;
// Combine and clamp
const raw = envelope + mid + noise;
return Math.max(0.06, Math.min(1, raw));
});
}, [seed]);
return (
<div className="flex items-center h-full overflow-hidden">
{bars.map((h, i) => (
<div
key={`w-${seed}-${i}`}
className="shrink-0 rounded-full opacity-50"
style={{
width: 2,
marginRight: 1,
height: `${h * 100}%`,
backgroundColor: color,
}}
/>
))}
</div>
);
}
// ─── Animation: Stories Editor ───────────────────────────────────────────────
// Clip shape: id, profile, track, left (px out of 220), width (px), waveform seed
type DemoClip = { id: string; profile: string; track: number; x: number; w: number; seed: number };
const INITIAL_CLIPS: DemoClip[] = [
{ id: 'n1', profile: 'Morgan', track: 0, x: 4, w: 70, seed: 1 },
{ id: 'n2', profile: 'Morgan', track: 0, x: 135, w: 35, seed: 2 },
{ id: 'a1', profile: 'Scarlett', track: 1, x: 25, w: 40, seed: 3 },
{ id: 'a2', profile: 'Scarlett', track: 1, x: 120, w: 35, seed: 4 },
{ id: 'b1', profile: 'Jarvis', track: 2, x: 70, w: 45, seed: 5 },
];
// Timeline width the clips live inside
const TL_W = 220;
// Each action returns a new clips array (or modifies in place)
type Action = { label: string; apply: (clips: DemoClip[]) => DemoClip[] };
const ACTIONS: Action[] = [
// 0 — move Jarvis clip earlier
{ label: 'Move clip', apply: (c) => c.map((cl) => (cl.id === 'b1' ? { ...cl, x: 55 } : cl)) },
// 1 — split Morgan's first clip into two with visible gap
{
label: 'Split clip',
apply: (c) => {
// Idempotent: if n1b already exists, the split already happened
if (c.some((cl) => cl.id === 'n1b')) return c;
const clip = c.find((cl) => cl.id === 'n1');
if (!clip) return c;
const leftW = 25;
const gap = 8;
const rightW = clip.w - leftW - gap;
return [
...c.filter((cl) => cl.id !== 'n1'),
{ ...clip, w: leftW, id: 'n1' },
{
id: 'n1b',
profile: clip.profile,
track: clip.track,
x: clip.x + leftW + gap,
w: rightW,
seed: 6,
},
];
},
},
// 2 — trim Scarlett's second clip shorter
{ label: 'Trim clip', apply: (c) => c.map((cl) => (cl.id === 'a2' ? { ...cl, w: 25 } : cl)) },
// 3 — duplicate Jarvis to track 0
{
label: 'Duplicate',
apply: (c) => {
// Idempotent: if b1d already exists, the duplicate already happened
if (c.some((cl) => cl.id === 'b1d')) return c;
const clip = c.find((cl) => cl.id === 'b1');
if (!clip) return c;
return [...c, { ...clip, id: 'b1d', track: 0, x: 180, w: 35, seed: 7 }];
},
},
// 4 — reset
{ label: '', apply: () => INITIAL_CLIPS },
];
function StoriesAnimation() {
const [clips, setClips] = useState<DemoClip[]>(INITIAL_CLIPS);
const [actionIndex, setActionIndex] = useState(-1);
const [playheadX, setPlayheadX] = useState(0);
const [selectedId, setSelectedId] = useState<string | null>(null);
const playheadRef = useRef<ReturnType<typeof requestAnimationFrame>>(0);
// Animate the playhead continuously
useEffect(() => {
let start: number | null = null;
const speed = 12; // px per second
const animate = (ts: number) => {
if (start === null) start = ts;
const elapsed = (ts - start) / 1000;
setPlayheadX((elapsed * speed) % TL_W);
playheadRef.current = requestAnimationFrame(animate);
};
playheadRef.current = requestAnimationFrame(animate);
return () => cancelAnimationFrame(playheadRef.current);
}, []);
// Step through actions
useEffect(() => {
const interval = setInterval(() => {
setActionIndex((prev) => {
const next = (prev + 1) % ACTIONS.length;
setClips((current) => ACTIONS[next].apply(current));
// Highlight the clip being acted on
if (next === 0) setSelectedId('b1');
else if (next === 1) setSelectedId('n1');
else if (next === 2) setSelectedId('a2');
else if (next === 3) setSelectedId('b1');
else setSelectedId(null);
return next;
});
}, 2600);
return () => clearInterval(interval);
}, []);
const trackLabels = ['1', '0', '-1'];
const timeMarkers = [0, 2, 4, 6, 8];
const accentColor = 'hsl(43 50% 45%)';
const accentFg = 'hsl(30 10% 94%)';
return (
<div className="h-40 w-full flex flex-col overflow-hidden rounded-md bg-app-darkerBox/50">
{/* Toolbar */}
<div className="flex items-center gap-1.5 px-2 py-1 border-b border-app-line bg-app-darkBox/60 shrink-0">
<div className="w-1.5 h-1.5 rounded-full bg-ink-faint/40" />
<div className="flex items-center gap-1">
<div className="w-4 h-4 rounded flex items-center justify-center bg-app-button">
<div className="border-l-[4px] border-l-ink-faint border-t-[3px] border-t-transparent border-b-[3px] border-b-transparent ml-0.5" />
</div>
<div className="w-4 h-4 rounded flex items-center justify-center bg-app-button">
<div className="w-2 h-2 rounded-sm bg-ink-faint/60" />
</div>
</div>
<span className="text-[8px] text-ink-faint font-mono ml-1 tabular-nums">0:03 / 0:10</span>
<div className="flex-1" />
{actionIndex >= 0 && actionIndex < ACTIONS.length - 1 && (
<motion.span
key={actionIndex}
className="text-[7px] font-medium px-1.5 py-0.5 rounded-full"
style={{
backgroundColor: `${accentColor.replace(')', ' / 0.15)')}`,
color: accentColor,
}}
initial={{ opacity: 0, y: 3 }}
animate={{ opacity: 1, y: 0 }}
exit={{ opacity: 0 }}
transition={{ duration: 0.25 }}
>
{ACTIONS[actionIndex].label}
</motion.span>
)}
<div className="flex items-center gap-0.5">
<span className="text-[7px] text-ink-faint">Zoom</span>
<div className="w-3 h-3 rounded flex items-center justify-center bg-app-button text-[8px] text-ink-faint">
-
</div>
<div className="w-3 h-3 rounded flex items-center justify-center bg-app-button text-[8px] text-ink-faint">
+
</div>
</div>
</div>
{/* Timeline */}
<div className="flex flex-1 min-h-0">
{/* Track labels sidebar */}
<div className="w-7 shrink-0 border-r border-app-line bg-app-darkBox/30 flex flex-col">
<div className="h-5 border-b border-app-line" />
{trackLabels.map((label) => (
<div
key={label}
className="flex-1 flex items-center justify-center border-b border-app-line"
>
<span className="text-[7px] text-ink-faint select-none">{label}</span>
</div>
))}
</div>
{/* Tracks area */}
<div className="flex-1 relative overflow-hidden flex flex-col">
{/* Time ruler */}
<div className="h-5 shrink-0 border-b border-app-line bg-app-darkBox/20 relative">
{timeMarkers.map((t) => (
<div
key={`tm-${t}`}
className="absolute top-0 h-full flex flex-col justify-end pb-0.5"
style={{ left: `${(t / 10) * 100}%` }}
>
<div className="h-1.5 w-px bg-app-line" />
<span className="text-[7px] text-ink-faint ml-0.5 select-none">{`0:0${t}`}</span>
</div>
))}
</div>
{/* Track rows + clips — same parent so percentages match */}
<div className="flex-1 relative min-h-0">
{/* Track rows background */}
{trackLabels.map((label, i) => (
<div
key={`bg-${label}`}
className="border-b border-app-line absolute left-0 right-0"
style={{
height: `${100 / 3}%`,
top: `${(i * 100) / 3}%`,
backgroundColor: i % 2 === 0 ? 'transparent' : 'rgba(255,255,255,0.01)',
}}
/>
))}
{/* Clips */}
{clips.map((clip) => {
const trackIdx = clip.track;
const isSelected = clip.id === selectedId;
const clipTop = `calc(${(trackIdx * 100) / 3}% + 2px)`;
const clipHeight = `calc(${100 / 3}% - 4px)`;
return (
<motion.div
key={clip.id}
className="absolute rounded overflow-hidden"
initial={false}
style={{
height: clipHeight,
left: `${(clip.x / TL_W) * 100}%`,
width: `${(clip.w / TL_W) * 100}%`,
top: clipTop,
}}
animate={{
left: `${(clip.x / TL_W) * 100}%`,
width: `${(clip.w / TL_W) * 100}%`,
top: clipTop,
}}
transition={{ type: 'spring', stiffness: 200, damping: 25 }}
>
<div
className="w-full h-full rounded overflow-hidden flex flex-col"
style={{
backgroundColor: isSelected ? 'hsl(43 50% 45%)' : 'hsl(43 45% 40%)',
boxShadow: isSelected
? 'inset 0 0 0 1px hsl(43 50% 55%), 0 0 0 1px hsl(30 10% 94% / 0.4)'
: 'inset 0 0 0 1px hsl(30 10% 94% / 0.1)',
}}
>
{/* Profile label — scaled to bypass browser min font size */}
<div className="shrink-0 relative" style={{ height: 9 }}>
<span
className="text-[10px] font-medium leading-none absolute top-0 left-0.5 origin-top-left opacity-80 whitespace-nowrap"
style={{ color: accentFg, transform: 'scale(0.75)' }}
>
{clip.profile}
</span>
</div>
{/* Waveform — absolutely positioned so it never affects clip width */}
<div className="absolute left-0 right-0 bottom-0" style={{ top: 9 }}>
<MiniWaveform seed={clip.seed} color={accentFg} />
</div>
</div>
{/* Trim handles on selected */}
{isSelected && (
<>
<div
className="absolute left-0 top-0 bottom-0 w-1 rounded-l"
style={{ backgroundColor: 'hsl(30 10% 94% / 0.25)' }}
/>
<div
className="absolute right-0 top-0 bottom-0 w-1 rounded-r"
style={{ backgroundColor: 'hsl(30 10% 94% / 0.25)' }}
/>
</>
)}
</motion.div>
);
})}
{/* Playhead */}
<motion.div
className="absolute top-0 bottom-0 w-[2px] rounded-full z-20 pointer-events-none"
style={{ backgroundColor: accentColor }}
animate={{ left: `${(playheadX / TL_W) * 100}%` }}
transition={{ duration: 0.05, ease: 'linear' }}
>
<div
className="absolute -top-0.5 left-1/2 -translate-x-1/2 w-2 h-2 rounded-full"
style={{ backgroundColor: accentColor }}
/>
</motion.div>
</div>
</div>
</div>
</div>
);
}
// ─── Animation: Effects Pipeline ────────────────────────────────────────────
function EffectsAnimation() {
const [activeEffect, setActiveEffect] = useState(0);
const effects = [
{ name: 'Pitch Shift', param: '-3 semitones', color: '#3b82f6' },
{ name: 'Reverb', param: 'Room 0.7', color: '#8b5cf6' },
{ name: 'Compressor', param: '-15 dB', color: '#ec4899' },
{ name: 'Low-Pass', param: '6000 Hz', color: '#14b8a6' },
];
// Waveform bars — original shape
const rawBars = [0.3, 0.6, 0.8, 0.5, 0.9, 0.4, 0.7, 0.3, 0.6, 0.5, 0.8, 0.4, 0.7, 0.9, 0.3];
useEffect(() => {
const interval = setInterval(() => {
setActiveEffect((p) => (p + 1) % effects.length);
}, 2200);
return () => clearInterval(interval);
}, [effects.length]);
return (
<div className="h-40 w-full flex flex-col items-center justify-center overflow-hidden rounded-md bg-app-darkerBox/50 p-4 gap-3">
{/* Effects chain */}
<div className="flex items-center gap-1">
{effects.map((fx, i) => (
<div key={fx.name} className="flex items-center gap-1">
<motion.div
className="text-[8px] px-2 py-0.5 rounded-full border font-medium"
animate={{
borderColor: i <= activeEffect ? `${fx.color}60` : 'rgba(255,255,255,0.06)',
backgroundColor: i <= activeEffect ? `${fx.color}15` : 'rgba(255,255,255,0.02)',
color: i <= activeEffect ? fx.color : 'rgba(255,255,255,0.3)',
}}
transition={{ duration: 0.3 }}
>
{fx.name}
</motion.div>
{i < effects.length - 1 && (
<motion.span
className="text-[8px]"
animate={{
color: i < activeEffect ? 'rgba(255,255,255,0.3)' : 'rgba(255,255,255,0.08)',
}}
transition={{ duration: 0.3 }}
>
&rarr;
</motion.span>
)}
</div>
))}
</div>
{/* Waveform that morphs as effects are applied */}
<div className="flex items-center gap-[2px] h-10 w-full max-w-[200px] justify-center">
{rawBars.map((h, i) => {
// Each effect stage progressively transforms the shape
const shifted = activeEffect >= 0 ? h * (0.7 + 0.3 * Math.sin(i * 0.8)) : h;
const dampened = activeEffect >= 1 ? shifted * (0.6 + 0.4 * Math.cos(i * 0.3)) : shifted;
const compressed = activeEffect >= 2 ? 0.3 + dampened * 0.5 : dampened;
const filtered = activeEffect >= 3 ? compressed * (1 - i * 0.03) : compressed;
const finalH = Math.max(0.08, Math.min(1, filtered));
return (
<motion.div
key={`bar-${i}`}
className="w-[3px] rounded-full"
animate={{
height: `${finalH * 100}%`,
backgroundColor: effects[activeEffect].color,
}}
transition={{
height: { duration: 0.5, delay: i * 0.02, ease: 'easeInOut' },
backgroundColor: { duration: 0.4 },
}}
/>
);
})}
</div>
{/* Active effect detail */}
<motion.div
className="text-[9px] font-mono text-ink-faint"
key={activeEffect}
initial={{ opacity: 0, y: 4 }}
animate={{ opacity: 1, y: 0 }}
transition={{ duration: 0.3 }}
>
{effects[activeEffect].name}: {effects[activeEffect].param}
</motion.div>
</div>
);
}
// ─── Animation: Local or Remote ─────────────────────────────────────────────
function LocalRemoteAnimation() {
const [mode, setMode] = useState(0);
const modes = ['Local GPU', 'Remote Server'];
useEffect(() => {
const interval = setInterval(() => {
setMode((p) => (p + 1) % 2);
}, 2800);
return () => clearInterval(interval);
}, []);
return (
<div className="h-40 w-full flex items-center justify-center overflow-hidden rounded-md bg-app-darkerBox/50 p-4">
<div className="flex flex-col items-center gap-4 w-full max-w-[180px]">
{/* Toggle */}
<div className="flex gap-1 p-0.5 rounded-full border border-app-line bg-app-darkerBox">
{modes.map((m, i) => (
<motion.div
key={m}
className="text-[9px] px-3 py-1 rounded-full font-medium"
animate={{
backgroundColor: i === mode ? 'hsl(43 50% 45%)' : 'transparent',
color: i === mode ? 'hsl(30 10% 94%)' : 'rgba(255,255,255,0.35)',
}}
transition={{ duration: 0.25 }}
>
{m}
</motion.div>
))}
</div>
{/* Status */}
<div className="flex flex-col items-center gap-2">
<motion.div
className="w-2 h-2 rounded-full"
animate={{
backgroundColor: mode === 0 ? '#4ade80' : '#3b82f6',
boxShadow: mode === 0 ? '0 0 8px #4ade80' : '0 0 8px #3b82f6',
}}
transition={{ duration: 0.3 }}
/>
<span className="text-[9px] text-ink-faint font-mono">
{mode === 0 ? 'Metal acceleration active' : 'Connected to 192.168.1.50'}
</span>
<span className="text-[8px] text-ink-faint/60 font-mono">
{mode === 0 ? 'VRAM: 8.2 / 16.0 GB' : 'Latency: 12ms | CUDA'}
</span>
</div>
</div>
</div>
);
}
// ─── Animation: Transcription ───────────────────────────────────────────────
function TranscriptionAnimation() {
const [charIndex, setCharIndex] = useState(0);
const text = 'The quick brown fox jumps over the lazy dog near the riverbank.';
useEffect(() => {
const interval = setInterval(() => {
setCharIndex((p) => {
if (p >= text.length) return 0;
return p + 1;
});
}, 80);
return () => clearInterval(interval);
}, [text.length]);
return (
<div className="h-40 w-full flex flex-col items-center justify-center overflow-hidden rounded-md bg-app-darkerBox/50 p-4 gap-3">
{/* Fake waveform */}
<div className="flex items-center gap-[1px] h-6 w-full max-w-[180px] justify-center">
{Array.from({ length: 30 }, (_, i) => {
const h = 0.2 + 0.8 * Math.abs(Math.sin(i * 0.5 + charIndex * 0.1));
const active = i < (charIndex / text.length) * 30;
return (
<div
key={i}
className={`w-[3px] rounded-full transition-colors duration-100 ${
active ? 'bg-accent' : 'bg-app-line'
}`}
style={{ height: `${h * 100}%` }}
/>
);
})}
</div>
{/* Transcribed text */}
<div className="text-[10px] text-ink-dull font-mono max-w-[200px] text-center leading-relaxed min-h-[32px]">
{text.slice(0, charIndex)}
{charIndex < text.length && (
<span className="inline-block w-[2px] h-3 bg-accent animate-pulse ml-[1px] align-middle" />
)}
</div>
</div>
);
}
// ─── Animation: Unlimited Length ─────────────────────────────────────────────
function UnlimitedLengthAnimation() {
const [phase, setPhase] = useState(0);
const chunks = [
'The morning sun crept over the mountains, casting long shadows across the valley below.',
'Birds stirred in the canopy, their songs weaving through the cool air like threads of gold.',
'Far below, a river wound its way through ancient stones, carrying whispers of the night.',
];
useEffect(() => {
const interval = setInterval(() => {
setPhase((p) => (p + 1) % 4); // 0-2 = processing chunks, 3 = crossfade/done
}, 2000);
return () => clearInterval(interval);
}, []);
return (
<div className="h-40 w-full flex flex-col items-center justify-center overflow-hidden rounded-md bg-app-darkerBox/50 p-4 gap-2.5">
{/* Chunk pills */}
<div className="flex flex-col gap-1 w-full max-w-[220px]">
{chunks.map((chunk, i) => (
<motion.div
key={`chunk-${i}`}
className="flex items-center gap-1.5 px-2 py-1 rounded border text-[8px]"
animate={{
borderColor:
phase === 3
? 'hsl(43 50% 45% / 0.3)'
: i === phase
? 'hsl(43 50% 45% / 0.5)'
: i < phase
? 'rgba(255,255,255,0.12)'
: 'rgba(255,255,255,0.06)',
backgroundColor:
phase === 3
? 'hsl(43 50% 45% / 0.04)'
: i === phase
? 'hsl(43 50% 45% / 0.08)'
: i < phase
? 'rgba(255,255,255,0.04)'
: 'rgba(255,255,255,0.02)',
}}
transition={{ duration: 0.4 }}
>
{/* Status indicator */}
<motion.div
className="w-1.5 h-1.5 rounded-full shrink-0"
animate={{
backgroundColor:
phase === 3
? 'hsl(43 50% 50%)'
: i === phase
? 'hsl(43 50% 50%)'
: i < phase
? 'rgba(255,255,255,0.3)'
: 'rgba(255,255,255,0.1)',
boxShadow:
i === phase && phase < 3 ? '0 0 6px hsl(43 50% 50%)' : '0 0 0px transparent',
}}
transition={{ duration: 0.3 }}
/>
<span
className={`truncate font-mono ${
phase === 3 || i <= phase ? 'text-ink-dull' : 'text-ink-faint/50'
}`}
>
{chunk}
</span>
</motion.div>
))}
</div>
{/* Crossfade / result bar */}
<div className="flex items-center gap-1 w-full max-w-[220px]">
{chunks.map((_, i) => (
<motion.div
key={`seg-${i}`}
className="h-1.5 flex-1 rounded-full"
animate={{
backgroundColor:
phase === 3
? 'hsl(43 50% 45%)'
: i < phase
? 'rgba(255,255,255,0.2)'
: i === phase
? 'hsl(43 50% 45% / 0.5)'
: 'rgba(255,255,255,0.06)',
}}
transition={{ duration: 0.4 }}
/>
))}
</div>
{/* Status text */}
<motion.div
className="text-[9px] font-mono"
key={phase}
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
transition={{ duration: 0.3 }}
>
<span className={phase === 3 ? 'text-accent' : 'text-ink-faint'}>
{phase < 3
? `generating chunk ${phase + 1} of ${chunks.length}...`
: 'crossfaded & ready'}
</span>
</motion.div>
</div>
);
}
// ─── Feature data ───────────────────────────────────────────────────────────
const FEATURES = [
{
title: 'Near-Perfect Voice Cloning',
description:
'Multiple TTS engines for exceptional voice quality. Clone any voice from a few seconds of audio with natural intonation and emotion.',
icon: Mic,
animation: VoiceCloningAnimation,
},
{
title: 'Stories Editor',
description:
'Create multi-voice narratives with a timeline-based editor. Arrange tracks, trim clips, and mix conversations between characters.',
icon: AudioLines,
animation: StoriesAnimation,
},
{
title: 'Audio Effects Pipeline',
description:
'Apply pitch shift, reverb, delay, compression, and more — then save as presets. Preview effects live and set defaults per voice profile.',
icon: Sparkles,
animation: EffectsAnimation,
},
{
title: 'Local or Remote',
description:
'Run GPU inference locally with Metal, CUDA, ROCm, Intel Arc, or DirectML — or connect to a remote machine. One-click server setup with automatic discovery.',
icon: Cloud,
animation: LocalRemoteAnimation,
},
{
title: 'Audio Transcription',
description:
'Powered by Whisper for accurate speech-to-text. Automatically extract reference text from voice samples.',
icon: MessageSquareText,
animation: TranscriptionAnimation,
},
{
title: 'Unlimited Generation Length',
description:
'Generate up to 50,000 characters in one go. Text is auto-split at sentence boundaries, generated per-chunk, and crossfaded seamlessly.',
icon: TextCursorInput,
animation: UnlimitedLengthAnimation,
},
];
// ─── Feature Card ───────────────────────────────────────────────────────────
function FeatureCard({ feature }: { feature: (typeof FEATURES)[number] }) {
const Icon = feature.icon;
const Animation = feature.animation;
return (
<div className="rounded-lg border border-app-line bg-app-darkBox overflow-hidden">
<LazyLoad>
<div className="pointer-events-none select-none">
<Animation />
</div>
</LazyLoad>
<div className="p-5">
<div className="flex items-center gap-2 mb-2">
<Icon className="h-4 w-4 text-accent" />
<h3 className="text-[15px] font-medium text-foreground">{feature.title}</h3>
</div>
<p className="text-sm leading-relaxed text-muted-foreground">{feature.description}</p>
</div>
</div>
);
}
// ─── Features Section ───────────────────────────────────────────────────────
export function Features() {
return (
<section id="features" className="border-t border-border py-24">
<div className="mx-auto max-w-7xl px-6">
<div className="mb-16 text-center">
<h2 className="text-3xl font-semibold tracking-tight text-foreground md:text-4xl mb-4">
Professional voice tools, zero compromise
</h2>
<p className="text-muted-foreground max-w-2xl mx-auto">
Everything you need to clone voices, generate speech, and produce multi-voice content —
running entirely on your machine.
</p>
</div>
<div className="grid gap-6 md:grid-cols-2 lg:grid-cols-3">
{FEATURES.map((feature) => (
<FeatureCard key={feature.title} feature={feature} />
))}
</div>
</div>
</section>
);
}
+61 -24
View File
@@ -1,22 +1,33 @@
import Image from 'next/image';
import Link from 'next/link';
import { Separator } from '@/components/ui/separator';
import { GITHUB_REPO } from '@/lib/constants';
export function Footer() {
return (
<footer className="border-t border-border pt-8 sm:pt-12 pb-6 sm:pb-8 mt-12 sm:mt-16 md:mt-20">
<div className="container mx-auto px-4">
<div className="grid grid-cols-1 sm:grid-cols-2 md:grid-cols-3 gap-6 sm:gap-8 mb-6 sm:mb-8">
<div>
<h3 className="font-bold text-lg mb-4">voicebox</h3>
<p className="text-muted-foreground text-sm">
Professional voice cloning powered by Qwen3-TTS. Desktop app for Mac, Windows, and
Linux.
<footer className="border-t border-border py-12">
<div className="mx-auto max-w-7xl px-6">
<div className="grid grid-cols-1 sm:grid-cols-2 md:grid-cols-4 gap-8 mb-10">
{/* Brand */}
<div className="md:col-span-1">
<div className="flex items-center gap-2.5 mb-4">
<Image
src="/voicebox-logo-app.webp"
alt="Voicebox"
width={24}
height={24}
className="h-6 w-6"
/>
<span className="text-sm font-semibold">Voicebox</span>
</div>
<p className="text-sm text-muted-foreground leading-relaxed">
Open source voice cloning studio. Local-first, free forever.
</p>
</div>
{/* Product */}
<div>
<h4 className="font-semibold mb-3">Product</h4>
<ul className="space-y-2 text-muted-foreground text-sm">
<h4 className="text-sm font-semibold mb-3">Product</h4>
<ul className="space-y-2 text-sm text-muted-foreground">
<li>
<a href="#features" className="hover:text-foreground transition-colors">
Features
@@ -28,20 +39,17 @@ export function Footer() {
</a>
</li>
<li>
<Link
href={GITHUB_REPO}
target="_blank"
rel="noopener noreferrer"
className="hover:text-foreground transition-colors"
>
GitHub
</Link>
<a href="#about" className="hover:text-foreground transition-colors">
About
</a>
</li>
</ul>
</div>
{/* Resources */}
<div>
<h4 className="font-semibold mb-3">Resources</h4>
<ul className="space-y-2 text-muted-foreground text-sm">
<h4 className="text-sm font-semibold mb-3">Resources</h4>
<ul className="space-y-2 text-sm text-muted-foreground">
<li>
<Link
href={GITHUB_REPO}
@@ -74,10 +82,39 @@ export function Footer() {
</li>
</ul>
</div>
{/* Also by */}
<div>
<h4 className="text-sm font-semibold mb-3">Also By</h4>
<ul className="space-y-2 text-sm text-muted-foreground">
<li>
<a
href="https://spacebot.sh"
target="_blank"
rel="noopener noreferrer"
className="hover:text-foreground transition-colors"
>
Spacebot
</a>
</li>
<li>
<a
href="https://spacedrive.com"
target="_blank"
rel="noopener noreferrer"
className="hover:text-foreground transition-colors"
>
Spacedrive
</a>
</li>
</ul>
</div>
</div>
<Separator className="my-8" />
<div className="text-center text-muted-foreground text-sm space-y-2">
<p>© 2026 voicebox. All rights reserved.</p>
<div className="border-t border-border pt-6">
<p className="text-center text-sm text-muted-foreground">
&copy; {new Date().getFullYear()} Voicebox. Open source under MIT license.
</p>
</div>
</div>
</footer>
@@ -0,0 +1,298 @@
'use client';
import { Pause, Play, Repeat, Volume2, VolumeX } from 'lucide-react';
import { useCallback, useEffect, useRef, useState } from 'react';
import WaveSurfer from 'wavesurfer.js';
function formatDuration(seconds: number): string {
const m = Math.floor(seconds / 60);
const s = Math.floor(seconds % 60);
return `${m}:${s.toString().padStart(2, '0')}`;
}
// Shared ref so the unmute button can unlock WaveSurfer's audio on iOS Safari
// Must call .play() on WaveSurfer's actual media element during a user gesture
let sharedWaveSurfer: WaveSurfer | null = null;
let audioUnlocked = false;
export function unlockAudioContext() {
if (audioUnlocked) return;
audioUnlocked = true;
// Unlock WaveSurfer's internal audio element
// Skip if already playing — the context is already unlocked and the
// play/pause/reset dance would destroy the active playback.
if (sharedWaveSurfer && !sharedWaveSurfer.isPlaying()) {
const media = sharedWaveSurfer.getMediaElement();
if (media) {
media.muted = true;
media
.play()
.then(() => {
media.pause();
media.muted = false;
media.currentTime = 0;
})
.catch(() => {});
}
}
// Also unlock a standalone AudioContext as fallback
try {
const ctx = new (
window.AudioContext ||
(window as unknown as { webkitAudioContext: typeof AudioContext }).webkitAudioContext
)();
const buffer = ctx.createBuffer(1, 1, 22050);
const source = ctx.createBufferSource();
source.buffer = buffer;
source.connect(ctx.destination);
source.start(0);
} catch {
// Silently fail
}
}
interface LandingAudioPlayerProps {
audioUrl: string;
title: string;
playing: boolean;
muted: boolean;
onFinish: () => void;
onClose: () => void;
}
export function LandingAudioPlayer({
audioUrl,
title,
playing,
muted,
onFinish,
onClose,
}: LandingAudioPlayerProps) {
const waveformRef = useRef<HTMLDivElement>(null);
const wavesurferRef = useRef<WaveSurfer | null>(null);
const [isPlaying, setIsPlaying] = useState(false);
const [currentTime, setCurrentTime] = useState(0);
const [duration, setDuration] = useState(0);
const [volume, setVolume] = useState(0.75);
const [isLooping, setIsLooping] = useState(false);
const [isReady, setIsReady] = useState(false);
const onFinishRef = useRef(onFinish);
onFinishRef.current = onFinish;
const playingRef = useRef(playing);
playingRef.current = playing;
const mutedRef = useRef(muted);
mutedRef.current = muted;
// Initialize WaveSurfer
useEffect(() => {
const initWaveSurfer = () => {
const container = waveformRef.current;
if (!container) {
setTimeout(initWaveSurfer, 50);
return;
}
const rect = container.getBoundingClientRect();
if (rect.width === 0 || rect.height === 0) {
setTimeout(initWaveSurfer, 50);
return;
}
// Clean up existing instance
if (wavesurferRef.current) {
wavesurferRef.current.destroy();
wavesurferRef.current = null;
}
const root = document.documentElement;
const getCSSVar = (varName: string) => {
const value = getComputedStyle(root).getPropertyValue(varName).trim();
return value ? `hsl(${value})` : '';
};
const ws = WaveSurfer.create({
container,
waveColor: getCSSVar('--muted'),
progressColor: getCSSVar('--accent'),
cursorColor: getCSSVar('--accent'),
barWidth: 2,
barRadius: 2,
height: 80,
normalize: true,
interact: true,
mediaControls: false,
});
ws.on('ready', () => {
setDuration(ws.getDuration());
ws.setVolume(mutedRef.current ? 0 : volume);
setIsReady(true);
});
ws.on('play', () => {
console.log('[Player] play event');
setIsPlaying(true);
});
ws.on('pause', () => {
console.log('[Player] pause event');
setIsPlaying(false);
});
ws.on('timeupdate', (time: number) => {
setCurrentTime(Math.min(time, ws.getDuration()));
});
let didFinish = false;
ws.on('finish', () => {
if (didFinish) return;
didFinish = true;
console.log(
'[Player] finish event, currentTime:',
ws.getCurrentTime(),
'duration:',
ws.getDuration(),
);
setIsPlaying(false);
onFinishRef.current();
});
ws.load(audioUrl);
wavesurferRef.current = ws;
sharedWaveSurfer = ws;
};
setIsReady(false);
setCurrentTime(0);
setDuration(0);
requestAnimationFrame(() => {
requestAnimationFrame(() => {
setTimeout(initWaveSurfer, 10);
});
});
return () => {
if (wavesurferRef.current) {
wavesurferRef.current.destroy();
wavesurferRef.current = null;
}
};
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [audioUrl]);
// Respond to external play/stop signals
useEffect(() => {
const ws = wavesurferRef.current;
console.log('[Player] effect', { playing, isReady, hasWs: !!ws });
if (!ws || !isReady) return;
if (playing) {
// Resume the AudioContext first (required for iOS Safari after unlock)
const backend = ws.getMediaElement();
if (backend && 'context' in backend) {
const ctx = (backend as unknown as { context: AudioContext }).context;
if (ctx?.state === 'suspended') ctx.resume();
}
ws.play()
.then(() => {
console.log('[Player] play succeeded');
})
.catch((e: Error) => {
if (e.name === 'NotAllowedError') {
console.warn('[Player] Autoplay blocked by browser — waiting for user gesture');
} else {
console.error('[Player] play failed', e);
}
});
} else {
ws.pause();
}
}, [playing, isReady]);
// Sync volume and muted state
useEffect(() => {
if (wavesurferRef.current) {
wavesurferRef.current.setVolume(muted ? 0 : volume);
}
}, [volume, muted]);
const handlePlayPause = useCallback(() => {
if (!wavesurferRef.current) return;
wavesurferRef.current.playPause();
}, []);
return (
<div className="absolute bottom-0 left-0 right-0 border-t border-border bg-background/95 backdrop-blur supports-backdrop-filter:bg-background/60 z-30">
<div className="px-4 py-3 flex flex-col md:flex-row md:items-center gap-2 md:gap-4">
{/* Waveform — full width row on mobile, inline on desktop */}
<div className="min-w-0 min-h-[60px] md:min-h-[80px] md:flex-1 md:order-2">
<div ref={waveformRef} className="w-full h-full min-h-[60px] md:min-h-[80px]" />
</div>
{/* Controls row */}
<div className="flex items-center gap-3 md:contents">
{/* Play/Pause */}
<button
onClick={handlePlayPause}
disabled={!isReady}
className="h-10 w-10 rounded-full bg-accent flex items-center justify-center shrink-0 disabled:opacity-50 md:order-1 shadow-lg"
>
{isPlaying ? (
<Pause className="h-5 w-5 text-accent-foreground fill-accent-foreground" />
) : (
<Play className="h-5 w-5 ml-0.5 text-accent-foreground fill-accent-foreground" />
)}
</button>
{/* Time */}
<div className="flex items-center gap-1 text-sm text-muted-foreground shrink-0 md:order-3">
<span className="font-mono text-xs">{formatDuration(currentTime)}</span>
<span className="text-xs">/</span>
<span className="font-mono text-xs">{formatDuration(duration)}</span>
</div>
{/* Title */}
{title && (
<div className="text-sm font-medium truncate max-w-[200px] shrink-0 hidden lg:block md:order-4">
{title}
</div>
)}
{/* Loop */}
<button
onClick={() => setIsLooping(!isLooping)}
className={`h-8 w-8 flex items-center justify-center rounded-sm shrink-0 hover:bg-muted md:order-5 ${
isLooping ? 'text-foreground' : 'text-muted-foreground'
}`}
>
<Repeat className="h-4 w-4" />
</button>
{/* Volume */}
<div className="flex items-center gap-2 shrink-0 w-[140px] md:order-6 mr-3">
<button
onClick={() => setVolume(volume > 0 ? 0 : 0.75)}
className="h-8 w-8 flex items-center justify-center hover:bg-muted rounded-sm"
>
{volume > 0 ? (
<Volume2 className="h-4 w-4 text-muted-foreground" />
) : (
<VolumeX className="h-4 w-4 text-muted-foreground" />
)}
</button>
<input
type="range"
min={0}
max={100}
value={volume * 100}
onChange={(e) => setVolume(Number(e.target.value) / 100)}
className="flex-1 h-1 appearance-none bg-muted rounded-full accent-foreground cursor-pointer [&::-webkit-slider-thumb]:appearance-none [&::-webkit-slider-thumb]:h-3 [&::-webkit-slider-thumb]:w-3 [&::-webkit-slider-thumb]:rounded-full [&::-webkit-slider-thumb]:bg-foreground"
/>
</div>
</div>
</div>
</div>
);
}
+88
View File
@@ -0,0 +1,88 @@
'use client';
import { Github } from 'lucide-react';
import Image from 'next/image';
import { useEffect, useState } from 'react';
import { GITHUB_REPO } from '@/lib/constants';
function formatStarCount(count: number): string {
if (count >= 1000) {
const k = count / 1000;
return k % 1 === 0 ? `${k}k` : `${k.toFixed(1)}k`;
}
return count.toString();
}
export function Navbar() {
const [starCount, setStarCount] = useState<number | null>(null);
useEffect(() => {
fetch('/api/stars')
.then((res) => {
if (!res.ok) throw new Error('Failed to fetch stars');
return res.json();
})
.then((data) => {
if (typeof data.count === 'number') setStarCount(data.count);
})
.catch((error) => {
console.error('Failed to fetch star count:', error);
});
}, []);
return (
<nav className="fixed inset-x-0 top-0 z-50 border-b border-border/50 bg-background/80 backdrop-blur-xl">
<div className="mx-auto flex max-w-7xl items-center justify-between px-6 py-3">
{/* Logo + wordmark */}
<a href="/" className="flex items-center gap-2.5">
<Image
src="/voicebox-logo-app.webp"
alt="Voicebox"
width={28}
height={28}
className="h-7 w-7"
/>
<span className="text-[15px] font-semibold text-foreground">Voicebox</span>
</a>
{/* Nav links */}
<div className="hidden sm:flex items-center gap-1">
<a
href="#features"
className="rounded-md px-3 py-1.5 text-sm font-medium text-muted-foreground transition-colors hover:text-foreground"
>
Features
</a>
<a
href="#about"
className="rounded-md px-3 py-1.5 text-sm font-medium text-muted-foreground transition-colors hover:text-foreground"
>
About
</a>
<a
href="#download"
className="rounded-md px-3 py-1.5 text-sm font-medium text-muted-foreground transition-colors hover:text-foreground"
>
Download
</a>
</div>
{/* GitHub star button */}
<a
href={GITHUB_REPO}
target="_blank"
rel="noopener noreferrer"
className="flex items-center gap-2 rounded-lg border border-border/60 bg-card/60 px-3 py-1.5 text-sm text-muted-foreground transition-colors hover:text-foreground hover:border-border"
>
<Github className="h-4 w-4" />
<span className="text-[13px] font-medium">Star</span>
{starCount !== null && (
<span className="border-l border-border/60 pl-2 text-[13px] font-semibold text-foreground">
{formatStarCount(starCount)}
</span>
)}
</a>
</div>
</nav>
);
}
+479
View File
@@ -0,0 +1,479 @@
'use client';
import { AnimatePresence, motion } from 'framer-motion';
import { Mic, Monitor, Upload } from 'lucide-react';
import { useEffect, useMemo, useState } from 'react';
// ─── Waveform bars generator ────────────────────────────────────────────────
function generateWaveformBars(count: number, seed: number): number[] {
const bars: number[] = [];
for (let i = 0; i < count; i++) {
const x = i / count;
// Speech-like envelope: ramp up, sustain, taper
const envelope = Math.sin(x * Math.PI) * 0.8 + 0.2;
// Layered pseudo-random noise
const n1 = Math.sin(seed * 127.1 + i * 43.7) * 0.5 + 0.5;
const n2 = Math.sin(seed * 269.5 + i * 17.3) * 0.3 + 0.5;
const n3 = Math.sin(seed * 53.9 + i * 97.1) * 0.2 + 0.5;
const noise = (n1 + n2 + n3) / 3;
bars.push(envelope * noise);
}
return bars;
}
// ─── Animated waveform background ───────────────────────────────────────────
function WaveformBackground({ active }: { active: boolean }) {
const bars = useMemo(() => generateWaveformBars(60, 42), []);
return (
<div className="absolute inset-0 pointer-events-none flex items-end justify-center overflow-hidden">
<div className="flex items-end gap-[2px] w-full h-full px-4 pb-4">
{bars.map((h, i) => {
const maxH = 120; // max bar height in px
const baseH = 4;
const activeH = baseH + h * maxH;
const idleH = baseH + h * maxH * 0.25;
return (
<motion.div
key={i}
className="flex-1 rounded-full bg-accent"
animate={{
opacity: active ? 0.35 : 0.1,
height: active ? [idleH, activeH, idleH * 1.5, activeH * 0.7, idleH] : idleH,
}}
transition={
active
? {
duration: 1.0 + (i % 5) * 0.12,
repeat: Infinity,
repeatType: 'mirror',
delay: (i % 7) * 0.04,
ease: 'easeInOut',
}
: { duration: 0.6 }
}
/>
);
})}
</div>
</div>
);
}
// ─── Tab content panels ─────────────────────────────────────────────────────
function UploadPanel() {
const [hasFile, setHasFile] = useState(false);
useEffect(() => {
// Simulate file drop after 2s
const t1 = setTimeout(() => setHasFile(true), 2000);
const t2 = setTimeout(() => setHasFile(false), 5000);
return () => {
clearTimeout(t1);
clearTimeout(t2);
};
}, []);
return (
<div
className={`relative flex flex-col items-center justify-center gap-3 p-6 border-2 rounded-lg min-h-[180px] transition-colors duration-300 ${
hasFile ? 'border-accent bg-accent/5' : 'border-dashed border-muted-foreground/25'
}`}
>
<AnimatePresence mode="wait">
{!hasFile ? (
<motion.div
key="idle"
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0 }}
className="flex flex-col items-center gap-3"
>
<div className="h-10 px-5 rounded-md bg-accent text-accent-foreground flex items-center gap-2 text-sm font-medium">
<Upload className="h-4 w-4" />
Choose File
</div>
<p className="text-xs text-muted-foreground text-center">
Drag and drop an audio file, or click to browse.
<br />
Maximum duration: 30 seconds.
</p>
</motion.div>
) : (
<motion.div
key="file"
initial={{ opacity: 0, scale: 0.95 }}
animate={{ opacity: 1, scale: 1 }}
exit={{ opacity: 0 }}
className="flex flex-col items-center gap-3"
>
<div className="flex items-center gap-2">
<Upload className="h-4 w-4 text-accent" />
<span className="text-sm font-medium">sample-voice-clip.wav</span>
</div>
<div className="flex gap-2">
<div className="h-8 px-3 rounded-md border border-border flex items-center gap-1.5 text-xs text-muted-foreground">
<span>0:04</span>
</div>
<div className="h-8 px-3 rounded-md border border-border flex items-center gap-1.5 text-xs text-muted-foreground">
<Mic className="h-3 w-3" />
Transcribe
</div>
</div>
</motion.div>
)}
</AnimatePresence>
</div>
);
}
function RecordPanel() {
const [state, setState] = useState<'idle' | 'recording' | 'done'>('idle');
const [elapsed, setElapsed] = useState(0);
useEffect(() => {
const t1 = setTimeout(() => setState('recording'), 1500);
const t2 = setTimeout(() => setState('done'), 5500);
const t3 = setTimeout(() => {
setState('idle');
setElapsed(0);
}, 8000);
return () => {
clearTimeout(t1);
clearTimeout(t2);
clearTimeout(t3);
};
}, []);
// Timer
useEffect(() => {
if (state !== 'recording') return;
setElapsed(0);
const interval = setInterval(() => setElapsed((e) => e + 1), 1000);
return () => clearInterval(interval);
}, [state]);
const formatTime = (s: number) => `${Math.floor(s / 60)}:${(s % 60).toString().padStart(2, '0')}`;
return (
<div
className={`relative flex flex-col items-center justify-center gap-3 p-6 border-2 rounded-lg min-h-[180px] overflow-hidden transition-colors duration-300 ${
state === 'recording'
? 'border-accent bg-accent/5'
: state === 'done'
? 'border-accent bg-accent/5'
: 'border-dashed border-muted-foreground/25'
}`}
>
<WaveformBackground active={state === 'recording'} />
<AnimatePresence mode="wait">
{state === 'idle' && (
<motion.div
key="idle"
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0 }}
className="relative z-10 flex flex-col items-center gap-3"
>
<div className="h-10 px-5 rounded-md bg-accent text-accent-foreground flex items-center gap-2 text-sm font-medium">
<Mic className="h-4 w-4" />
Start Recording
</div>
<p className="text-xs text-muted-foreground text-center">
Click to record from your microphone.
<br />
Maximum duration: 30 seconds.
</p>
</motion.div>
)}
{state === 'recording' && (
<motion.div
key="recording"
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0 }}
className="relative z-10 flex flex-col items-center gap-3"
>
<div className="flex items-center gap-3">
<div className="h-3 w-3 rounded-full bg-accent animate-pulse" />
<span className="text-lg font-mono font-semibold">{formatTime(elapsed)}</span>
</div>
<div className="h-9 px-4 rounded-md bg-accent text-accent-foreground flex items-center gap-2 text-sm font-medium">
<div className="h-3 w-3 rounded-sm bg-accent-foreground" />
Stop Recording
</div>
<p className="text-xs text-muted-foreground">{formatTime(30 - elapsed)} remaining</p>
</motion.div>
)}
{state === 'done' && (
<motion.div
key="done"
initial={{ opacity: 0, scale: 0.95 }}
animate={{ opacity: 1, scale: 1 }}
exit={{ opacity: 0 }}
className="relative z-10 flex flex-col items-center gap-3"
>
<div className="flex items-center gap-2">
<Mic className="h-4 w-4 text-accent" />
<span className="text-sm font-medium">Recording complete</span>
</div>
<div className="flex gap-2">
<div className="h-8 px-3 rounded-md border border-border flex items-center gap-1.5 text-xs text-muted-foreground">
<span>0:04</span>
</div>
<div className="h-8 px-3 rounded-md border border-border flex items-center gap-1.5 text-xs text-muted-foreground">
<Mic className="h-3 w-3" />
Transcribe
</div>
</div>
</motion.div>
)}
</AnimatePresence>
</div>
);
}
function SystemPanel() {
const [state, setState] = useState<'idle' | 'capturing' | 'done'>('idle');
const [elapsed, setElapsed] = useState(0);
useEffect(() => {
const t1 = setTimeout(() => setState('capturing'), 1500);
const t2 = setTimeout(() => setState('done'), 5500);
const t3 = setTimeout(() => {
setState('idle');
setElapsed(0);
}, 8000);
return () => {
clearTimeout(t1);
clearTimeout(t2);
clearTimeout(t3);
};
}, []);
useEffect(() => {
if (state !== 'capturing') return;
setElapsed(0);
const interval = setInterval(() => setElapsed((e) => e + 1), 1000);
return () => clearInterval(interval);
}, [state]);
const formatTime = (s: number) => `${Math.floor(s / 60)}:${(s % 60).toString().padStart(2, '0')}`;
return (
<div
className={`relative flex flex-col items-center justify-center gap-3 p-6 border-2 rounded-lg min-h-[180px] overflow-hidden transition-colors duration-300 ${
state === 'capturing'
? 'border-accent bg-accent/5'
: state === 'done'
? 'border-accent bg-accent/5'
: 'border-dashed border-muted-foreground/25'
}`}
>
<WaveformBackground active={state === 'capturing'} />
<AnimatePresence mode="wait">
{state === 'idle' && (
<motion.div
key="idle"
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0 }}
className="relative z-10 flex flex-col items-center gap-3"
>
<div className="h-10 px-5 rounded-md bg-accent text-accent-foreground flex items-center gap-2 text-sm font-medium">
<Monitor className="h-4 w-4" />
Start Capture
</div>
<p className="text-xs text-muted-foreground text-center">
Capture audio playing on your system.
<br />
Maximum duration: 30 seconds.
</p>
</motion.div>
)}
{state === 'capturing' && (
<motion.div
key="capturing"
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0 }}
className="relative z-10 flex flex-col items-center gap-3"
>
<div className="flex items-center gap-3">
<div className="h-3 w-3 rounded-full bg-accent animate-pulse" />
<span className="text-lg font-mono font-semibold">{formatTime(elapsed)}</span>
</div>
<div className="h-9 px-4 rounded-md bg-accent text-accent-foreground flex items-center gap-2 text-sm font-medium">
<div className="h-3 w-3 rounded-sm bg-accent-foreground" />
Stop Capture
</div>
<p className="text-xs text-muted-foreground">{formatTime(30 - elapsed)} remaining</p>
</motion.div>
)}
{state === 'done' && (
<motion.div
key="done"
initial={{ opacity: 0, scale: 0.95 }}
animate={{ opacity: 1, scale: 1 }}
exit={{ opacity: 0 }}
className="relative z-10 flex flex-col items-center gap-3"
>
<div className="flex items-center gap-2">
<Monitor className="h-4 w-4 text-accent" />
<span className="text-sm font-medium">Capture complete</span>
</div>
<div className="flex gap-2">
<div className="h-8 px-3 rounded-md border border-border flex items-center gap-1.5 text-xs text-muted-foreground">
<span>0:04</span>
</div>
<div className="h-8 px-3 rounded-md border border-border flex items-center gap-1.5 text-xs text-muted-foreground">
<Mic className="h-3 w-3" />
Transcribe
</div>
</div>
</motion.div>
)}
</AnimatePresence>
</div>
);
}
// ─── Tab selector ───────────────────────────────────────────────────────────
const TABS = [
{ id: 'upload' as const, label: 'Upload', icon: Upload },
{ id: 'record' as const, label: 'Microphone', icon: Mic },
{ id: 'system' as const, label: 'System Audio', icon: Monitor },
];
type TabId = (typeof TABS)[number]['id'];
// ─── Main section ───────────────────────────────────────────────────────────
export function VoiceCreator() {
const [activeTab, setActiveTab] = useState<TabId>('record');
const [cycleKey, setCycleKey] = useState(0);
// Auto-cycle tabs
useEffect(() => {
const tabOrder: TabId[] = ['record', 'upload', 'system'];
let idx = tabOrder.indexOf(activeTab);
const interval = setInterval(() => {
idx = (idx + 1) % tabOrder.length;
setActiveTab(tabOrder[idx]);
setCycleKey((k) => k + 1);
}, 9000);
return () => clearInterval(interval);
}, [activeTab]);
return (
<section className="border-t border-border py-24">
<div className="mx-auto max-w-5xl px-6">
<div className="grid grid-cols-1 md:grid-cols-2 gap-12 md:gap-16 items-center">
{/* Left: Copy */}
<div>
<h2 className="text-3xl font-semibold tracking-tight text-foreground md:text-4xl mb-4">
Clone any voice in seconds
</h2>
<p className="text-muted-foreground mb-6">
Three ways to capture a voice sample. Upload a clip, record from your microphone, or
capture audio playing on your system. Voicebox clones the voice from as little as 3
seconds of audio.
</p>
<div className="space-y-3">
<div className="flex items-start gap-3">
<div className="h-8 w-8 rounded-lg bg-accent/10 flex items-center justify-center shrink-0 mt-0.5">
<Upload className="h-4 w-4 text-accent" />
</div>
<div>
<div className="text-sm font-medium">Upload a clip</div>
<div className="text-xs text-muted-foreground">
Drag and drop any audio file — WAV, MP3, FLAC, or WebM.
</div>
</div>
</div>
<div className="flex items-start gap-3">
<div className="h-8 w-8 rounded-lg bg-accent/10 flex items-center justify-center shrink-0 mt-0.5">
<Mic className="h-4 w-4 text-accent" />
</div>
<div>
<div className="text-sm font-medium">Record from microphone</div>
<div className="text-xs text-muted-foreground">
Live waveform preview while you record. Up to 30 seconds.
</div>
</div>
</div>
<div className="flex items-start gap-3">
<div className="h-8 w-8 rounded-lg bg-accent/10 flex items-center justify-center shrink-0 mt-0.5">
<Monitor className="h-4 w-4 text-accent" />
</div>
<div>
<div className="text-sm font-medium">System audio capture</div>
<div className="text-xs text-muted-foreground">
Clone a voice from a YouTube video, podcast, or any app playing audio.
</div>
</div>
</div>
</div>
</div>
{/* Right: Animated UI mock */}
<div className="rounded-xl border border-app-line bg-app-darkBox overflow-hidden pointer-events-none select-none">
<div className="p-5">
{/* Tab bar */}
<div className="flex rounded-lg border border-border bg-card/50 p-1 mb-4">
{TABS.map((tab) => {
const Icon = tab.icon;
const isActive = activeTab === tab.id;
return (
<button
key={tab.id}
onClick={() => {
setActiveTab(tab.id);
setCycleKey((k) => k + 1);
}}
className={`flex-1 flex items-center justify-center gap-1.5 px-3 py-1.5 rounded-md text-xs font-medium transition-colors ${
isActive
? 'bg-background text-foreground shadow-sm'
: 'text-muted-foreground hover:text-foreground'
}`}
>
<Icon className="h-3.5 w-3.5" />
<span className="hidden sm:inline">{tab.label}</span>
</button>
);
})}
</div>
{/* Panel */}
<AnimatePresence mode="wait">
<motion.div
key={`${activeTab}-${cycleKey}`}
initial={{ opacity: 0, y: 6 }}
animate={{ opacity: 1, y: 0 }}
exit={{ opacity: 0, y: -6 }}
transition={{ duration: 0.2 }}
>
{activeTab === 'upload' && <UploadPanel />}
{activeTab === 'record' && <RecordPanel />}
{activeTab === 'system' && <SystemPanel />}
</motion.div>
</AnimatePresence>
</div>
</div>
</div>
</div>
</section>
);
}
+97 -2
View File
@@ -9,6 +9,7 @@ export interface DownloadLinks {
export interface ReleaseInfo {
version: string;
downloadLinks: DownloadLinks;
totalDownloads: number;
}
const GITHUB_REPO = 'jamiepine/voicebox';
@@ -17,7 +18,11 @@ const GITHUB_API_BASE = 'https://api.github.com';
// Cache for release info (in-memory cache, resets on server restart)
let cachedReleaseInfo: ReleaseInfo | null = null;
let cacheTimestamp: number = 0;
const CACHE_DURATION = 1000 * 60 * 10; // 10 minutes
const CACHE_DURATION = 1000 * 60 * 5; // 5 minutes
// Cache for star count
let cachedStarCount: number | null = null;
let starCacheTimestamp: number = 0;
/**
* Fetches the latest release from GitHub and extracts download links
@@ -31,7 +36,7 @@ export async function getLatestRelease(): Promise<ReleaseInfo> {
try {
const response = await fetch(`${GITHUB_API_BASE}/repos/${GITHUB_REPO}/releases/latest`, {
next: { revalidate: 600 }, // Revalidate every 10 minutes
cache: 'no-store',
headers: {
Accept: 'application/vnd.github.v3+json',
},
@@ -68,11 +73,15 @@ export async function getLatestRelease(): Promise<ReleaseInfo> {
}
}
// Fetch total downloads across ALL releases
const totalDownloads = await getTotalDownloads();
// Fallback: construct URLs if not found in assets
const baseUrl = `https://github.com/${GITHUB_REPO}/releases/download/${version}`;
const releaseInfo: ReleaseInfo = {
version,
totalDownloads,
downloadLinks: {
macArm: downloadLinks.macArm || `${baseUrl}/voicebox_aarch64.app.tar.gz`,
macIntel: downloadLinks.macIntel || `${baseUrl}/voicebox_x64.app.tar.gz`,
@@ -92,3 +101,89 @@ export async function getLatestRelease(): Promise<ReleaseInfo> {
throw error;
}
}
// Cache for total download count
let cachedTotalDownloads: number | null = null;
let downloadsCacheTimestamp: number = 0;
/**
* Fetches download counts across ALL releases (paginated)
*/
async function getTotalDownloads(): Promise<number> {
const now = Date.now();
if (cachedTotalDownloads !== null && now - downloadsCacheTimestamp < CACHE_DURATION) {
return cachedTotalDownloads;
}
let total = 0;
let page = 1;
try {
while (true) {
const response = await fetch(
`${GITHUB_API_BASE}/repos/${GITHUB_REPO}/releases?per_page=100&page=${page}`,
{
cache: 'no-store',
headers: { Accept: 'application/vnd.github.v3+json' },
},
);
if (!response.ok) break;
const releases = await response.json();
if (!Array.isArray(releases) || releases.length === 0) break;
for (const release of releases) {
for (const asset of release.assets || []) {
total += asset.download_count || 0;
}
}
if (releases.length < 100) break;
page++;
}
cachedTotalDownloads = total;
downloadsCacheTimestamp = now;
} catch (error) {
console.error('Failed to fetch total downloads:', error);
if (cachedTotalDownloads !== null) return cachedTotalDownloads;
}
return total;
}
/**
* Fetches the star count for the repo from GitHub
*/
export async function getStarCount(): Promise<number> {
const now = Date.now();
if (cachedStarCount !== null && now - starCacheTimestamp < CACHE_DURATION) {
return cachedStarCount;
}
try {
const response = await fetch(`${GITHUB_API_BASE}/repos/${GITHUB_REPO}`, {
next: { revalidate: 600 },
headers: {
Accept: 'application/vnd.github.v3+json',
},
});
if (!response.ok) {
throw new Error(`GitHub API error: ${response.status}`);
}
const repo = await response.json();
const count = repo.stargazers_count ?? 0;
cachedStarCount = count;
starCacheTimestamp = now;
return count;
} catch (error) {
console.error('Failed to fetch star count:', error);
if (cachedStarCount !== null) return cachedStarCount;
throw error;
}
}
+27
View File
@@ -8,6 +8,9 @@ module.exports = {
],
theme: {
extend: {
fontFamily: {
sans: ['var(--font-sans)', 'system-ui', 'sans-serif'],
},
colors: {
border: 'hsl(var(--border))',
input: 'hsl(var(--input))',
@@ -33,6 +36,9 @@ module.exports = {
accent: {
DEFAULT: 'hsl(var(--accent))',
foreground: 'hsl(var(--accent-foreground))',
faint: 'hsl(var(--accent-faint))',
deep: 'hsl(var(--accent-deep))',
glow: 'hsl(var(--accent-glow))',
},
popover: {
DEFAULT: 'hsl(var(--popover))',
@@ -42,6 +48,27 @@ module.exports = {
DEFAULT: 'hsl(var(--card))',
foreground: 'hsl(var(--card-foreground))',
},
// App surface tokens
app: {
DEFAULT: 'hsl(var(--app))',
box: 'hsl(var(--app-box))',
darkBox: 'hsl(var(--app-dark-box))',
darkerBox: 'hsl(var(--app-darker-box))',
lightBox: 'hsl(var(--app-light-box))',
line: 'hsl(var(--app-line))',
button: 'hsl(var(--app-button))',
hover: 'hsl(var(--app-hover))',
selected: 'hsl(var(--app-selected))',
},
ink: {
DEFAULT: 'hsl(var(--ink))',
dull: 'hsl(var(--ink-dull))',
faint: 'hsl(var(--ink-faint))',
},
sidebar: {
DEFAULT: 'hsl(var(--sidebar))',
line: 'hsl(var(--sidebar-line))',
},
},
borderRadius: {
lg: 'var(--radius)',
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "voicebox",
"version": "0.2.0",
"version": "0.2.3",
"private": true,
"workspaces": [
"app",
+3 -1
View File
@@ -9,12 +9,14 @@ PLATFORM=$(rustc --print host-tuple 2>/dev/null || echo "unknown")
echo "Building voicebox-server for platform: $PLATFORM"
# Build Python binary
# Resolve PATH to absolute paths before changing directory
export PATH="$(cd "$(dirname "$0")/.." && pwd)/backend/venv/bin:$PATH"
cd backend
# Check if PyInstaller is installed
if ! python -c "import PyInstaller" 2>/dev/null; then
echo "Installing PyInstaller..."
pip install pyinstaller
python -m pip install pyinstaller
fi
# Build binary
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "@voicebox/tauri",
"private": true,
"version": "0.2.0",
"version": "0.2.3",
"type": "module",
"scripts": {
"dev": "vite",
+1 -1
View File
@@ -5041,7 +5041,7 @@ checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a"
[[package]]
name = "voicebox"
version = "0.2.0"
version = "0.2.1"
dependencies = [
"base64 0.22.1",
"core-foundation-sys",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "voicebox"
version = "0.2.0"
version = "0.2.3"
description = "A production-quality desktop app for Qwen3-TTS voice cloning and generation"
authors = ["you"]
license = ""
Binary file not shown.
Binary file not shown.
+9 -2
View File
@@ -572,6 +572,7 @@ async fn restart_server(
#[command]
fn set_keep_server_running(state: State<'_, ServerState>, keep_running: bool) {
println!("set_keep_server_running called with: {}", keep_running);
*state.keep_running_on_close.lock().unwrap() = keep_running;
}
@@ -762,6 +763,8 @@ pub fn run() {
RunEvent::Exit => {
let state = app.state::<ServerState>();
let keep_running = *state.keep_running_on_close.lock().unwrap();
let has_pid = state.server_pid.lock().unwrap().is_some();
println!("RunEvent::Exit — keep_running={}, has_pid={}", keep_running, has_pid);
if keep_running {
// Tell the server to disable its watchdog so it survives
@@ -771,9 +774,13 @@ pub fn run() {
.timeout(std::time::Duration::from_secs(2))
.build()
.unwrap();
let _ = client
match client
.post(&format!("http://127.0.0.1:{}/watchdog/disable", SERVER_PORT))
.send();
.send()
{
Ok(resp) => println!("Watchdog disable response: {}", resp.status()),
Err(e) => eprintln!("Failed to disable watchdog: {}", e),
}
} else {
// Server will self-terminate via parent-pid watchdog when
// this process exits. On Unix, also send SIGTERM for
+2 -2
View File
@@ -1,7 +1,7 @@
{
"$schema": "https://schema.tauri.app/config/2",
"productName": "Voicebox",
"version": "0.2.0",
"version": "0.2.3",
"identifier": "sh.voicebox.app",
"build": {
"beforeDevCommand": "bun run dev",
@@ -12,7 +12,7 @@
"bundle": {
"active": true,
"targets": "all",
"createUpdaterArtifacts": false,
"createUpdaterArtifacts": "v1Compatible",
"externalBin": ["binaries/voicebox-server"],
"icon": [
"icons/32x32.png",
+6
View File
@@ -64,6 +64,12 @@ class TauriLifecycle implements PlatformLifecycle {
// @ts-expect-error - accessing module-level variable from another module
const serverStartedByApp = window.__voiceboxServerStartedByApp ?? false;
console.log(
'[lifecycle] window-close-requested: keepRunning=%s, serverStartedByApp=%s',
keepRunning,
serverStartedByApp,
);
if (!keepRunning && serverStartedByApp) {
// Stop server before closing (only if we started it)
try {
+5 -1
View File
@@ -64,13 +64,17 @@ class TauriUpdater implements PlatformUpdater {
}
this.notifySubscribers();
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
// Tauri updater throws on 404 / no published release / network errors.
// Treat "no update available" style errors as up-to-date, not failures.
const isNoUpdate = /404|not found|no update|up.to.date/i.test(message);
this.status = {
checking: false,
available: false,
downloading: false,
installing: false,
readyToInstall: false,
error: error instanceof Error ? error.message : 'Failed to check for updates',
error: isNoUpdate ? undefined : message,
};
this.notifySubscribers();
}
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "@voicebox/web",
"private": true,
"version": "0.2.0",
"version": "0.2.3",
"type": "module",
"scripts": {
"dev": "vite",