mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-09-15 04:40:40 -07:00
- Add backend/STYLE_GUIDE.md covering formatting, imports, types, docstrings, comments, error handling, async, logging, and naming conventions - Add pyproject.toml with ruff linter/formatter config (ERA, FIX, isort, pyupgrade) - Extract generation service (Phase 3): unified run_generation() replaces three duplicated closures, serial queue moved to services/task_queue.py - Delete Makefile in favor of justfile; update all references - Add Python lint/format/test commands to justfile (check-python, fix-python, test) - Install ruff, pytest, pytest-asyncio as dev tools in setup-python - Update REFACTOR_PLAN.md with Phase 3 and Phase 7 completion
3.5 KiB
3.5 KiB
Changelog
All notable changes to Voicebox will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Fixed
- Profile Name Validation - Added proper validation to prevent duplicate profile names (#134)
- Users now receive clear error messages when attempting to create or update profiles with duplicate names
- Improved error handling in create and update profile API endpoints
- Added comprehensive test suite for duplicate name validation
0.1.0 - 2026-01-25
Added
Core Features
- Voice Cloning - Clone voices from audio samples using Qwen3-TTS (1.7B and 0.6B models)
- Voice Profile Management - Create, edit, and organize voice profiles with multiple samples
- Speech Generation - Generate high-quality speech from text using cloned voices
- Generation History - Track all generations with search and filtering capabilities
- Audio Transcription - Automatic transcription powered by Whisper
- In-App Recording - Record audio samples directly in the app with waveform visualization
Desktop App
- Tauri Desktop App - Native desktop application for macOS, Windows, and Linux
- Local Server Mode - Embedded Python server runs automatically
- Remote Server Mode - Connect to a remote Voicebox server on your network
- Auto-Updates - Automatic update notifications and installation
API
- REST API - Full REST API for voice synthesis and profile management
- OpenAPI Documentation - Interactive API docs at
/docsendpoint - Type-Safe Client - Auto-generated TypeScript client from OpenAPI schema
Technical
- Voice Prompt Caching - Fast regeneration with cached voice prompts
- Multi-Sample Support - Combine multiple audio samples for better voice quality
- GPU/CPU/MPS Support - Automatic device detection and optimization
- Model Management - Lazy loading and VRAM management
- SQLite Database - Local data persistence
Technical Details
- Built with Tauri v2 (Rust + React)
- FastAPI backend with async Python
- TypeScript frontend with React Query and Zustand
- Qwen3-TTS for voice cloning
- Whisper for transcription
Platform Support
- macOS (Apple Silicon and Intel)
- Windows
- Linux (AppImage)
[Unreleased]
Fixed
- Audio export failing when Tauri save dialog returns object instead of string path
- OpenAPI client generator script now documents the local backend port and avoids an unused loop variable warning
Added
- justfile - Comprehensive development workflow automation with commands for setup, development, building, testing, and code quality checks
- Cross-platform support (macOS, Linux, Windows)
- Python version detection and compatibility warnings
- Self-documenting help system with
just --list
Changed
- README - Updated Quick Start with justfile-based setup instructions
Removed
- Makefile - Replaced by justfile (cross-platform, simpler syntax)
[Unreleased - Planned]
Planned
- Real-time streaming synthesis
- Conversation mode with multiple speakers
- Voice effects (pitch shift, reverb, M3GAN-style)
- Timeline-based audio editor
- Additional voice models (XTTS, Bark)
- Voice design from text descriptions
- Project system for saving sessions
- Plugin architecture