mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-09-28 06:35:18 -07:00
Compare commits
81
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e5f4606a6c | ||
|
|
146ef5aaeb | ||
|
|
971604d14f | ||
|
|
953e6ec7d8 | ||
|
|
d3c65fc6c2 | ||
|
|
b6e772c6ac | ||
|
|
a6b070201b | ||
|
|
30352e2419 | ||
|
|
bfa38b36b7 | ||
|
|
1b66a528d1 | ||
|
|
bef4092e6e | ||
|
|
9654f7b642 | ||
|
|
eba1244add | ||
|
|
94487f32a5 | ||
|
|
081f45e680 | ||
|
|
86768288ce | ||
|
|
0fd063442a | ||
|
|
a0c2493e98 | ||
|
|
b39f48cc81 | ||
|
|
4ff775bc98 | ||
|
|
6351aa75e9 | ||
|
|
43873a883b | ||
|
|
60012b81c0 | ||
|
|
89f3127c37 | ||
|
|
ef3c3a7f8c | ||
|
|
7b5e73cfa8 | ||
|
|
462f104494 | ||
|
|
fadb57164e | ||
|
|
e870d65136 | ||
|
|
3df40278cc | ||
|
|
deeef5a474 | ||
|
|
236e464525 | ||
|
|
cf3cf3f002 | ||
|
|
9d98e1e768 | ||
|
|
3be8980f48 | ||
|
|
76bc070f5b | ||
|
|
01838f4773 | ||
|
|
39e4f9d08c | ||
|
|
341d71470c | ||
|
|
bb6cea24ba | ||
|
|
fa7ac88abc | ||
|
|
c68ddc45b1 | ||
|
|
f89dc66d0c | ||
|
|
cba7d7bc23 | ||
|
|
3c89b068f3 | ||
|
|
229841e05e | ||
|
|
123e8215e4 | ||
|
|
2a3afec2ca | ||
|
|
99ddd5a0b4 | ||
|
|
8d730621bc | ||
|
|
e23118f610 | ||
|
|
d4bfdc0d68 | ||
|
|
116c108906 | ||
|
|
2d23c8e06a | ||
|
|
3973a59ba3 | ||
|
|
d9aa75253a | ||
|
|
b22bf36565 | ||
|
|
33f4ed9b44 | ||
|
|
cc37e04221 | ||
|
|
2b4fbe5173 | ||
|
|
423d69b7cc | ||
|
|
ea943876dc | ||
|
|
b55d8cc567 | ||
|
|
036d90dc8e | ||
|
|
c513451277 | ||
|
|
27ae6dfbab | ||
|
|
51b9e2fd3d | ||
|
|
be25ddbe0e | ||
|
|
9cd4921291 | ||
|
|
2349bd24ba | ||
|
|
cd82ed0664 | ||
|
|
c4884a0443 | ||
|
|
232d231788 | ||
|
|
1cf90c81dd | ||
|
|
3204e193fa | ||
|
|
9d5d6cb56a | ||
|
|
153eaba5f3 | ||
|
|
3370e3b419 | ||
|
|
615bd188a0 | ||
|
|
7208f51eee | ||
|
|
07a91a2381 |
+4
-4
@@ -1,5 +1,5 @@
|
|||||||
[bumpversion]
|
[bumpversion]
|
||||||
current_version = 0.1.4
|
current_version = 0.1.11
|
||||||
commit = True
|
commit = True
|
||||||
tag = True
|
tag = True
|
||||||
tag_name = v{new_version}
|
tag_name = v{new_version}
|
||||||
@@ -34,6 +34,6 @@ replace = "version": "{new_version}"
|
|||||||
search = "version": "{current_version}"
|
search = "version": "{current_version}"
|
||||||
replace = "version": "{new_version}"
|
replace = "version": "{new_version}"
|
||||||
|
|
||||||
[bumpversion:file:backend/main.py]
|
[bumpversion:file:backend/__init__.py]
|
||||||
search = "version": "{current_version}"
|
search = __version__ = "{current_version}"
|
||||||
replace = "version": "{new_version}"
|
replace = __version__ = "{new_version}"
|
||||||
|
|||||||
@@ -17,15 +17,19 @@ jobs:
|
|||||||
- platform: 'macos-latest'
|
- platform: 'macos-latest'
|
||||||
args: '--target aarch64-apple-darwin'
|
args: '--target aarch64-apple-darwin'
|
||||||
python-version: '3.12'
|
python-version: '3.12'
|
||||||
|
backend: 'mlx'
|
||||||
- platform: 'macos-15-intel'
|
- platform: 'macos-15-intel'
|
||||||
args: '--target x86_64-apple-darwin'
|
args: '--target x86_64-apple-darwin'
|
||||||
python-version: '3.12'
|
python-version: '3.12'
|
||||||
|
backend: 'pytorch'
|
||||||
# - platform: 'ubuntu-22.04'
|
# - platform: 'ubuntu-22.04'
|
||||||
# args: ''
|
# args: ''
|
||||||
# python-version: '3.12'
|
# python-version: '3.12'
|
||||||
|
# backend: 'pytorch'
|
||||||
- platform: 'windows-latest'
|
- platform: 'windows-latest'
|
||||||
args: ''
|
args: ''
|
||||||
python-version: '3.12'
|
python-version: '3.12'
|
||||||
|
backend: 'pytorch'
|
||||||
|
|
||||||
runs-on: ${{ matrix.platform }}
|
runs-on: ${{ matrix.platform }}
|
||||||
|
|
||||||
@@ -57,6 +61,11 @@ jobs:
|
|||||||
pip install pyinstaller
|
pip install pyinstaller
|
||||||
pip install -r backend/requirements.txt
|
pip install -r backend/requirements.txt
|
||||||
|
|
||||||
|
- name: Install MLX dependencies (Apple Silicon only)
|
||||||
|
if: matrix.backend == 'mlx'
|
||||||
|
run: |
|
||||||
|
pip install -r backend/requirements-mlx.txt
|
||||||
|
|
||||||
- name: Build Python server (Linux/macOS)
|
- name: Build Python server (Linux/macOS)
|
||||||
if: matrix.platform != 'windows-latest'
|
if: matrix.platform != 'windows-latest'
|
||||||
run: |
|
run: |
|
||||||
@@ -133,7 +142,8 @@ jobs:
|
|||||||
See the assets below to download and install this version.
|
See the assets below to download and install this version.
|
||||||
|
|
||||||
### Installation
|
### Installation
|
||||||
- **macOS**: Download the `.dmg` file
|
- **macOS (Apple Silicon)**: Download the `aarch64.dmg` file - uses MLX for fast native inference
|
||||||
|
- **macOS (Intel)**: Download the `x64.dmg` file - uses PyTorch
|
||||||
- **Windows**: Download the `.msi` installer
|
- **Windows**: Download the `.msi` installer
|
||||||
- **Linux**: Download the `.AppImage` or `.deb` package
|
- **Linux**: Download the `.AppImage` or `.deb` package
|
||||||
|
|
||||||
|
|||||||
@@ -53,6 +53,20 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|||||||
|
|
||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
|
||||||
|
### Added
|
||||||
|
- **Makefile** - Comprehensive development workflow automation with commands for setup, development, building, testing, and code quality checks
|
||||||
|
- Includes Python version detection and compatibility warnings
|
||||||
|
- Self-documenting help system with `make help`
|
||||||
|
- Colored output for better readability
|
||||||
|
- Supports parallel development server execution
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
- **README** - Added Makefile reference and updated Quick Start with Makefile-based setup instructions alongside manual setup
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## [Unreleased - Planned]
|
||||||
|
|
||||||
### Planned
|
### Planned
|
||||||
- Real-time streaming synthesis
|
- Real-time streaming synthesis
|
||||||
- Conversation mode with multiple speakers
|
- Conversation mode with multiple speakers
|
||||||
|
|||||||
+65
-17
@@ -32,6 +32,10 @@ Thank you for your interest in contributing to Voicebox! This document provides
|
|||||||
|
|
||||||
### Development Setup
|
### Development Setup
|
||||||
|
|
||||||
|
**Using the Makefile (recommended for macOS/Linux):** Run `make setup` to install all dependencies, then `make dev` to start development servers. See `make help` for all available commands.
|
||||||
|
|
||||||
|
**Manual setup (required for Windows):**
|
||||||
|
|
||||||
1. **Fork and clone the repository**
|
1. **Fork and clone the repository**
|
||||||
```bash
|
```bash
|
||||||
git clone https://github.com/YOUR_USERNAME/voicebox.git
|
git clone https://github.com/YOUR_USERNAME/voicebox.git
|
||||||
@@ -62,37 +66,43 @@ Thank you for your interest in contributing to Voicebox! This document provides
|
|||||||
# Install Python dependencies
|
# Install Python dependencies
|
||||||
pip install -r requirements.txt
|
pip install -r requirements.txt
|
||||||
|
|
||||||
|
# Install MLX dependencies (Apple Silicon only - for faster inference)
|
||||||
|
# On Apple Silicon, this enables native Metal acceleration
|
||||||
|
if [[ $(uname -m) == "arm64" ]]; then
|
||||||
|
pip install -r requirements-mlx.txt
|
||||||
|
fi
|
||||||
|
|
||||||
# Install Qwen3-TTS (required for voice synthesis)
|
# Install Qwen3-TTS (required for voice synthesis)
|
||||||
pip install git+https://github.com/QwenLM/Qwen3-TTS.git
|
pip install git+https://github.com/QwenLM/Qwen3-TTS.git
|
||||||
```
|
```
|
||||||
|
|
||||||
4. **Initialize database**
|
4. **Start development servers**
|
||||||
```bash
|
|
||||||
cd backend
|
|
||||||
python -c "from database import init_db; init_db()"
|
|
||||||
```
|
|
||||||
This creates the SQLite database at `data/voicebox.db`.
|
|
||||||
|
|
||||||
5. **Start development servers**
|
Development requires two terminals: one for the Python backend, one for the Tauri app.
|
||||||
|
|
||||||
**Terminal 1: Backend server**
|
**Terminal 1: Backend server** (start this first)
|
||||||
```bash
|
```bash
|
||||||
cd backend
|
cd backend
|
||||||
source venv/bin/activate # Activate venv if not already active
|
source venv/bin/activate # Activate venv if not already active
|
||||||
bun run dev:server
|
bun run dev:server
|
||||||
# Or manually: uvicorn main:app --reload --port 8000
|
# Or manually: uvicorn main:app --reload --port 17493
|
||||||
```
|
```
|
||||||
Backend will be available at `http://localhost:8000`
|
Backend will be available at `http://localhost:17493`
|
||||||
|
|
||||||
**Terminal 2: Desktop app**
|
**Terminal 2: Desktop app**
|
||||||
```bash
|
```bash
|
||||||
bun run dev
|
bun run dev
|
||||||
```
|
```
|
||||||
This will:
|
This will:
|
||||||
|
- Create a placeholder sidecar binary (for Tauri compilation)
|
||||||
- Start Vite dev server on port 5173
|
- Start Vite dev server on port 5173
|
||||||
- Launch Tauri window pointing to localhost:5173
|
- Launch Tauri window pointing to localhost:5173
|
||||||
|
- Connect to the Python server you started in Terminal 1
|
||||||
- Enable hot reload
|
- Enable hot reload
|
||||||
|
|
||||||
|
> **Note:** In dev mode, the app connects to your manually-started Python server.
|
||||||
|
> The bundled server binary is only used in production builds.
|
||||||
|
|
||||||
**Optional: Web app**
|
**Optional: Web app**
|
||||||
```bash
|
```bash
|
||||||
bun run dev:web
|
bun run dev:web
|
||||||
@@ -109,18 +119,36 @@ First-time usage will be slower due to model downloads, but subsequent runs will
|
|||||||
|
|
||||||
### Building
|
### Building
|
||||||
|
|
||||||
**Build Python server binary:**
|
**Build everything (recommended):**
|
||||||
```bash
|
```bash
|
||||||
|
bun run build
|
||||||
|
```
|
||||||
|
This automatically:
|
||||||
|
1. Builds the Python server binary (`./scripts/build-server.sh`)
|
||||||
|
2. Builds the Tauri desktop app (`cd tauri && bun run tauri build`)
|
||||||
|
|
||||||
|
Creates platform-specific installers (`.dmg`, `.msi`, `.AppImage`) in `tauri/src-tauri/target/release/bundle/`.
|
||||||
|
|
||||||
|
**Note:** The build process detects your platform and includes the appropriate backend (MLX for Apple Silicon, PyTorch for others).
|
||||||
|
|
||||||
|
**Build server binary only:**
|
||||||
|
```bash
|
||||||
|
bun run build:server
|
||||||
|
# or
|
||||||
./scripts/build-server.sh
|
./scripts/build-server.sh
|
||||||
```
|
```
|
||||||
Creates platform-specific binary in `tauri/src-tauri/binaries/`
|
Creates platform-specific binary in `tauri/src-tauri/binaries/`
|
||||||
|
|
||||||
**Build Tauri desktop app:**
|
**Building with local Qwen3-TTS development version:**
|
||||||
|
|
||||||
|
If you're actively developing or modifying the Qwen3-TTS library, set the `QWEN_TTS_PATH` environment variable to point to your local clone:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd tauri
|
export QWEN_TTS_PATH=~/path/to/your/Qwen3-TTS
|
||||||
bun run tauri build
|
bun run build:server
|
||||||
```
|
```
|
||||||
Creates platform-specific installers (`.dmg`, `.msi`, `.AppImage`)
|
|
||||||
|
This makes PyInstaller use your local qwen-tts version instead of the pip-installed package. Useful when testing changes to the TTS library before they're published to PyPI or when using an editable install (`pip install -e`).
|
||||||
|
|
||||||
**Build web app:**
|
**Build web app:**
|
||||||
```bash
|
```bash
|
||||||
@@ -137,6 +165,26 @@ After starting the backend server:
|
|||||||
```
|
```
|
||||||
This downloads the OpenAPI schema and generates the TypeScript client in `app/src/lib/api/`
|
This downloads the OpenAPI schema and generates the TypeScript client in `app/src/lib/api/`
|
||||||
|
|
||||||
|
### Convert Assets to Web Formats
|
||||||
|
|
||||||
|
To optimize images and videos for the web, run:
|
||||||
|
```bash
|
||||||
|
bun run convert:assets
|
||||||
|
```
|
||||||
|
|
||||||
|
This script:
|
||||||
|
- Converts PNG → WebP (better compression, same quality)
|
||||||
|
- Converts MOV → WebM (VP9 codec, smaller file size)
|
||||||
|
- Processes files in `landing/public/` and `docs/public/`
|
||||||
|
- **Deletes original files** after successful conversion
|
||||||
|
|
||||||
|
**Requirements:** Install `webp` and `ffmpeg`:
|
||||||
|
```bash
|
||||||
|
brew install webp ffmpeg
|
||||||
|
```
|
||||||
|
|
||||||
|
> **Note:** Run this before committing new images or videos to keep the repository size small.
|
||||||
|
|
||||||
## Development Workflow
|
## Development Workflow
|
||||||
|
|
||||||
### 1. Create a Branch
|
### 1. Create a Branch
|
||||||
|
|||||||
@@ -0,0 +1,245 @@
|
|||||||
|
# Voicebox Makefile
|
||||||
|
# Unix-only (macOS/Linux). Windows users should use WSL.
|
||||||
|
|
||||||
|
SHELL := /bin/bash
|
||||||
|
.DEFAULT_GOAL := help
|
||||||
|
|
||||||
|
# Directories
|
||||||
|
BACKEND_DIR := backend
|
||||||
|
TAURI_DIR := tauri
|
||||||
|
WEB_DIR := web
|
||||||
|
APP_DIR := app
|
||||||
|
|
||||||
|
# Python (prefer 3.12, fallback to 3.13, then python3)
|
||||||
|
PYTHON := $(shell command -v python3.12 2>/dev/null || command -v python3.13 2>/dev/null || echo python3)
|
||||||
|
VENV := $(CURDIR)/$(BACKEND_DIR)/venv
|
||||||
|
VENV_BIN := $(VENV)/bin
|
||||||
|
PIP := $(VENV_BIN)/pip
|
||||||
|
PYTHON_VENV := $(VENV_BIN)/python
|
||||||
|
|
||||||
|
# Colors for output
|
||||||
|
BLUE := \033[0;34m
|
||||||
|
GREEN := \033[0;32m
|
||||||
|
YELLOW := \033[0;33m
|
||||||
|
NC := \033[0m # No Color
|
||||||
|
|
||||||
|
.PHONY: help
|
||||||
|
help: ## Show this help message
|
||||||
|
@echo -e "$(BLUE)Voicebox$(NC) - Development Commands"
|
||||||
|
@echo ""
|
||||||
|
@grep -E '^[a-zA-Z_-]+:.*?## .*$$' $(MAKEFILE_LIST) | sort | \
|
||||||
|
awk 'BEGIN {FS = ":.*?## "}; {printf " $(GREEN)%-20s$(NC) %s\n", $$1, $$2}'
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# SETUP
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: setup setup-js setup-python setup-rust
|
||||||
|
|
||||||
|
setup: setup-js setup-python ## Full project setup (all dependencies)
|
||||||
|
@echo -e "$(GREEN)✓ Setup complete!$(NC)"
|
||||||
|
@echo -e " Run $(YELLOW)make dev$(NC) to start development servers"
|
||||||
|
|
||||||
|
setup-js: ## Install JavaScript dependencies (bun)
|
||||||
|
@echo -e "$(BLUE)Installing JavaScript dependencies...$(NC)"
|
||||||
|
bun install
|
||||||
|
|
||||||
|
setup-python: $(VENV)/bin/activate ## Set up Python virtual environment and dependencies
|
||||||
|
@echo -e "$(BLUE)Installing Python dependencies...$(NC)"
|
||||||
|
$(PIP) install --upgrade pip
|
||||||
|
$(PIP) install -r $(BACKEND_DIR)/requirements.txt
|
||||||
|
@if [ "$$(uname -m)" = "arm64" ] && [ "$$(uname)" = "Darwin" ]; then \
|
||||||
|
echo -e "$(BLUE)Detected Apple Silicon - installing MLX dependencies...$(NC)"; \
|
||||||
|
$(PIP) install -r $(BACKEND_DIR)/requirements-mlx.txt; \
|
||||||
|
echo -e "$(GREEN)✓ MLX backend enabled (native Metal acceleration)$(NC)"; \
|
||||||
|
fi
|
||||||
|
$(PIP) install git+https://github.com/QwenLM/Qwen3-TTS.git
|
||||||
|
@echo -e "$(GREEN)✓ Python environment ready$(NC)"
|
||||||
|
|
||||||
|
$(VENV)/bin/activate:
|
||||||
|
@echo -e "$(BLUE)Creating Python virtual environment...$(NC)"
|
||||||
|
@PY_MINOR=$$($(PYTHON) -c "import sys; print(sys.version_info[1])"); \
|
||||||
|
if [ "$$PY_MINOR" -gt 13 ]; then \
|
||||||
|
echo -e "$(YELLOW)Warning: Python 3.$$PY_MINOR detected. ML packages may not be compatible.$(NC)"; \
|
||||||
|
echo -e "$(YELLOW)Recommended: Use Python 3.12 or 3.13 (brew install [email protected])$(NC)"; \
|
||||||
|
fi
|
||||||
|
$(PYTHON) -m venv $(VENV)
|
||||||
|
|
||||||
|
setup-rust: ## Install Rust toolchain (if not present)
|
||||||
|
@command -v rustc >/dev/null 2>&1 || curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# DEVELOPMENT
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: dev dev-backend dev-frontend dev-web kill-dev
|
||||||
|
|
||||||
|
dev: ## Start backend + desktop app (parallel)
|
||||||
|
@echo -e "$(BLUE)Starting development servers...$(NC)"
|
||||||
|
@echo -e "$(YELLOW)Note: If Tauri fails, run 'make build-server' first or use separate terminals$(NC)"
|
||||||
|
@trap 'kill 0' EXIT; \
|
||||||
|
$(MAKE) dev-backend & \
|
||||||
|
sleep 2 && $(MAKE) dev-frontend & \
|
||||||
|
wait
|
||||||
|
|
||||||
|
dev-backend: ## Start FastAPI backend server
|
||||||
|
@echo -e "$(BLUE)Starting backend server on http://localhost:17493$(NC)"
|
||||||
|
$(VENV_BIN)/uvicorn backend.main:app --reload --port 17493
|
||||||
|
|
||||||
|
dev-frontend: ## Start Tauri desktop app
|
||||||
|
@echo -e "$(BLUE)Starting Tauri desktop app...$(NC)"
|
||||||
|
bun run dev
|
||||||
|
|
||||||
|
dev-web: ## Start backend + web app (parallel)
|
||||||
|
@echo -e "$(BLUE)Starting web development servers...$(NC)"
|
||||||
|
@trap 'kill 0' EXIT; \
|
||||||
|
$(MAKE) dev-backend & \
|
||||||
|
sleep 2 && cd $(WEB_DIR) && bun run dev & \
|
||||||
|
wait
|
||||||
|
|
||||||
|
kill-dev: ## Kill all development processes
|
||||||
|
@echo -e "$(YELLOW)Killing development processes...$(NC)"
|
||||||
|
-pkill -f "uvicorn main:app" 2>/dev/null || true
|
||||||
|
-pkill -f "vite" 2>/dev/null || true
|
||||||
|
@echo -e "$(GREEN)✓ Processes killed$(NC)"
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# BUILD
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: build build-server build-tauri build-web
|
||||||
|
|
||||||
|
build: build-server build-tauri ## Build everything (server binary + desktop app)
|
||||||
|
@echo -e "$(GREEN)✓ Build complete!$(NC)"
|
||||||
|
|
||||||
|
build-server: ## Build Python server binary
|
||||||
|
@echo -e "$(BLUE)Building server binary...$(NC)"
|
||||||
|
PATH="$(VENV_BIN):$$PATH" ./scripts/build-server.sh
|
||||||
|
|
||||||
|
build-tauri: ## Build Tauri desktop app
|
||||||
|
@echo -e "$(BLUE)Building Tauri desktop app...$(NC)"
|
||||||
|
cd $(TAURI_DIR) && bun run tauri build
|
||||||
|
|
||||||
|
build-web: ## Build web app
|
||||||
|
@echo -e "$(BLUE)Building web app...$(NC)"
|
||||||
|
cd $(WEB_DIR) && bun run build
|
||||||
|
@echo -e "$(GREEN)✓ Web build output in $(WEB_DIR)/dist/$(NC)"
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# DATABASE & API
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: db-init db-reset generate-api
|
||||||
|
|
||||||
|
db-init: $(VENV)/bin/activate ## Initialize SQLite database
|
||||||
|
@echo -e "$(BLUE)Initializing database...$(NC)"
|
||||||
|
cd $(BACKEND_DIR) && $(PYTHON_VENV) -c "from database import init_db; init_db()"
|
||||||
|
@echo -e "$(GREEN)✓ Database created at $(BACKEND_DIR)/data/voicebox.db$(NC)"
|
||||||
|
|
||||||
|
db-reset: ## Reset database (delete and reinitialize)
|
||||||
|
@echo -e "$(YELLOW)Resetting database...$(NC)"
|
||||||
|
rm -f $(BACKEND_DIR)/data/voicebox.db
|
||||||
|
$(MAKE) db-init
|
||||||
|
|
||||||
|
generate-api: ## Generate TypeScript API client from OpenAPI schema
|
||||||
|
@echo -e "$(BLUE)Generating API client...$(NC)"
|
||||||
|
@echo -e "$(YELLOW)Note: Backend must be running (make dev-backend)$(NC)"
|
||||||
|
./scripts/generate-api.sh
|
||||||
|
@echo -e "$(GREEN)✓ API client generated in $(APP_DIR)/src/lib/api/$(NC)"
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# CODE QUALITY
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: lint format typecheck check
|
||||||
|
|
||||||
|
lint: ## Run linter (Biome)
|
||||||
|
@echo -e "$(BLUE)Linting...$(NC)"
|
||||||
|
bun run lint
|
||||||
|
|
||||||
|
format: ## Format code (Biome)
|
||||||
|
@echo -e "$(BLUE)Formatting...$(NC)"
|
||||||
|
bun run format
|
||||||
|
|
||||||
|
typecheck: ## Run TypeScript type checking
|
||||||
|
@echo -e "$(BLUE)Type checking...$(NC)"
|
||||||
|
bun run tsc --noEmit
|
||||||
|
|
||||||
|
check: ## Run all checks (Biome lint + format + type check)
|
||||||
|
@echo -e "$(BLUE)Running all checks...$(NC)"
|
||||||
|
bun run check
|
||||||
|
@echo -e "$(GREEN)✓ All checks passed$(NC)"
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# TESTING
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: test test-backend test-frontend
|
||||||
|
|
||||||
|
test: test-backend test-frontend ## Run all tests
|
||||||
|
@echo -e "$(GREEN)✓ All tests passed$(NC)"
|
||||||
|
|
||||||
|
test-backend: ## Run Python backend tests (requires pytest)
|
||||||
|
@echo -e "$(BLUE)Running backend tests...$(NC)"
|
||||||
|
@if [ -f "$(VENV_BIN)/pytest" ]; then \
|
||||||
|
cd $(BACKEND_DIR) && $(VENV_BIN)/pytest -v; \
|
||||||
|
else \
|
||||||
|
echo -e "$(YELLOW)pytest not installed. Run: $(PIP) install pytest$(NC)"; \
|
||||||
|
exit 1; \
|
||||||
|
fi
|
||||||
|
|
||||||
|
test-frontend: ## Run frontend tests (requires test script in package.json)
|
||||||
|
@echo -e "$(BLUE)Running frontend tests...$(NC)"
|
||||||
|
@if bun run test --help >/dev/null 2>&1; then \
|
||||||
|
bun run test; \
|
||||||
|
else \
|
||||||
|
echo -e "$(YELLOW)No test script configured$(NC)"; \
|
||||||
|
exit 1; \
|
||||||
|
fi
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# LOGS & DEBUGGING
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: logs docs
|
||||||
|
|
||||||
|
logs: ## Tail backend logs
|
||||||
|
@echo -e "$(BLUE)Tailing logs (Ctrl+C to stop)...$(NC)"
|
||||||
|
tail -f $(BACKEND_DIR)/logs/*.log 2>/dev/null || echo "No log files found"
|
||||||
|
|
||||||
|
docs: ## Open API documentation (backend must be running)
|
||||||
|
@echo -e "$(BLUE)Opening API docs...$(NC)"
|
||||||
|
open http://localhost:17493/docs 2>/dev/null || xdg-open http://localhost:17493/docs
|
||||||
|
|
||||||
|
# =============================================================================
|
||||||
|
# CLEAN
|
||||||
|
# =============================================================================
|
||||||
|
|
||||||
|
.PHONY: clean clean-python clean-build clean-all
|
||||||
|
|
||||||
|
clean: ## Clean build artifacts
|
||||||
|
@echo -e "$(BLUE)Cleaning build artifacts...$(NC)"
|
||||||
|
rm -rf $(TAURI_DIR)/src-tauri/target/release
|
||||||
|
rm -rf $(WEB_DIR)/dist
|
||||||
|
rm -rf $(APP_DIR)/dist
|
||||||
|
@echo -e "$(GREEN)✓ Build artifacts cleaned$(NC)"
|
||||||
|
|
||||||
|
clean-python: ## Clean Python cache and virtual environment
|
||||||
|
@echo -e "$(BLUE)Cleaning Python files...$(NC)"
|
||||||
|
rm -rf $(VENV)
|
||||||
|
find $(BACKEND_DIR) -type d -name "__pycache__" -exec rm -rf {} + 2>/dev/null || true
|
||||||
|
find $(BACKEND_DIR) -type f -name "*.pyc" -delete 2>/dev/null || true
|
||||||
|
@echo -e "$(GREEN)✓ Python environment cleaned$(NC)"
|
||||||
|
|
||||||
|
clean-build: ## Clean Rust/Tauri build cache
|
||||||
|
@echo -e "$(BLUE)Cleaning Rust build cache...$(NC)"
|
||||||
|
cd $(TAURI_DIR)/src-tauri && cargo clean
|
||||||
|
@echo -e "$(GREEN)✓ Rust cache cleaned$(NC)"
|
||||||
|
|
||||||
|
clean-all: clean clean-python clean-build ## Nuclear clean (everything)
|
||||||
|
@echo -e "$(BLUE)Cleaning node_modules...$(NC)"
|
||||||
|
rm -rf node_modules
|
||||||
|
rm -rf $(APP_DIR)/node_modules
|
||||||
|
rm -rf $(TAURI_DIR)/node_modules
|
||||||
|
rm -rf $(WEB_DIR)/node_modules
|
||||||
|
@echo -e "$(GREEN)✓ Full clean complete$(NC)"
|
||||||
@@ -10,6 +10,21 @@
|
|||||||
All running locally on your machine.
|
All running locally on your machine.
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<a href="https://github.com/jamiepine/voicebox/releases">
|
||||||
|
<img src="https://img.shields.io/github/downloads/jamiepine/voicebox/total?style=flat&color=blue" alt="Downloads" />
|
||||||
|
</a>
|
||||||
|
<a href="https://github.com/jamiepine/voicebox/releases/latest">
|
||||||
|
<img src="https://img.shields.io/github/v/release/jamiepine/voicebox?style=flat" alt="Release" />
|
||||||
|
</a>
|
||||||
|
<a href="https://github.com/jamiepine/voicebox/stargazers">
|
||||||
|
<img src="https://img.shields.io/github/stars/jamiepine/voicebox?style=flat" alt="Stars" />
|
||||||
|
</a>
|
||||||
|
<a href="https://github.com/jamiepine/voicebox/blob/main/LICENSE">
|
||||||
|
<img src="https://img.shields.io/github/license/jamiepine/voicebox?style=flat" alt="License" />
|
||||||
|
</a>
|
||||||
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<a href="https://voicebox.sh">voicebox.sh</a> •
|
<a href="https://voicebox.sh">voicebox.sh</a> •
|
||||||
<a href="#download">Download</a> •
|
<a href="#download">Download</a> •
|
||||||
@@ -22,7 +37,7 @@
|
|||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<a href="https://voicebox.sh">
|
<a href="https://voicebox.sh">
|
||||||
<img src=".github/assets/screenshot.webp" alt="Voicebox App Screenshot" width="800" />
|
<img src="landing/public/assets/app-screenshot-1.webp" alt="Voicebox App Screenshot" width="800" />
|
||||||
</a>
|
</a>
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
@@ -32,17 +47,30 @@
|
|||||||
|
|
||||||
<br/>
|
<br/>
|
||||||
|
|
||||||
## Why Voicebox?
|
<p align="center">
|
||||||
|
<img src="landing/public/assets/app-screenshot-2.webp" alt="Voicebox Screenshot 2" width="800" />
|
||||||
|
</p>
|
||||||
|
|
||||||
Voice AI is exploding, but most tools are either cloud-locked, expensive, or a nightmare to set up. Voicebox is different:
|
<p align="center">
|
||||||
|
<img src="landing/public/assets/app-screenshot-3.webp" alt="Voicebox Screenshot 3" width="800" />
|
||||||
|
</p>
|
||||||
|
|
||||||
- **100% Local** — Your voice data never leaves your machine
|
<br/>
|
||||||
- **Lightweight** — No bloated Electron, native Tauri performance
|
|
||||||
- **Fast** — Near-instant on CUDA, optimized for Apple Silicon
|
|
||||||
- **Flexible** — Use the app, integrate the API, or both
|
|
||||||
- **Open Source** — No subscriptions, no limits, no lock-in
|
|
||||||
|
|
||||||
Built with **Tauri** (Rust), **TypeScript**, **React**, and **Python**. Native performance meets modern DX.
|
## What is Voicebox?
|
||||||
|
|
||||||
|
Voicebox is a **local-first voice cloning studio** with DAW-like features for professional voice synthesis. Think of it as the **Ollama for voice** — download models, clone voices, and generate speech entirely on your machine.
|
||||||
|
|
||||||
|
Unlike cloud services that lock your voice data behind subscriptions, Voicebox gives you:
|
||||||
|
|
||||||
|
- **Complete privacy** — models and voice data stay on your machine
|
||||||
|
- **Professional tools** — multi-track timeline editor, audio trimming, conversation mixing
|
||||||
|
- **Model flexibility** — currently powered by Qwen3-TTS, with support for XTTS, Bark, and other models coming soon
|
||||||
|
- **API-first** — use the desktop app or integrate voice synthesis into your own projects
|
||||||
|
- **Native performance** — built with Tauri (Rust), not Electron
|
||||||
|
- **Super fast on Mac** — MLX backend with native Metal acceleration for 4-5x faster inference on Apple Silicon
|
||||||
|
|
||||||
|
Download a voice model, clone any voice from a few seconds of audio, and compose multi-voice projects with studio-grade editing tools. No Python install required, no cloud dependency, no limits.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -70,11 +98,13 @@ Powered by Alibaba's **Qwen3-TTS** — a breakthrough model that achieves near-p
|
|||||||
- **Instant cloning** — Upload a sample, get a voice profile
|
- **Instant cloning** — Upload a sample, get a voice profile
|
||||||
- **High fidelity** — Natural prosody, emotion, and cadence
|
- **High fidelity** — Natural prosody, emotion, and cadence
|
||||||
- **Multi-language** — English, Chinese, and more coming
|
- **Multi-language** — English, Chinese, and more coming
|
||||||
|
- **Lightning fast on Mac** — MLX backend leverages Apple Silicon's Neural Engine for super fast generation
|
||||||
|
|
||||||
### Voice Profile Management
|
### Voice Profile Management
|
||||||
|
|
||||||
- **Create profiles** from audio files or record directly in-app
|
- **Create profiles** from audio files or record directly in-app
|
||||||
- **Import/Export** profiles to share or backup
|
- **Import/Export** profiles to share or backup
|
||||||
|
- **Multi-sample support** — combine multiple samples for higher quality cloning
|
||||||
- **Organize** with descriptions and language tags
|
- **Organize** with descriptions and language tags
|
||||||
|
|
||||||
### Speech Generation
|
### Speech Generation
|
||||||
@@ -83,9 +113,19 @@ Powered by Alibaba's **Qwen3-TTS** — a breakthrough model that achieves near-p
|
|||||||
- **Batch generation** for long-form content
|
- **Batch generation** for long-form content
|
||||||
- **Smart caching** — regenerate instantly with voice prompt caching
|
- **Smart caching** — regenerate instantly with voice prompt caching
|
||||||
|
|
||||||
|
### Stories Editor
|
||||||
|
|
||||||
|
Create multi-voice narratives, podcasts, and conversations with a timeline-based editor.
|
||||||
|
|
||||||
|
- **Multi-track composition** — arrange multiple voice tracks in a single project
|
||||||
|
- **Inline audio editing** — trim and split clips directly in the timeline
|
||||||
|
- **Auto-playback** — preview stories with synchronized playhead
|
||||||
|
- **Voice mixing** — build conversations with multiple participants
|
||||||
|
|
||||||
### Recording & Transcription
|
### Recording & Transcription
|
||||||
|
|
||||||
- **In-app recording** with waveform visualization
|
- **In-app recording** with waveform visualization
|
||||||
|
- **System audio capture** — record desktop audio on macOS and Windows
|
||||||
- **Automatic transcription** powered by Whisper
|
- **Automatic transcription** powered by Whisper
|
||||||
- **Export recordings** in multiple formats
|
- **Export recordings** in multiple formats
|
||||||
|
|
||||||
@@ -109,17 +149,17 @@ Voicebox exposes a full REST API, so you can integrate voice synthesis into your
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Generate speech
|
# Generate speech
|
||||||
curl -X POST http://localhost:8000/api/generate \
|
curl -X POST http://localhost:8000/generate \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{"text": "Hello world", "profile_id": "abc123"}'
|
-d '{"text": "Hello world", "profile_id": "abc123", "language": "en"}'
|
||||||
|
|
||||||
# List voice profiles
|
# List voice profiles
|
||||||
curl http://localhost:8000/api/profiles
|
curl http://localhost:8000/profiles
|
||||||
|
|
||||||
# Create a profile from audio
|
# Create a profile
|
||||||
curl -X POST http://localhost:8000/api/profiles \
|
curl -X POST http://localhost:8000/profiles \
|
||||||
-F "[email protected]" \
|
-H "Content-Type: application/json" \
|
||||||
-F "name=My Voice"
|
-d '{"name": "My Voice", "language": "en"}'
|
||||||
```
|
```
|
||||||
|
|
||||||
**Use cases:**
|
**Use cases:**
|
||||||
@@ -142,8 +182,9 @@ Full API documentation available at `http://localhost:8000/docs` when running.
|
|||||||
| Frontend | React, TypeScript, Tailwind CSS |
|
| Frontend | React, TypeScript, Tailwind CSS |
|
||||||
| State | Zustand, React Query |
|
| State | Zustand, React Query |
|
||||||
| Backend | FastAPI (Python) |
|
| Backend | FastAPI (Python) |
|
||||||
| Voice Model | Qwen3-TTS |
|
| Voice Model | Qwen3-TTS (PyTorch or MLX) |
|
||||||
| Transcription | Whisper |
|
| Transcription | Whisper (PyTorch or MLX) |
|
||||||
|
| Inference Engine | MLX (Apple Silicon) / PyTorch (Windows/Linux/Intel) |
|
||||||
| Database | SQLite |
|
| Database | SQLite |
|
||||||
| Audio | WaveSurfer.js, librosa |
|
| Audio | WaveSurfer.js, librosa |
|
||||||
|
|
||||||
@@ -184,8 +225,26 @@ Voicebox aims to be the **one-stop shop for everything voice** — cloning, synt
|
|||||||
|
|
||||||
See [CONTRIBUTING.md](CONTRIBUTING.md) for detailed setup and contribution guidelines.
|
See [CONTRIBUTING.md](CONTRIBUTING.md) for detailed setup and contribution guidelines.
|
||||||
|
|
||||||
|
**Using the Makefile (recommended):** Run `make help` to see all available commands for setup, development, building, and testing.
|
||||||
|
|
||||||
### Quick Start
|
### Quick Start
|
||||||
|
|
||||||
|
**With Makefile (Unix/macOS/Linux):**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Clone the repo
|
||||||
|
git clone https://github.com/voicebox-sh/voicebox.git
|
||||||
|
cd voicebox
|
||||||
|
|
||||||
|
# Setup everything
|
||||||
|
make setup
|
||||||
|
|
||||||
|
# Start development
|
||||||
|
make dev
|
||||||
|
```
|
||||||
|
|
||||||
|
**Manual setup (all platforms):**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Clone the repo
|
# Clone the repo
|
||||||
git clone https://github.com/voicebox-sh/voicebox.git
|
git clone https://github.com/voicebox-sh/voicebox.git
|
||||||
@@ -201,7 +260,11 @@ cd backend && pip install -r requirements.txt && cd ..
|
|||||||
bun run dev
|
bun run dev
|
||||||
```
|
```
|
||||||
|
|
||||||
**Prerequisites:** [Bun](https://bun.sh), [Rust](https://rustup.rs), [Python 3.11+](https://python.org). CUDA-capable GPU recommended (CPU inference supported but slower).
|
**Prerequisites:** [Bun](https://bun.sh), [Rust](https://rustup.rs), [Python 3.11+](https://python.org).
|
||||||
|
|
||||||
|
**Performance:**
|
||||||
|
- **Apple Silicon (M1/M2/M3)**: Uses MLX backend with native Metal acceleration for 4-5x faster inference
|
||||||
|
- **Windows/Linux/Intel Mac**: Uses PyTorch backend (CUDA GPU recommended, CPU supported but slower)
|
||||||
|
|
||||||
### Project Structure
|
### Project Structure
|
||||||
|
|
||||||
|
|||||||
+6
-1
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"name": "@voicebox/app",
|
"name": "@voicebox/app",
|
||||||
"version": "0.1.4",
|
"version": "0.1.11",
|
||||||
"private": true,
|
"private": true,
|
||||||
"type": "module",
|
"type": "module",
|
||||||
"scripts": {
|
"scripts": {
|
||||||
@@ -13,6 +13,9 @@
|
|||||||
"check": "biome check --write src"
|
"check": "biome check --write src"
|
||||||
},
|
},
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
|
"@dnd-kit/core": "^6.3.1",
|
||||||
|
"@dnd-kit/sortable": "^10.0.0",
|
||||||
|
"@dnd-kit/utilities": "^3.2.2",
|
||||||
"@hookform/resolvers": "^3.9.0",
|
"@hookform/resolvers": "^3.9.0",
|
||||||
"@radix-ui/react-alert-dialog": "^1.1.1",
|
"@radix-ui/react-alert-dialog": "^1.1.1",
|
||||||
"@radix-ui/react-avatar": "^1.1.0",
|
"@radix-ui/react-avatar": "^1.1.0",
|
||||||
@@ -30,6 +33,7 @@
|
|||||||
"@radix-ui/react-toast": "^1.2.1",
|
"@radix-ui/react-toast": "^1.2.1",
|
||||||
"@tanstack/react-query": "^5.0.0",
|
"@tanstack/react-query": "^5.0.0",
|
||||||
"@tanstack/react-query-devtools": "^5.0.0",
|
"@tanstack/react-query-devtools": "^5.0.0",
|
||||||
|
"@tanstack/react-router": "^1.157.16",
|
||||||
"@tauri-apps/api": "^2.0.0",
|
"@tauri-apps/api": "^2.0.0",
|
||||||
"@tauri-apps/plugin-dialog": "^2.0.0",
|
"@tauri-apps/plugin-dialog": "^2.0.0",
|
||||||
"@tauri-apps/plugin-fs": "^2.0.0",
|
"@tauri-apps/plugin-fs": "^2.0.0",
|
||||||
@@ -44,6 +48,7 @@
|
|||||||
"react": "^18.3.0",
|
"react": "^18.3.0",
|
||||||
"react-dom": "^18.3.0",
|
"react-dom": "^18.3.0",
|
||||||
"react-hook-form": "^7.53.0",
|
"react-hook-form": "^7.53.0",
|
||||||
|
"react-sound-visualizer": "^1.4.0",
|
||||||
"tailwind-merge": "^2.5.4",
|
"tailwind-merge": "^2.5.4",
|
||||||
"wavesurfer.js": "^7.0.0",
|
"wavesurfer.js": "^7.0.0",
|
||||||
"zod": "^3.23.8",
|
"zod": "^3.23.8",
|
||||||
|
|||||||
+30
-91
@@ -1,31 +1,13 @@
|
|||||||
import { useEffect, useState } from 'react';
|
import { useEffect, useRef, useState } from 'react';
|
||||||
|
import { RouterProvider } from '@tanstack/react-router';
|
||||||
import voiceboxLogo from '@/assets/voicebox-logo.png';
|
import voiceboxLogo from '@/assets/voicebox-logo.png';
|
||||||
import { AppFrame } from '@/components/AppFrame/AppFrame';
|
|
||||||
import { AudioTab } from '@/components/AudioTab/AudioTab';
|
|
||||||
import { MainEditor } from '@/components/MainEditor/MainEditor';
|
|
||||||
import { ModelsTab } from '@/components/ModelsTab/ModelsTab';
|
|
||||||
import { ServerTab } from '@/components/ServerTab/ServerTab';
|
|
||||||
// import { GenerationForm } from '@/components/Generation/GenerationForm';
|
|
||||||
import ShinyText from '@/components/ShinyText';
|
import ShinyText from '@/components/ShinyText';
|
||||||
import { Sidebar } from '@/components/Sidebar';
|
|
||||||
import { TitleBarDragRegion } from '@/components/TitleBarDragRegion';
|
import { TitleBarDragRegion } from '@/components/TitleBarDragRegion';
|
||||||
import { Toaster } from '@/components/ui/toaster';
|
|
||||||
import { VoicesTab } from '@/components/VoicesTab/VoicesTab';
|
|
||||||
import { TOP_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
import { TOP_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
||||||
import { useModelDownloadToast } from '@/lib/hooks/useModelDownloadToast';
|
|
||||||
import { MODEL_DISPLAY_NAMES, useRestoreActiveTasks } from '@/lib/hooks/useRestoreActiveTasks';
|
|
||||||
import {
|
|
||||||
isMacOS,
|
|
||||||
isTauri,
|
|
||||||
setKeepServerRunning,
|
|
||||||
setupWindowCloseHandler,
|
|
||||||
startServer,
|
|
||||||
} from '@/lib/tauri';
|
|
||||||
import { cn } from '@/lib/utils/cn';
|
import { cn } from '@/lib/utils/cn';
|
||||||
|
import { router } from '@/router';
|
||||||
import { useServerStore } from '@/stores/serverStore';
|
import { useServerStore } from '@/stores/serverStore';
|
||||||
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
// Track if server is starting to prevent duplicate starts
|
|
||||||
let serverStarting = false;
|
|
||||||
|
|
||||||
const LOADING_MESSAGES = [
|
const LOADING_MESSAGES = [
|
||||||
'Warming up tensors...',
|
'Warming up tensors...',
|
||||||
@@ -51,32 +33,38 @@ const LOADING_MESSAGES = [
|
|||||||
];
|
];
|
||||||
|
|
||||||
function App() {
|
function App() {
|
||||||
const [activeTab, setActiveTab] = useState('main');
|
const platform = usePlatform();
|
||||||
const [serverReady, setServerReady] = useState(false);
|
const [serverReady, setServerReady] = useState(false);
|
||||||
const [loadingMessageIndex, setLoadingMessageIndex] = useState(0);
|
const [loadingMessageIndex, setLoadingMessageIndex] = useState(0);
|
||||||
|
const serverStartingRef = useRef(false);
|
||||||
// Monitor active downloads/generations and show toasts for them
|
|
||||||
const activeDownloads = useRestoreActiveTasks();
|
|
||||||
|
|
||||||
// Sync stored setting to Rust on startup
|
// Sync stored setting to Rust on startup
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (isTauri()) {
|
if (platform.metadata.isTauri) {
|
||||||
const keepRunning = useServerStore.getState().keepServerRunningOnClose;
|
const keepRunning = useServerStore.getState().keepServerRunningOnClose;
|
||||||
setKeepServerRunning(keepRunning).catch((error) => {
|
platform.lifecycle.setKeepServerRunning(keepRunning).catch((error) => {
|
||||||
console.error('Failed to sync initial setting to Rust:', error);
|
console.error('Failed to sync initial setting to Rust:', error);
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
}, []);
|
}, [platform]);
|
||||||
|
|
||||||
|
// Setup lifecycle callbacks
|
||||||
|
useEffect(() => {
|
||||||
|
platform.lifecycle.onServerReady = () => {
|
||||||
|
setServerReady(true);
|
||||||
|
};
|
||||||
|
}, [platform]);
|
||||||
|
|
||||||
// Setup window close handler and auto-start server when running in Tauri (production only)
|
// Setup window close handler and auto-start server when running in Tauri (production only)
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!isTauri()) {
|
if (!platform.metadata.isTauri) {
|
||||||
|
setServerReady(true); // Web assumes server is running
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Setup window close handler to check setting and stop server if needed
|
// Setup window close handler to check setting and stop server if needed
|
||||||
// This works in both dev and prod, but will only stop server if it was started by the app
|
// This works in both dev and prod, but will only stop server if it was started by the app
|
||||||
setupWindowCloseHandler().catch((error) => {
|
platform.lifecycle.setupWindowCloseHandler().catch((error) => {
|
||||||
console.error('Failed to setup window close handler:', error);
|
console.error('Failed to setup window close handler:', error);
|
||||||
});
|
});
|
||||||
|
|
||||||
@@ -92,14 +80,15 @@ function App() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Auto-start server in production
|
// Auto-start server in production
|
||||||
if (serverStarting) {
|
if (serverStartingRef.current) {
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
serverStarting = true;
|
serverStartingRef.current = true;
|
||||||
console.log('Production mode: Starting bundled server...');
|
console.log('Production mode: Starting bundled server...');
|
||||||
|
|
||||||
startServer(false)
|
platform.lifecycle
|
||||||
|
.startServer(false)
|
||||||
.then((serverUrl) => {
|
.then((serverUrl) => {
|
||||||
console.log('Server is ready at:', serverUrl);
|
console.log('Server is ready at:', serverUrl);
|
||||||
// Update the server URL in the store with the dynamically assigned port
|
// Update the server URL in the store with the dynamically assigned port
|
||||||
@@ -111,7 +100,7 @@ function App() {
|
|||||||
})
|
})
|
||||||
.catch((error) => {
|
.catch((error) => {
|
||||||
console.error('Failed to auto-start server:', error);
|
console.error('Failed to auto-start server:', error);
|
||||||
serverStarting = false;
|
serverStartingRef.current = false;
|
||||||
// @ts-expect-error - adding property to window
|
// @ts-expect-error - adding property to window
|
||||||
window.__voiceboxServerStartedByApp = false;
|
window.__voiceboxServerStartedByApp = false;
|
||||||
});
|
});
|
||||||
@@ -120,13 +109,13 @@ function App() {
|
|||||||
// Note: Window close is handled separately in Tauri Rust code
|
// Note: Window close is handled separately in Tauri Rust code
|
||||||
return () => {
|
return () => {
|
||||||
// Window close event handles server shutdown based on setting
|
// Window close event handles server shutdown based on setting
|
||||||
serverStarting = false;
|
serverStartingRef.current = false;
|
||||||
};
|
};
|
||||||
}, []);
|
}, [platform]);
|
||||||
|
|
||||||
// Cycle through loading messages every 3 seconds
|
// Cycle through loading messages every 3 seconds
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!isTauri() || serverReady) {
|
if (!platform.metadata.isTauri || serverReady) {
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -135,10 +124,10 @@ function App() {
|
|||||||
}, 3000);
|
}, 3000);
|
||||||
|
|
||||||
return () => clearInterval(interval);
|
return () => clearInterval(interval);
|
||||||
}, [serverReady]);
|
}, [serverReady, platform.metadata.isTauri]);
|
||||||
|
|
||||||
// Show loading screen while server is starting in Tauri
|
// Show loading screen while server is starting in Tauri
|
||||||
if (isTauri() && !serverReady) {
|
if (platform.metadata.isTauri && !serverReady) {
|
||||||
return (
|
return (
|
||||||
<div
|
<div
|
||||||
className={cn(
|
className={cn(
|
||||||
@@ -172,57 +161,7 @@ function App() {
|
|||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
return (
|
return <RouterProvider router={router} />;
|
||||||
<AppFrame>
|
|
||||||
<div className="flex flex-1 min-h-0 overflow-hidden">
|
|
||||||
<Sidebar activeTab={activeTab} onTabChange={setActiveTab} isMacOS={isMacOS()} />
|
|
||||||
|
|
||||||
<main className="flex-1 ml-20 overflow-hidden flex flex-col">
|
|
||||||
<div className="container mx-auto px-8 max-w-[1800px] h-full overflow-hidden flex flex-col">
|
|
||||||
{activeTab === 'main' && <MainEditor />}
|
|
||||||
{activeTab === 'voices' && <VoicesTab />}
|
|
||||||
{activeTab === 'audio' && <AudioTab />}
|
|
||||||
{activeTab === 'server' && <ServerTab />}
|
|
||||||
{activeTab === 'models' && <ModelsTab />}
|
|
||||||
</div>
|
|
||||||
</main>
|
|
||||||
</div>
|
|
||||||
|
|
||||||
{/* Show download toasts for any active downloads (from anywhere) */}
|
|
||||||
{activeDownloads.map((download) => {
|
|
||||||
const displayName = MODEL_DISPLAY_NAMES[download.model_name] || download.model_name;
|
|
||||||
return (
|
|
||||||
<DownloadToastRestorer
|
|
||||||
key={download.model_name}
|
|
||||||
modelName={download.model_name}
|
|
||||||
displayName={displayName}
|
|
||||||
/>
|
|
||||||
);
|
|
||||||
})}
|
|
||||||
|
|
||||||
<Toaster />
|
|
||||||
</AppFrame>
|
|
||||||
);
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Component that restores a download toast for a specific model.
|
|
||||||
*/
|
|
||||||
function DownloadToastRestorer({
|
|
||||||
modelName,
|
|
||||||
displayName,
|
|
||||||
}: {
|
|
||||||
modelName: string;
|
|
||||||
displayName: string;
|
|
||||||
}) {
|
|
||||||
// Use the download toast hook to restore the toast
|
|
||||||
useModelDownloadToast({
|
|
||||||
modelName,
|
|
||||||
displayName,
|
|
||||||
enabled: true,
|
|
||||||
});
|
|
||||||
|
|
||||||
return null;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
export default App;
|
export default App;
|
||||||
|
|||||||
@@ -1,18 +1,35 @@
|
|||||||
|
import { useRouterState } from '@tanstack/react-router';
|
||||||
import { TitleBarDragRegion } from '@/components/TitleBarDragRegion';
|
import { TitleBarDragRegion } from '@/components/TitleBarDragRegion';
|
||||||
import { AudioPlayer } from '@/components/AudioPlayer/AudioPlayer';
|
import { AudioPlayer } from '@/components/AudioPlayer/AudioPlayer';
|
||||||
|
import { StoryTrackEditor } from '@/components/StoriesTab/StoryTrackEditor';
|
||||||
import { TOP_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
import { TOP_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
||||||
import { cn } from '@/lib/utils/cn';
|
import { cn } from '@/lib/utils/cn';
|
||||||
|
import { useStoryStore } from '@/stores/storyStore';
|
||||||
|
import { useStory } from '@/lib/hooks/useStories';
|
||||||
|
|
||||||
interface AppFrameProps {
|
interface AppFrameProps {
|
||||||
children: React.ReactNode;
|
children: React.ReactNode;
|
||||||
}
|
}
|
||||||
|
|
||||||
export function AppFrame({ children }: AppFrameProps) {
|
export function AppFrame({ children }: AppFrameProps) {
|
||||||
|
const routerState = useRouterState();
|
||||||
|
const isStoriesRoute = routerState.location.pathname === '/stories';
|
||||||
|
|
||||||
|
const selectedStoryId = useStoryStore((state) => state.selectedStoryId);
|
||||||
|
const { data: story } = useStory(selectedStoryId);
|
||||||
|
|
||||||
|
// Show track editor when on stories route with a selected story that has items
|
||||||
|
const showTrackEditor = isStoriesRoute && selectedStoryId && story && story.items.length > 0;
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<div className={cn('h-screen bg-background flex flex-col overflow-hidden', TOP_SAFE_AREA_PADDING)}>
|
<div className={cn('h-screen bg-background flex flex-col overflow-hidden', TOP_SAFE_AREA_PADDING)}>
|
||||||
<TitleBarDragRegion />
|
<TitleBarDragRegion />
|
||||||
{children}
|
{children}
|
||||||
<AudioPlayer />
|
{showTrackEditor ? (
|
||||||
|
<StoryTrackEditor storyId={story.id} items={story.items} />
|
||||||
|
) : (
|
||||||
|
<AudioPlayer />
|
||||||
|
)}
|
||||||
</div>
|
</div>
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,16 +1,17 @@
|
|||||||
import { useQuery } from '@tanstack/react-query';
|
import { useQuery } from '@tanstack/react-query';
|
||||||
import { invoke } from '@tauri-apps/api/core';
|
|
||||||
import { Pause, Play, Repeat, Volume2, VolumeX, X } from 'lucide-react';
|
import { Pause, Play, Repeat, Volume2, VolumeX, X } from 'lucide-react';
|
||||||
import { useEffect, useMemo, useRef, useState } from 'react';
|
import { useEffect, useMemo, useRef, useState } from 'react';
|
||||||
import WaveSurfer from 'wavesurfer.js';
|
import WaveSurfer from 'wavesurfer.js';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { Slider } from '@/components/ui/slider';
|
import { Slider } from '@/components/ui/slider';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { apiClient } from '@/lib/api/client';
|
||||||
import { isTauri } from '@/lib/tauri';
|
|
||||||
import { formatAudioDuration } from '@/lib/utils/audio';
|
import { formatAudioDuration } from '@/lib/utils/audio';
|
||||||
|
import { debug } from '@/lib/utils/debug';
|
||||||
import { usePlayerStore } from '@/stores/playerStore';
|
import { usePlayerStore } from '@/stores/playerStore';
|
||||||
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
export function AudioPlayer() {
|
export function AudioPlayer() {
|
||||||
|
const platform = usePlatform();
|
||||||
const {
|
const {
|
||||||
audioUrl,
|
audioUrl,
|
||||||
audioId,
|
audioId,
|
||||||
@@ -38,7 +39,7 @@ export function AudioPlayer() {
|
|||||||
if (!profileId) return { channel_ids: [] };
|
if (!profileId) return { channel_ids: [] };
|
||||||
return apiClient.getProfileChannels(profileId);
|
return apiClient.getProfileChannels(profileId);
|
||||||
},
|
},
|
||||||
enabled: !!profileId && isTauri(),
|
enabled: !!profileId && platform.metadata.isTauri,
|
||||||
});
|
});
|
||||||
|
|
||||||
const { data: channels } = useQuery({
|
const { data: channels } = useQuery({
|
||||||
@@ -49,28 +50,17 @@ export function AudioPlayer() {
|
|||||||
|
|
||||||
// Determine if we should use native playback
|
// Determine if we should use native playback
|
||||||
const useNativePlayback = useMemo(() => {
|
const useNativePlayback = useMemo(() => {
|
||||||
console.log('useNativePlayback memo:', {
|
if (!platform.metadata.isTauri || !profileChannels || !channels) {
|
||||||
isTauri: isTauri(),
|
|
||||||
profileId,
|
|
||||||
profileChannels,
|
|
||||||
channels,
|
|
||||||
});
|
|
||||||
|
|
||||||
if (!isTauri() || !profileChannels || !channels) {
|
|
||||||
console.log('useNativePlayback: false - missing requirements');
|
|
||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
|
|
||||||
const assignedChannels = channels.filter((ch) => profileChannels.channel_ids.includes(ch.id));
|
const assignedChannels = channels.filter((ch) => profileChannels.channel_ids.includes(ch.id));
|
||||||
|
|
||||||
console.log('Assigned channels:', assignedChannels);
|
|
||||||
|
|
||||||
// Use native playback if any assigned channel has non-default devices
|
// Use native playback if any assigned channel has non-default devices
|
||||||
const shouldUseNative = assignedChannels.some(
|
const shouldUseNative = assignedChannels.some(
|
||||||
(ch) => ch.device_ids.length > 0 && !ch.is_default,
|
(ch) => ch.device_ids.length > 0 && !ch.is_default,
|
||||||
);
|
);
|
||||||
|
|
||||||
console.log('useNativePlayback result:', shouldUseNative);
|
|
||||||
return shouldUseNative;
|
return shouldUseNative;
|
||||||
}, [profileChannels, channels, profileId]);
|
}, [profileChannels, channels, profileId]);
|
||||||
|
|
||||||
@@ -91,11 +81,11 @@ export function AudioPlayer() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
if (wavesurferRef.current) {
|
if (wavesurferRef.current) {
|
||||||
console.log('WaveSurfer already initialized, skipping');
|
debug.log('WaveSurfer already initialized, skipping');
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
console.log('Creating NEW WaveSurfer instance');
|
debug.log('Creating NEW WaveSurfer instance');
|
||||||
|
|
||||||
// Wait for container to be properly rendered
|
// Wait for container to be properly rendered
|
||||||
const initWaveSurfer = () => {
|
const initWaveSurfer = () => {
|
||||||
@@ -121,7 +111,7 @@ export function AudioPlayer() {
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
console.log('Initializing WaveSurfer...', {
|
debug.log('Initializing WaveSurfer...', {
|
||||||
container,
|
container,
|
||||||
width: rect.width,
|
width: rect.width,
|
||||||
height: rect.height,
|
height: rect.height,
|
||||||
@@ -154,9 +144,9 @@ export function AudioPlayer() {
|
|||||||
});
|
});
|
||||||
|
|
||||||
wavesurferRef.current = wavesurfer;
|
wavesurferRef.current = wavesurfer;
|
||||||
console.log('WaveSurfer created successfully');
|
debug.log('WaveSurfer created successfully');
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Failed to create WaveSurfer:', error);
|
debug.error('Failed to create WaveSurfer:', error);
|
||||||
setError(
|
setError(
|
||||||
`Failed to initialize waveform: ${error instanceof Error ? error.message : String(error)}`,
|
`Failed to initialize waveform: ${error instanceof Error ? error.message : String(error)}`,
|
||||||
);
|
);
|
||||||
@@ -178,8 +168,8 @@ export function AudioPlayer() {
|
|||||||
loadingRef.current = false;
|
loadingRef.current = false;
|
||||||
setIsLoading(false);
|
setIsLoading(false);
|
||||||
setError(null);
|
setError(null);
|
||||||
console.log('Audio ready, duration:', dur);
|
debug.log('Audio ready, duration:', dur);
|
||||||
console.log('Waveform should be visible now');
|
debug.log('Waveform should be visible now');
|
||||||
|
|
||||||
// Ensure volume is set
|
// Ensure volume is set
|
||||||
const currentVolume = usePlayerStore.getState().volume;
|
const currentVolume = usePlayerStore.getState().volume;
|
||||||
@@ -191,7 +181,7 @@ export function AudioPlayer() {
|
|||||||
if (mediaElement && !isUsingNativePlaybackRef.current) {
|
if (mediaElement && !isUsingNativePlaybackRef.current) {
|
||||||
mediaElement.volume = currentVolume;
|
mediaElement.volume = currentVolume;
|
||||||
mediaElement.muted = false;
|
mediaElement.muted = false;
|
||||||
console.log('Audio element volume:', mediaElement.volume, 'muted:', mediaElement.muted);
|
debug.log('Audio element volume:', mediaElement.volume, 'muted:', mediaElement.muted);
|
||||||
}
|
}
|
||||||
|
|
||||||
// Auto-play when ready - check if we should use native playback
|
// Auto-play when ready - check if we should use native playback
|
||||||
@@ -199,28 +189,28 @@ export function AudioPlayer() {
|
|||||||
const currentAudioUrl = usePlayerStore.getState().audioUrl;
|
const currentAudioUrl = usePlayerStore.getState().audioUrl;
|
||||||
const currentProfileId = usePlayerStore.getState().profileId;
|
const currentProfileId = usePlayerStore.getState().profileId;
|
||||||
|
|
||||||
console.log('Auto-play check - capturing runtime values...');
|
debug.log('Auto-play check - capturing runtime values...');
|
||||||
|
|
||||||
// Fetch profile channels at runtime (not using captured value)
|
// Fetch profile channels at runtime (not using captured value)
|
||||||
let runtimeProfileChannels = null;
|
let runtimeProfileChannels = null;
|
||||||
let runtimeChannels = null;
|
let runtimeChannels = null;
|
||||||
|
|
||||||
if (isTauri() && currentProfileId) {
|
if (platform.metadata.isTauri && currentProfileId) {
|
||||||
try {
|
try {
|
||||||
runtimeProfileChannels = await apiClient.getProfileChannels(currentProfileId);
|
runtimeProfileChannels = await apiClient.getProfileChannels(currentProfileId);
|
||||||
console.log('Runtime profileChannels:', runtimeProfileChannels);
|
debug.log('Runtime profileChannels:', runtimeProfileChannels);
|
||||||
|
|
||||||
if (runtimeProfileChannels && runtimeProfileChannels.channel_ids.length > 0) {
|
if (runtimeProfileChannels && runtimeProfileChannels.channel_ids.length > 0) {
|
||||||
runtimeChannels = await apiClient.listChannels();
|
runtimeChannels = await apiClient.listChannels();
|
||||||
console.log('Runtime channels:', runtimeChannels);
|
debug.log('Runtime channels:', runtimeChannels);
|
||||||
}
|
}
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Failed to fetch runtime channel data:', error);
|
debug.error('Failed to fetch runtime channel data:', error);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
console.log('Auto-play check:', {
|
debug.log('Auto-play check:', {
|
||||||
isTauri: isTauri(),
|
isTauri: platform.metadata.isTauri,
|
||||||
currentAudioUrl,
|
currentAudioUrl,
|
||||||
currentProfileId,
|
currentProfileId,
|
||||||
hasProfileChannels: !!runtimeProfileChannels,
|
hasProfileChannels: !!runtimeProfileChannels,
|
||||||
@@ -228,21 +218,21 @@ export function AudioPlayer() {
|
|||||||
});
|
});
|
||||||
|
|
||||||
if (
|
if (
|
||||||
isTauri() &&
|
platform.metadata.isTauri &&
|
||||||
currentAudioUrl &&
|
currentAudioUrl &&
|
||||||
currentProfileId &&
|
currentProfileId &&
|
||||||
runtimeProfileChannels &&
|
runtimeProfileChannels &&
|
||||||
runtimeChannels
|
runtimeChannels
|
||||||
) {
|
) {
|
||||||
console.log('Attempting native audio playback...');
|
debug.log('Attempting native audio playback...');
|
||||||
|
|
||||||
// Stop any existing native playback first
|
// Stop any existing native playback first
|
||||||
if (isUsingNativePlaybackRef.current) {
|
if (isUsingNativePlaybackRef.current) {
|
||||||
try {
|
try {
|
||||||
await invoke('stop_audio_playback');
|
platform.audio.stopPlayback();
|
||||||
console.log('Stopped existing native playback before starting new one');
|
debug.log('Stopped existing native playback before starting new one');
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Failed to stop existing playback:', error);
|
debug.error('Failed to stop existing playback:', error);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -251,16 +241,16 @@ export function AudioPlayer() {
|
|||||||
const assignedChannels = runtimeChannels.filter((ch: any) =>
|
const assignedChannels = runtimeChannels.filter((ch: any) =>
|
||||||
runtimeProfileChannels.channel_ids.includes(ch.id),
|
runtimeProfileChannels.channel_ids.includes(ch.id),
|
||||||
);
|
);
|
||||||
console.log('Assigned channels for playback:', assignedChannels);
|
debug.log('Assigned channels for playback:', assignedChannels);
|
||||||
|
|
||||||
// Check if any assigned channel has non-default devices
|
// Check if any assigned channel has non-default devices
|
||||||
const shouldUseNative = assignedChannels.some(
|
const shouldUseNative = assignedChannels.some(
|
||||||
(ch: any) => ch.device_ids.length > 0 && !ch.is_default,
|
(ch: any) => ch.device_ids.length > 0 && !ch.is_default,
|
||||||
);
|
);
|
||||||
console.log('Should use native playback:', shouldUseNative);
|
debug.log('Should use native playback:', shouldUseNative);
|
||||||
|
|
||||||
if (!shouldUseNative) {
|
if (!shouldUseNative) {
|
||||||
console.log('No custom devices assigned, falling back to WaveSurfer');
|
debug.log('No custom devices assigned, falling back to WaveSurfer');
|
||||||
// Reset native playback flag and unmute WaveSurfer
|
// Reset native playback flag and unmute WaveSurfer
|
||||||
isUsingNativePlaybackRef.current = false;
|
isUsingNativePlaybackRef.current = false;
|
||||||
const mediaElement = wavesurfer.getMediaElement();
|
const mediaElement = wavesurfer.getMediaElement();
|
||||||
@@ -268,7 +258,7 @@ export function AudioPlayer() {
|
|||||||
const currentVolume = usePlayerStore.getState().volume;
|
const currentVolume = usePlayerStore.getState().volume;
|
||||||
mediaElement.volume = currentVolume;
|
mediaElement.volume = currentVolume;
|
||||||
mediaElement.muted = false;
|
mediaElement.muted = false;
|
||||||
console.log(
|
debug.log(
|
||||||
'WaveSurfer unmuted for normal playback - volume:',
|
'WaveSurfer unmuted for normal playback - volume:',
|
||||||
mediaElement.volume,
|
mediaElement.volume,
|
||||||
'muted:',
|
'muted:',
|
||||||
@@ -277,23 +267,20 @@ export function AudioPlayer() {
|
|||||||
}
|
}
|
||||||
} else {
|
} else {
|
||||||
const deviceIds = assignedChannels.flatMap((ch: any) => ch.device_ids);
|
const deviceIds = assignedChannels.flatMap((ch: any) => ch.device_ids);
|
||||||
console.log('Device IDs to play to:', deviceIds);
|
debug.log('Device IDs to play to:', deviceIds);
|
||||||
|
|
||||||
if (deviceIds.length > 0) {
|
if (deviceIds.length > 0) {
|
||||||
console.log('Fetching audio data from:', currentAudioUrl);
|
debug.log('Fetching audio data from:', currentAudioUrl);
|
||||||
// Fetch audio data
|
// Fetch audio data
|
||||||
const response = await fetch(currentAudioUrl);
|
const response = await fetch(currentAudioUrl);
|
||||||
const audioData = new Uint8Array(await response.arrayBuffer());
|
const audioData = new Uint8Array(await response.arrayBuffer());
|
||||||
console.log('Audio data size:', audioData.length);
|
debug.log('Audio data size:', audioData.length);
|
||||||
|
|
||||||
// Play via native audio
|
// Play via native audio
|
||||||
console.log('Invoking play_audio_to_devices...');
|
debug.log('Invoking play_audio_to_devices...');
|
||||||
try {
|
try {
|
||||||
const result = await invoke('play_audio_to_devices', {
|
await platform.audio.playToDevices(audioData, deviceIds);
|
||||||
audioData: Array.from(audioData),
|
debug.log('play_audio_to_devices completed successfully');
|
||||||
deviceIds: deviceIds,
|
|
||||||
});
|
|
||||||
console.log('play_audio_to_devices completed successfully, result:', result);
|
|
||||||
|
|
||||||
// Mark that we're using native playback
|
// Mark that we're using native playback
|
||||||
isUsingNativePlaybackRef.current = true;
|
isUsingNativePlaybackRef.current = true;
|
||||||
@@ -304,7 +291,7 @@ export function AudioPlayer() {
|
|||||||
if (mediaElement) {
|
if (mediaElement) {
|
||||||
mediaElement.volume = 0;
|
mediaElement.volume = 0;
|
||||||
mediaElement.muted = true;
|
mediaElement.muted = true;
|
||||||
console.log(
|
debug.log(
|
||||||
'WaveSurfer muted for native playback - volume:',
|
'WaveSurfer muted for native playback - volume:',
|
||||||
mediaElement.volume,
|
mediaElement.volume,
|
||||||
'muted:',
|
'muted:',
|
||||||
@@ -314,22 +301,22 @@ export function AudioPlayer() {
|
|||||||
|
|
||||||
// Start WaveSurfer playback for visualization (muted)
|
// Start WaveSurfer playback for visualization (muted)
|
||||||
wavesurfer.play().catch((error) => {
|
wavesurfer.play().catch((error) => {
|
||||||
console.error('Failed to start WaveSurfer visualization:', error);
|
debug.error('Failed to start WaveSurfer visualization:', error);
|
||||||
});
|
});
|
||||||
|
|
||||||
setIsPlaying(true);
|
setIsPlaying(true);
|
||||||
console.log('Auto-playing via native audio routing - SUCCESS');
|
debug.log('Auto-playing via native audio routing - SUCCESS');
|
||||||
return;
|
return;
|
||||||
} catch (invokeError) {
|
} catch (invokeError) {
|
||||||
console.error('play_audio_to_devices invoke failed:', invokeError);
|
debug.error('play_audio_to_devices invoke failed:', invokeError);
|
||||||
throw invokeError;
|
throw invokeError;
|
||||||
}
|
}
|
||||||
} else {
|
} else {
|
||||||
console.log('No device IDs found, falling back to WaveSurfer');
|
debug.log('No device IDs found, falling back to WaveSurfer');
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error(
|
debug.error(
|
||||||
'Native playback failed during auto-play, falling back to WaveSurfer:',
|
'Native playback failed during auto-play, falling back to WaveSurfer:',
|
||||||
error,
|
error,
|
||||||
);
|
);
|
||||||
@@ -340,7 +327,7 @@ export function AudioPlayer() {
|
|||||||
const currentVolume = usePlayerStore.getState().volume;
|
const currentVolume = usePlayerStore.getState().volume;
|
||||||
mediaElement.volume = currentVolume;
|
mediaElement.volume = currentVolume;
|
||||||
mediaElement.muted = false;
|
mediaElement.muted = false;
|
||||||
console.log(
|
debug.log(
|
||||||
'WaveSurfer unmuted after native playback failure - volume:',
|
'WaveSurfer unmuted after native playback failure - volume:',
|
||||||
mediaElement.volume,
|
mediaElement.volume,
|
||||||
'muted:',
|
'muted:',
|
||||||
@@ -350,7 +337,7 @@ export function AudioPlayer() {
|
|||||||
// Fall through to WaveSurfer playback
|
// Fall through to WaveSurfer playback
|
||||||
}
|
}
|
||||||
} else {
|
} else {
|
||||||
console.log('Not using native playback, using WaveSurfer');
|
debug.log('Not using native playback, using WaveSurfer');
|
||||||
// Reset native playback flag and unmute WaveSurfer
|
// Reset native playback flag and unmute WaveSurfer
|
||||||
isUsingNativePlaybackRef.current = false;
|
isUsingNativePlaybackRef.current = false;
|
||||||
const mediaElement = wavesurfer.getMediaElement();
|
const mediaElement = wavesurfer.getMediaElement();
|
||||||
@@ -358,7 +345,7 @@ export function AudioPlayer() {
|
|||||||
const currentVolume = usePlayerStore.getState().volume;
|
const currentVolume = usePlayerStore.getState().volume;
|
||||||
mediaElement.volume = currentVolume;
|
mediaElement.volume = currentVolume;
|
||||||
mediaElement.muted = false;
|
mediaElement.muted = false;
|
||||||
console.log(
|
debug.log(
|
||||||
'WaveSurfer unmuted for normal playback - volume:',
|
'WaveSurfer unmuted for normal playback - volume:',
|
||||||
mediaElement.volume,
|
mediaElement.volume,
|
||||||
'muted:',
|
'muted:',
|
||||||
@@ -367,14 +354,22 @@ export function AudioPlayer() {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// Standard WaveSurfer auto-play
|
// Only auto-play if shouldAutoPlay flag is set (user explicitly clicked to play)
|
||||||
// Use a small delay to ensure audio element is fully ready
|
const shouldAutoPlayNow = usePlayerStore.getState().shouldAutoPlay;
|
||||||
setTimeout(() => {
|
if (shouldAutoPlayNow) {
|
||||||
wavesurfer.play().catch((error) => {
|
// Clear the flag first
|
||||||
console.error('Failed to autoplay:', error);
|
usePlayerStore.getState().clearAutoPlayFlag();
|
||||||
// Don't show error for autoplay failures (browser restrictions)
|
|
||||||
});
|
// Use a small delay to ensure audio element is fully ready
|
||||||
}, 100);
|
setTimeout(() => {
|
||||||
|
wavesurfer.play().catch((error) => {
|
||||||
|
debug.error('Failed to autoplay:', error);
|
||||||
|
// Don't show error for autoplay failures (browser restrictions)
|
||||||
|
});
|
||||||
|
}, 100);
|
||||||
|
} else {
|
||||||
|
debug.log('Skipping auto-play - shouldAutoPlay is false');
|
||||||
|
}
|
||||||
});
|
});
|
||||||
|
|
||||||
// Handle play/pause
|
// Handle play/pause
|
||||||
@@ -388,13 +383,13 @@ export function AudioPlayer() {
|
|||||||
if (isUsingNativePlaybackRef.current) {
|
if (isUsingNativePlaybackRef.current) {
|
||||||
mediaElement.volume = 0;
|
mediaElement.volume = 0;
|
||||||
mediaElement.muted = true;
|
mediaElement.muted = true;
|
||||||
console.log('Playing (native mode) - WaveSurfer muted for visualization only');
|
debug.log('Playing (native mode) - WaveSurfer muted for visualization only');
|
||||||
} else {
|
} else {
|
||||||
// Ensure WaveSurfer is unmuted for normal playback
|
// Ensure WaveSurfer is unmuted for normal playback
|
||||||
const currentVolume = usePlayerStore.getState().volume;
|
const currentVolume = usePlayerStore.getState().volume;
|
||||||
mediaElement.volume = currentVolume;
|
mediaElement.volume = currentVolume;
|
||||||
mediaElement.muted = false;
|
mediaElement.muted = false;
|
||||||
console.log(
|
debug.log(
|
||||||
'Playing (normal mode) - volume:',
|
'Playing (normal mode) - volume:',
|
||||||
mediaElement.volume,
|
mediaElement.volume,
|
||||||
'muted:',
|
'muted:',
|
||||||
@@ -412,12 +407,17 @@ export function AudioPlayer() {
|
|||||||
wavesurfer.play();
|
wavesurfer.play();
|
||||||
} else {
|
} else {
|
||||||
setIsPlaying(false);
|
setIsPlaying(false);
|
||||||
|
// Trigger finish callback if set
|
||||||
|
const onFinish = usePlayerStore.getState().onFinish;
|
||||||
|
if (onFinish) {
|
||||||
|
onFinish();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
});
|
});
|
||||||
|
|
||||||
// Handle errors
|
// Handle errors
|
||||||
wavesurfer.on('error', (error) => {
|
wavesurfer.on('error', (error) => {
|
||||||
console.error('WaveSurfer error:', error);
|
debug.error('WaveSurfer error:', error);
|
||||||
setIsLoading(false);
|
setIsLoading(false);
|
||||||
setError(`Audio error: ${error instanceof Error ? error.message : String(error)}`);
|
setError(`Audio error: ${error instanceof Error ? error.message : String(error)}`);
|
||||||
});
|
});
|
||||||
@@ -432,7 +432,7 @@ export function AudioPlayer() {
|
|||||||
|
|
||||||
// Load audio immediately if audioUrl is already set
|
// Load audio immediately if audioUrl is already set
|
||||||
if (audioUrl) {
|
if (audioUrl) {
|
||||||
console.log('WaveSurfer ready, loading audio:', audioUrl);
|
debug.log('WaveSurfer ready, loading audio:', audioUrl);
|
||||||
loadingRef.current = true;
|
loadingRef.current = true;
|
||||||
setIsLoading(true);
|
setIsLoading(true);
|
||||||
// Stop any current playback before loading new audio
|
// Stop any current playback before loading new audio
|
||||||
@@ -442,11 +442,11 @@ export function AudioPlayer() {
|
|||||||
wavesurfer
|
wavesurfer
|
||||||
.load(audioUrl)
|
.load(audioUrl)
|
||||||
.then(() => {
|
.then(() => {
|
||||||
console.log('Audio loaded into WaveSurfer');
|
debug.log('Audio loaded into WaveSurfer');
|
||||||
loadingRef.current = false;
|
loadingRef.current = false;
|
||||||
})
|
})
|
||||||
.catch((error) => {
|
.catch((error) => {
|
||||||
console.error('Failed to load audio into WaveSurfer:', error);
|
debug.error('Failed to load audio into WaveSurfer:', error);
|
||||||
loadingRef.current = false;
|
loadingRef.current = false;
|
||||||
setIsLoading(false);
|
setIsLoading(false);
|
||||||
setError(
|
setError(
|
||||||
@@ -471,12 +471,12 @@ export function AudioPlayer() {
|
|||||||
});
|
});
|
||||||
|
|
||||||
return () => {
|
return () => {
|
||||||
console.log('Cleaning up WaveSurfer initialization effect');
|
debug.log('Cleaning up WaveSurfer initialization effect');
|
||||||
if (rafId1) cancelAnimationFrame(rafId1);
|
if (rafId1) cancelAnimationFrame(rafId1);
|
||||||
if (rafId2) cancelAnimationFrame(rafId2);
|
if (rafId2) cancelAnimationFrame(rafId2);
|
||||||
if (timeoutId) clearTimeout(timeoutId);
|
if (timeoutId) clearTimeout(timeoutId);
|
||||||
if (wavesurferRef.current) {
|
if (wavesurferRef.current) {
|
||||||
console.log('Destroying WaveSurfer instance');
|
debug.log('Destroying WaveSurfer instance');
|
||||||
try {
|
try {
|
||||||
const mediaElement = wavesurferRef.current.getMediaElement();
|
const mediaElement = wavesurferRef.current.getMediaElement();
|
||||||
if (mediaElement) {
|
if (mediaElement) {
|
||||||
@@ -485,7 +485,7 @@ export function AudioPlayer() {
|
|||||||
}
|
}
|
||||||
wavesurferRef.current.destroy();
|
wavesurferRef.current.destroy();
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Error destroying WaveSurfer:', error);
|
debug.error('Error destroying WaveSurfer:', error);
|
||||||
}
|
}
|
||||||
wavesurferRef.current = null;
|
wavesurferRef.current = null;
|
||||||
}
|
}
|
||||||
@@ -513,15 +513,13 @@ export function AudioPlayer() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Stop native playback if it was active
|
// Stop native playback if it was active
|
||||||
if (isUsingNativePlaybackRef.current && isTauri()) {
|
if (isUsingNativePlaybackRef.current && platform.metadata.isTauri) {
|
||||||
(async () => {
|
try {
|
||||||
try {
|
platform.audio.stopPlayback();
|
||||||
await invoke('stop_audio_playback');
|
debug.log('Stopped native audio playback');
|
||||||
console.log('Stopped native audio playback');
|
} catch (error) {
|
||||||
} catch (error) {
|
debug.error('Failed to stop native playback:', error);
|
||||||
console.error('Failed to stop native playback:', error);
|
}
|
||||||
}
|
|
||||||
})();
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// Reset native playback flag when loading new audio
|
// Reset native playback flag when loading new audio
|
||||||
@@ -537,30 +535,30 @@ export function AudioPlayer() {
|
|||||||
|
|
||||||
// CRITICAL: Force stop any current playback and cancel any pending loads
|
// CRITICAL: Force stop any current playback and cancel any pending loads
|
||||||
// This must happen BEFORE any early returns
|
// This must happen BEFORE any early returns
|
||||||
console.log('Audio URL changed to:', audioUrl);
|
debug.log('Audio URL changed to:', audioUrl);
|
||||||
|
|
||||||
// COMPLETELY stop and destroy the current audio
|
// COMPLETELY stop and destroy the current audio
|
||||||
try {
|
try {
|
||||||
// First pause if playing
|
// First pause if playing
|
||||||
if (wavesurfer.isPlaying()) {
|
if (wavesurfer.isPlaying()) {
|
||||||
console.log('Pausing current playback');
|
debug.log('Pausing current playback');
|
||||||
wavesurfer.pause();
|
wavesurfer.pause();
|
||||||
}
|
}
|
||||||
|
|
||||||
// Stop the media element explicitly
|
// Stop the media element explicitly
|
||||||
const mediaElement = wavesurfer.getMediaElement();
|
const mediaElement = wavesurfer.getMediaElement();
|
||||||
if (mediaElement) {
|
if (mediaElement) {
|
||||||
console.log('Stopping media element');
|
debug.log('Stopping media element');
|
||||||
mediaElement.pause();
|
mediaElement.pause();
|
||||||
mediaElement.currentTime = 0;
|
mediaElement.currentTime = 0;
|
||||||
mediaElement.src = '';
|
mediaElement.src = '';
|
||||||
}
|
}
|
||||||
|
|
||||||
// Use empty() to completely destroy the waveform and media element
|
// Use empty() to completely destroy the waveform and media element
|
||||||
console.log('Calling wavesurfer.empty() to destroy audio');
|
debug.log('Calling wavesurfer.empty() to destroy audio');
|
||||||
wavesurfer.empty();
|
wavesurfer.empty();
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Error stopping previous audio:', error);
|
debug.error('Error stopping previous audio:', error);
|
||||||
// Continue anyway to load new audio
|
// Continue anyway to load new audio
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -575,16 +573,16 @@ export function AudioPlayer() {
|
|||||||
setDuration(0);
|
setDuration(0);
|
||||||
|
|
||||||
// Load new audio
|
// Load new audio
|
||||||
console.log('Starting new audio load for:', audioUrl);
|
debug.log('Starting new audio load for:', audioUrl);
|
||||||
wavesurfer
|
wavesurfer
|
||||||
.load(audioUrl)
|
.load(audioUrl)
|
||||||
.then(() => {
|
.then(() => {
|
||||||
console.log('Audio load promise resolved');
|
debug.log('Audio load promise resolved');
|
||||||
// Don't set loading to false here - wait for 'ready' event
|
// Don't set loading to false here - wait for 'ready' event
|
||||||
})
|
})
|
||||||
.catch((error) => {
|
.catch((error) => {
|
||||||
console.error('Failed to load audio:', error);
|
debug.error('Failed to load audio:', error);
|
||||||
console.error('Audio URL:', audioUrl);
|
debug.error('Audio URL:', audioUrl);
|
||||||
loadingRef.current = false;
|
loadingRef.current = false;
|
||||||
setIsLoading(false);
|
setIsLoading(false);
|
||||||
setError(`Failed to load audio: ${error instanceof Error ? error.message : String(error)}`);
|
setError(`Failed to load audio: ${error instanceof Error ? error.message : String(error)}`);
|
||||||
@@ -599,7 +597,7 @@ export function AudioPlayer() {
|
|||||||
if (isPlaying && wavesurferRef.current.isPlaying() === false) {
|
if (isPlaying && wavesurferRef.current.isPlaying() === false) {
|
||||||
// Only auto-play if audio is ready
|
// Only auto-play if audio is ready
|
||||||
wavesurferRef.current.play().catch((error) => {
|
wavesurferRef.current.play().catch((error) => {
|
||||||
console.error('Failed to play:', error);
|
debug.error('Failed to play:', error);
|
||||||
setIsPlaying(false);
|
setIsPlaying(false);
|
||||||
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
||||||
});
|
});
|
||||||
@@ -619,11 +617,11 @@ export function AudioPlayer() {
|
|||||||
if (isUsingNativePlaybackRef.current) {
|
if (isUsingNativePlaybackRef.current) {
|
||||||
mediaElement.volume = 0;
|
mediaElement.volume = 0;
|
||||||
mediaElement.muted = true;
|
mediaElement.muted = true;
|
||||||
console.log('Volume sync: Using native playback, keeping WaveSurfer muted');
|
debug.log('Volume sync: Using native playback, keeping WaveSurfer muted');
|
||||||
} else {
|
} else {
|
||||||
mediaElement.volume = volume;
|
mediaElement.volume = volume;
|
||||||
mediaElement.muted = volume === 0;
|
mediaElement.muted = volume === 0;
|
||||||
console.log('Volume synced:', volume, 'muted:', mediaElement.muted);
|
debug.log('Volume synced:', volume, 'muted:', mediaElement.muted);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -651,10 +649,10 @@ export function AudioPlayer() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Reset to beginning and play
|
// Reset to beginning and play
|
||||||
console.log('Restarting current audio from beginning');
|
debug.log('Restarting current audio from beginning');
|
||||||
wavesurfer.seekTo(0);
|
wavesurfer.seekTo(0);
|
||||||
wavesurfer.play().catch((error) => {
|
wavesurfer.play().catch((error) => {
|
||||||
console.error('Failed to play after restart:', error);
|
debug.error('Failed to play after restart:', error);
|
||||||
setIsPlaying(false);
|
setIsPlaying(false);
|
||||||
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
||||||
});
|
});
|
||||||
@@ -663,19 +661,42 @@ export function AudioPlayer() {
|
|||||||
clearRestartFlag();
|
clearRestartFlag();
|
||||||
}, [shouldRestart, duration, setIsPlaying, clearRestartFlag]);
|
}, [shouldRestart, duration, setIsPlaying, clearRestartFlag]);
|
||||||
|
|
||||||
|
// Handle shouldAutoPlay flag - for story mode auto-advance
|
||||||
|
const shouldAutoPlay = usePlayerStore((state) => state.shouldAutoPlay);
|
||||||
|
const clearAutoPlayFlag = usePlayerStore((state) => state.clearAutoPlayFlag);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
const wavesurfer = wavesurferRef.current;
|
||||||
|
if (!wavesurfer || !shouldAutoPlay || duration === 0) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Auto-play the newly loaded audio
|
||||||
|
debug.log('Auto-playing next track in story mode');
|
||||||
|
wavesurfer.seekTo(0);
|
||||||
|
wavesurfer.play().catch((error) => {
|
||||||
|
debug.error('Failed to auto-play:', error);
|
||||||
|
setIsPlaying(false);
|
||||||
|
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
||||||
|
});
|
||||||
|
|
||||||
|
// Clear the auto-play flag
|
||||||
|
clearAutoPlayFlag();
|
||||||
|
}, [shouldAutoPlay, duration, setIsPlaying, clearAutoPlayFlag]);
|
||||||
|
|
||||||
// Handle loop - WaveSurfer handles this via the 'finish' event
|
// Handle loop - WaveSurfer handles this via the 'finish' event
|
||||||
|
|
||||||
const handlePlayPause = async () => {
|
const handlePlayPause = async () => {
|
||||||
// Standard WaveSurfer playback (works for both normal and native playback modes)
|
// Standard WaveSurfer playback (works for both normal and native playback modes)
|
||||||
// When using native playback, WaveSurfer is muted but still controls visualization
|
// When using native playback, WaveSurfer is muted but still controls visualization
|
||||||
if (!wavesurferRef.current) {
|
if (!wavesurferRef.current) {
|
||||||
console.error('WaveSurfer not initialized');
|
debug.error('WaveSurfer not initialized');
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Check if audio is loaded
|
// Check if audio is loaded
|
||||||
if (duration === 0 && !isLoading) {
|
if (duration === 0 && !isLoading) {
|
||||||
console.error('Audio not loaded yet');
|
debug.error('Audio not loaded yet');
|
||||||
setError('Audio not loaded. Please wait...');
|
setError('Audio not loaded. Please wait...');
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
@@ -685,10 +706,10 @@ export function AudioPlayer() {
|
|||||||
if (isPlaying) {
|
if (isPlaying) {
|
||||||
// Pause: stop native playback and pause WaveSurfer visualization
|
// Pause: stop native playback and pause WaveSurfer visualization
|
||||||
try {
|
try {
|
||||||
await invoke('stop_audio_playback');
|
platform.audio.stopPlayback();
|
||||||
console.log('Stopped native audio playback');
|
debug.log('Stopped native audio playback');
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Failed to stop native playback:', error);
|
debug.error('Failed to stop native playback:', error);
|
||||||
}
|
}
|
||||||
wavesurferRef.current.pause();
|
wavesurferRef.current.pause();
|
||||||
return;
|
return;
|
||||||
@@ -698,10 +719,10 @@ export function AudioPlayer() {
|
|||||||
try {
|
try {
|
||||||
// Stop any existing native playback first
|
// Stop any existing native playback first
|
||||||
try {
|
try {
|
||||||
await invoke('stop_audio_playback');
|
platform.audio.stopPlayback();
|
||||||
} catch (_error) {
|
} catch (_error) {
|
||||||
// Ignore errors when stopping (might not be playing)
|
// Ignore errors when stopping (might not be playing)
|
||||||
console.log('No existing playback to stop');
|
debug.log('No existing playback to stop');
|
||||||
}
|
}
|
||||||
|
|
||||||
// Collect all device IDs from assigned channels
|
// Collect all device IDs from assigned channels
|
||||||
@@ -716,10 +737,7 @@ export function AudioPlayer() {
|
|||||||
const audioData = new Uint8Array(await response.arrayBuffer());
|
const audioData = new Uint8Array(await response.arrayBuffer());
|
||||||
|
|
||||||
// Play via native audio
|
// Play via native audio
|
||||||
await invoke('play_audio_to_devices', {
|
await platform.audio.playToDevices(audioData, deviceIds);
|
||||||
audioData: Array.from(audioData),
|
|
||||||
deviceIds: deviceIds,
|
|
||||||
});
|
|
||||||
|
|
||||||
// Mark that we're using native playback
|
// Mark that we're using native playback
|
||||||
isUsingNativePlaybackRef.current = true;
|
isUsingNativePlaybackRef.current = true;
|
||||||
@@ -733,7 +751,7 @@ export function AudioPlayer() {
|
|||||||
|
|
||||||
// Start WaveSurfer for visualization (muted)
|
// Start WaveSurfer for visualization (muted)
|
||||||
wavesurferRef.current.play().catch((error) => {
|
wavesurferRef.current.play().catch((error) => {
|
||||||
console.error('Failed to start WaveSurfer visualization:', error);
|
debug.error('Failed to start WaveSurfer visualization:', error);
|
||||||
setIsPlaying(false);
|
setIsPlaying(false);
|
||||||
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
||||||
});
|
});
|
||||||
@@ -741,7 +759,7 @@ export function AudioPlayer() {
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Native playback failed, falling back to WaveSurfer:', error);
|
debug.error('Native playback failed, falling back to WaveSurfer:', error);
|
||||||
// Fall through to WaveSurfer playback
|
// Fall through to WaveSurfer playback
|
||||||
isUsingNativePlaybackRef.current = false;
|
isUsingNativePlaybackRef.current = false;
|
||||||
}
|
}
|
||||||
@@ -761,7 +779,7 @@ export function AudioPlayer() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
wavesurferRef.current.play().catch((error) => {
|
wavesurferRef.current.play().catch((error) => {
|
||||||
console.error('Failed to play:', error);
|
debug.error('Failed to play:', error);
|
||||||
setIsPlaying(false);
|
setIsPlaying(false);
|
||||||
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
setError(`Playback error: ${error instanceof Error ? error.message : String(error)}`);
|
||||||
});
|
});
|
||||||
@@ -780,10 +798,12 @@ export function AudioPlayer() {
|
|||||||
|
|
||||||
const handleClose = () => {
|
const handleClose = () => {
|
||||||
// Stop any native playback
|
// Stop any native playback
|
||||||
if (isUsingNativePlaybackRef.current && isTauri()) {
|
if (isUsingNativePlaybackRef.current && platform.metadata.isTauri) {
|
||||||
invoke('stop_audio_playback').catch((error) => {
|
try {
|
||||||
console.error('Failed to stop native playback:', error);
|
platform.audio.stopPlayback();
|
||||||
});
|
} catch (error) {
|
||||||
|
debug.error('Failed to stop native playback:', error);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
// Stop WaveSurfer
|
// Stop WaveSurfer
|
||||||
if (wavesurferRef.current) {
|
if (wavesurferRef.current) {
|
||||||
|
|||||||
@@ -1,5 +1,4 @@
|
|||||||
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query';
|
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query';
|
||||||
import { invoke } from '@tauri-apps/api/core';
|
|
||||||
import { Check, CheckCircle2, Edit, Plus, Speaker, Trash2 } from 'lucide-react';
|
import { Check, CheckCircle2, Edit, Plus, Speaker, Trash2 } from 'lucide-react';
|
||||||
import { useState } from 'react';
|
import { useState } from 'react';
|
||||||
import { Badge } from '@/components/ui/badge';
|
import { Badge } from '@/components/ui/badge';
|
||||||
@@ -23,9 +22,9 @@ import {
|
|||||||
} from '@/components/ui/select';
|
} from '@/components/ui/select';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { apiClient } from '@/lib/api/client';
|
||||||
import { BOTTOM_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
import { BOTTOM_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
||||||
import { isTauri } from '@/lib/tauri';
|
|
||||||
import { cn } from '@/lib/utils/cn';
|
import { cn } from '@/lib/utils/cn';
|
||||||
import { usePlayerStore } from '@/stores/playerStore';
|
import { usePlayerStore } from '@/stores/playerStore';
|
||||||
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
interface AudioDevice {
|
interface AudioDevice {
|
||||||
id: string;
|
id: string;
|
||||||
@@ -34,6 +33,7 @@ interface AudioDevice {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export function AudioTab() {
|
export function AudioTab() {
|
||||||
|
const platform = usePlatform();
|
||||||
const [createDialogOpen, setCreateDialogOpen] = useState(false);
|
const [createDialogOpen, setCreateDialogOpen] = useState(false);
|
||||||
const [editingChannel, setEditingChannel] = useState<string | null>(null);
|
const [editingChannel, setEditingChannel] = useState<string | null>(null);
|
||||||
const [selectedChannelId, setSelectedChannelId] = useState<string | null>(null);
|
const [selectedChannelId, setSelectedChannelId] = useState<string | null>(null);
|
||||||
@@ -49,18 +49,17 @@ export function AudioTab() {
|
|||||||
const { data: devices, isLoading: devicesLoading } = useQuery({
|
const { data: devices, isLoading: devicesLoading } = useQuery({
|
||||||
queryKey: ['audio-devices'],
|
queryKey: ['audio-devices'],
|
||||||
queryFn: async () => {
|
queryFn: async () => {
|
||||||
if (!isTauri()) {
|
if (!platform.metadata.isTauri) {
|
||||||
return [];
|
return [];
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
const result = await invoke<AudioDevice[]>('list_audio_output_devices');
|
return await platform.audio.listOutputDevices();
|
||||||
return result;
|
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Failed to list audio devices:', error);
|
console.error('Failed to list audio devices:', error);
|
||||||
return [];
|
return [];
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
enabled: isTauri(),
|
enabled: platform.metadata.isTauri,
|
||||||
});
|
});
|
||||||
|
|
||||||
const { data: profiles } = useQuery({
|
const { data: profiles } = useQuery({
|
||||||
@@ -342,7 +341,7 @@ export function AudioTab() {
|
|||||||
<div className="flex flex-col items-center justify-center py-12 border-2 border-dashed border-muted rounded-md">
|
<div className="flex flex-col items-center justify-center py-12 border-2 border-dashed border-muted rounded-md">
|
||||||
<CheckCircle2 className="h-12 w-12 text-muted-foreground mb-4" />
|
<CheckCircle2 className="h-12 w-12 text-muted-foreground mb-4" />
|
||||||
<p className="text-muted-foreground text-center">
|
<p className="text-muted-foreground text-center">
|
||||||
{isTauri() ? 'No audio devices found' : 'Audio device selection requires Tauri'}
|
{platform.metadata.isTauri ? 'No audio devices found' : 'Audio device selection requires Tauri'}
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
)}
|
)}
|
||||||
|
|||||||
@@ -1,9 +1,7 @@
|
|||||||
import { zodResolver } from '@hookform/resolvers/zod';
|
import { useMatchRoute } from '@tanstack/react-router';
|
||||||
import { AnimatePresence, motion } from 'framer-motion';
|
import { AnimatePresence, motion } from 'framer-motion';
|
||||||
import { Loader2, Sparkles } from 'lucide-react';
|
import { Loader2, MessageSquare, Sparkles } from 'lucide-react';
|
||||||
import { useEffect, useRef, useState } from 'react';
|
import { useEffect, useRef, useState } from 'react';
|
||||||
import { useForm } from 'react-hook-form';
|
|
||||||
import * as z from 'zod';
|
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { Form, FormControl, FormField, FormItem, FormMessage } from '@/components/ui/form';
|
import { Form, FormControl, FormField, FormItem, FormMessage } from '@/components/ui/form';
|
||||||
import {
|
import {
|
||||||
@@ -15,51 +13,65 @@ import {
|
|||||||
} from '@/components/ui/select';
|
} from '@/components/ui/select';
|
||||||
import { Textarea } from '@/components/ui/textarea';
|
import { Textarea } from '@/components/ui/textarea';
|
||||||
import { useToast } from '@/components/ui/use-toast';
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { LANGUAGE_OPTIONS } from '@/lib/constants/languages';
|
||||||
import { LANGUAGE_CODES, LANGUAGE_OPTIONS, type LanguageCode } from '@/lib/constants/languages';
|
import { useGenerationForm } from '@/lib/hooks/useGenerationForm';
|
||||||
import { useGeneration } from '@/lib/hooks/useGeneration';
|
import { useProfile, useProfiles } from '@/lib/hooks/useProfiles';
|
||||||
import { useModelDownloadToast } from '@/lib/hooks/useModelDownloadToast';
|
import { useAddStoryItem, useStory } from '@/lib/hooks/useStories';
|
||||||
import { useProfile } from '@/lib/hooks/useProfiles';
|
import { cn } from '@/lib/utils/cn';
|
||||||
import { useGenerationStore } from '@/stores/generationStore';
|
import { useStoryStore } from '@/stores/storyStore';
|
||||||
import { usePlayerStore } from '@/stores/playerStore';
|
|
||||||
import { useUIStore } from '@/stores/uiStore';
|
import { useUIStore } from '@/stores/uiStore';
|
||||||
|
|
||||||
const generationSchema = z.object({
|
|
||||||
text: z.string().min(1, 'Text is required').max(5000),
|
|
||||||
language: z.enum(LANGUAGE_CODES as [LanguageCode, ...LanguageCode[]]),
|
|
||||||
modelSize: z.enum(['1.7B', '0.6B']).optional(),
|
|
||||||
});
|
|
||||||
|
|
||||||
type GenerationFormValues = z.infer<typeof generationSchema>;
|
|
||||||
|
|
||||||
interface FloatingGenerateBoxProps {
|
interface FloatingGenerateBoxProps {
|
||||||
isPlayerOpen: boolean;
|
isPlayerOpen?: boolean;
|
||||||
|
showVoiceSelector?: boolean;
|
||||||
}
|
}
|
||||||
|
|
||||||
export function FloatingGenerateBox({ isPlayerOpen }: FloatingGenerateBoxProps) {
|
export function FloatingGenerateBox({
|
||||||
|
isPlayerOpen = false,
|
||||||
|
showVoiceSelector = false,
|
||||||
|
}: FloatingGenerateBoxProps) {
|
||||||
const selectedProfileId = useUIStore((state) => state.selectedProfileId);
|
const selectedProfileId = useUIStore((state) => state.selectedProfileId);
|
||||||
|
const setSelectedProfileId = useUIStore((state) => state.setSelectedProfileId);
|
||||||
const { data: selectedProfile } = useProfile(selectedProfileId || '');
|
const { data: selectedProfile } = useProfile(selectedProfileId || '');
|
||||||
const generation = useGeneration();
|
const { data: profiles } = useProfiles();
|
||||||
const { toast } = useToast();
|
|
||||||
const setAudio = usePlayerStore((state) => state.setAudio);
|
|
||||||
const setIsGenerating = useGenerationStore((state) => state.setIsGenerating);
|
|
||||||
const [downloadingModelName, setDownloadingModelName] = useState<string | null>(null);
|
|
||||||
const [downloadingDisplayName, setDownloadingDisplayName] = useState<string | null>(null);
|
|
||||||
const [isExpanded, setIsExpanded] = useState(false);
|
const [isExpanded, setIsExpanded] = useState(false);
|
||||||
|
const [isInstructMode, setIsInstructMode] = useState(false);
|
||||||
const containerRef = useRef<HTMLDivElement>(null);
|
const containerRef = useRef<HTMLDivElement>(null);
|
||||||
|
const textareaRef = useRef<HTMLTextAreaElement | null>(null);
|
||||||
|
const matchRoute = useMatchRoute();
|
||||||
|
const isStoriesRoute = matchRoute({ to: '/stories' });
|
||||||
|
const selectedStoryId = useStoryStore((state) => state.selectedStoryId);
|
||||||
|
const trackEditorHeight = useStoryStore((state) => state.trackEditorHeight);
|
||||||
|
const { data: currentStory } = useStory(selectedStoryId);
|
||||||
|
const addStoryItem = useAddStoryItem();
|
||||||
|
const { toast } = useToast();
|
||||||
|
|
||||||
useModelDownloadToast({
|
// Calculate if track editor is visible (on stories route with items)
|
||||||
modelName: downloadingModelName || '',
|
const hasTrackEditor = isStoriesRoute && currentStory && currentStory.items.length > 0;
|
||||||
displayName: downloadingDisplayName || '',
|
|
||||||
enabled: !!downloadingModelName,
|
|
||||||
});
|
|
||||||
|
|
||||||
const form = useForm<GenerationFormValues>({
|
const { form, handleSubmit, isPending } = useGenerationForm({
|
||||||
resolver: zodResolver(generationSchema),
|
onSuccess: async (generationId) => {
|
||||||
defaultValues: {
|
setIsExpanded(false);
|
||||||
text: '',
|
// If on stories route and a story is selected, add generation to story
|
||||||
language: 'en',
|
if (isStoriesRoute && selectedStoryId && generationId) {
|
||||||
modelSize: '1.7B',
|
try {
|
||||||
|
await addStoryItem.mutateAsync({
|
||||||
|
storyId: selectedStoryId,
|
||||||
|
data: { generation_id: generationId },
|
||||||
|
});
|
||||||
|
toast({
|
||||||
|
title: 'Added to story',
|
||||||
|
description: `Generation added to "${currentStory?.name || 'story'}"`,
|
||||||
|
});
|
||||||
|
} catch (error) {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to add to story',
|
||||||
|
description:
|
||||||
|
error instanceof Error ? error.message : 'Could not add generation to story',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
@@ -93,70 +105,85 @@ export function FloatingGenerateBox({ isPlayerOpen }: FloatingGenerateBoxProps)
|
|||||||
};
|
};
|
||||||
}, [isExpanded]);
|
}, [isExpanded]);
|
||||||
|
|
||||||
async function onSubmit(data: GenerationFormValues) {
|
// Set first voice as default if none selected
|
||||||
if (!selectedProfileId) {
|
useEffect(() => {
|
||||||
toast({
|
if (!selectedProfileId && profiles && profiles.length > 0) {
|
||||||
title: 'No profile selected',
|
setSelectedProfileId(profiles[0].id);
|
||||||
description: 'Please select a voice profile from the cards above.',
|
|
||||||
variant: 'destructive',
|
|
||||||
});
|
|
||||||
return;
|
|
||||||
}
|
}
|
||||||
|
}, [selectedProfileId, profiles, setSelectedProfileId]);
|
||||||
|
|
||||||
try {
|
// Auto-resize textarea based on content (only when expanded)
|
||||||
setIsGenerating(true);
|
useEffect(() => {
|
||||||
|
if (!isExpanded) {
|
||||||
const modelName = `qwen-tts-${data.modelSize}`;
|
// Reset textarea height after collapse animation completes
|
||||||
const displayName = data.modelSize === '1.7B' ? 'Qwen TTS 1.7B' : 'Qwen TTS 0.6B';
|
const timeoutId = setTimeout(() => {
|
||||||
|
const textarea = textareaRef.current;
|
||||||
try {
|
if (textarea) {
|
||||||
const modelStatus = await apiClient.getModelStatus();
|
textarea.style.height = '32px';
|
||||||
const model = modelStatus.models.find((m) => m.model_name === modelName);
|
textarea.style.overflowY = 'hidden';
|
||||||
|
|
||||||
if (model && !model.downloaded) {
|
|
||||||
setDownloadingModelName(modelName);
|
|
||||||
setDownloadingDisplayName(displayName);
|
|
||||||
}
|
}
|
||||||
} catch (error) {
|
}, 200); // Wait for animation to complete
|
||||||
console.error('Failed to check model status:', error);
|
return () => clearTimeout(timeoutId);
|
||||||
}
|
|
||||||
|
|
||||||
const result = await generation.mutateAsync({
|
|
||||||
profile_id: selectedProfileId,
|
|
||||||
text: data.text,
|
|
||||||
language: data.language,
|
|
||||||
model_size: data.modelSize,
|
|
||||||
});
|
|
||||||
|
|
||||||
toast({
|
|
||||||
title: 'Generation complete!',
|
|
||||||
description: `Audio generated (${result.duration.toFixed(2)}s)`,
|
|
||||||
});
|
|
||||||
|
|
||||||
const audioUrl = apiClient.getAudioUrl(result.id);
|
|
||||||
setAudio(audioUrl, result.id, selectedProfileId, data.text.substring(0, 50));
|
|
||||||
|
|
||||||
form.reset();
|
|
||||||
setIsExpanded(false);
|
|
||||||
} catch (error) {
|
|
||||||
toast({
|
|
||||||
title: 'Generation failed',
|
|
||||||
description: error instanceof Error ? error.message : 'Failed to generate audio',
|
|
||||||
variant: 'destructive',
|
|
||||||
});
|
|
||||||
} finally {
|
|
||||||
setIsGenerating(false);
|
|
||||||
setDownloadingModelName(null);
|
|
||||||
setDownloadingDisplayName(null);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
|
const textarea = textareaRef.current;
|
||||||
|
if (!textarea) return;
|
||||||
|
|
||||||
|
const adjustHeight = () => {
|
||||||
|
textarea.style.height = 'auto';
|
||||||
|
const scrollHeight = textarea.scrollHeight;
|
||||||
|
const minHeight = 100; // Expanded minimum
|
||||||
|
const maxHeight = 300; // Max height in pixels
|
||||||
|
const targetHeight = Math.max(minHeight, Math.min(scrollHeight, maxHeight));
|
||||||
|
textarea.style.height = `${targetHeight}px`;
|
||||||
|
|
||||||
|
// Show scrollbar if content exceeds max height
|
||||||
|
if (scrollHeight > maxHeight) {
|
||||||
|
textarea.style.overflowY = 'auto';
|
||||||
|
} else {
|
||||||
|
textarea.style.overflowY = 'hidden';
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
// Small delay to let framer animation complete
|
||||||
|
const timeoutId = setTimeout(() => {
|
||||||
|
adjustHeight();
|
||||||
|
}, 200);
|
||||||
|
|
||||||
|
// Adjust on mount and when value changes
|
||||||
|
adjustHeight();
|
||||||
|
|
||||||
|
// Watch for input changes
|
||||||
|
textarea.addEventListener('input', adjustHeight);
|
||||||
|
|
||||||
|
return () => {
|
||||||
|
clearTimeout(timeoutId);
|
||||||
|
textarea.removeEventListener('input', adjustHeight);
|
||||||
|
};
|
||||||
|
}, [isExpanded]);
|
||||||
|
|
||||||
|
async function onSubmit(data: Parameters<typeof handleSubmit>[0]) {
|
||||||
|
await handleSubmit(data, selectedProfileId);
|
||||||
}
|
}
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<motion.div
|
<motion.div
|
||||||
ref={containerRef}
|
ref={containerRef}
|
||||||
className="fixed left-[calc(5rem+2rem)] right-auto w-[calc((100%-5rem-4rem)/2-1rem)]"
|
className={cn(
|
||||||
|
'fixed right-auto',
|
||||||
|
isStoriesRoute
|
||||||
|
? // Position aligned with story list: after sidebar + padding, width 360px
|
||||||
|
'left-[calc(5rem+2rem)] w-[360px]'
|
||||||
|
: 'left-[calc(5rem+2rem)] w-[calc((100%-5rem-4rem)/2-1rem)]',
|
||||||
|
)}
|
||||||
style={{
|
style={{
|
||||||
bottom: isPlayerOpen ? 'calc(7rem + 1.5rem)' : '1.5rem',
|
// On stories route: offset by track editor height when visible
|
||||||
|
// On other routes: offset by audio player height when visible
|
||||||
|
bottom: hasTrackEditor
|
||||||
|
? `${trackEditorHeight + 24}px`
|
||||||
|
: isPlayerOpen
|
||||||
|
? 'calc(7rem + 1.5rem)'
|
||||||
|
: '1.5rem',
|
||||||
}}
|
}}
|
||||||
>
|
>
|
||||||
<motion.div
|
<motion.div
|
||||||
@@ -167,51 +194,145 @@ export function FloatingGenerateBox({ isPlayerOpen }: FloatingGenerateBoxProps)
|
|||||||
<form onSubmit={form.handleSubmit(onSubmit)}>
|
<form onSubmit={form.handleSubmit(onSubmit)}>
|
||||||
<div className="flex gap-2">
|
<div className="flex gap-2">
|
||||||
<motion.div
|
<motion.div
|
||||||
className="flex-1"
|
className={cn('flex-1', isExpanded && 'mr-12')}
|
||||||
// animate={{ marginBottom: isExpanded ? '0.75rem' : '0' }}
|
|
||||||
transition={{ duration: 0.3, ease: 'easeOut' }}
|
transition={{ duration: 0.3, ease: 'easeOut' }}
|
||||||
>
|
>
|
||||||
<FormField
|
{/* Text field - hidden when in instruct mode */}
|
||||||
control={form.control}
|
<div style={{ display: isInstructMode ? 'none' : 'block' }}>
|
||||||
name="text"
|
<FormField
|
||||||
render={({ field }) => (
|
control={form.control}
|
||||||
<FormItem>
|
name="text"
|
||||||
<FormControl>
|
render={({ field }) => (
|
||||||
<Textarea
|
<FormItem>
|
||||||
placeholder={
|
<FormControl>
|
||||||
selectedProfile
|
<motion.div
|
||||||
? `Generate speech using ${selectedProfile.name}...`
|
animate={{
|
||||||
: 'Select a voice profile above...'
|
height: isExpanded ? 'auto' : '32px',
|
||||||
}
|
}}
|
||||||
className="resize-none bg-transparent border-none focus-visible:ring-0 focus-visible:ring-offset-0 focus:outline-none focus:ring-0 outline-none ring-0 rounded-2xl text-sm placeholder:text-muted-foreground/60 overflow-hidden transition-all"
|
transition={{ duration: 0.15, ease: 'easeOut' }}
|
||||||
style={{
|
style={{ overflow: 'hidden' }}
|
||||||
minHeight: isExpanded ? '100px' : '32px',
|
>
|
||||||
height: isExpanded ? '100px' : '32px',
|
<Textarea
|
||||||
}}
|
{...field}
|
||||||
disabled={!selectedProfileId}
|
ref={(node: HTMLTextAreaElement | null) => {
|
||||||
onClick={() => setIsExpanded(true)}
|
// Store ref for auto-resize (only for active field)
|
||||||
onFocus={() => setIsExpanded(true)}
|
if (!isInstructMode) {
|
||||||
{...field}
|
textareaRef.current = node;
|
||||||
/>
|
}
|
||||||
</FormControl>
|
// Forward ref to react-hook-form
|
||||||
<FormMessage className="text-xs" />
|
if (typeof field.ref === 'function') {
|
||||||
</FormItem>
|
field.ref(node);
|
||||||
)}
|
}
|
||||||
/>
|
}}
|
||||||
|
placeholder={
|
||||||
|
isStoriesRoute && currentStory
|
||||||
|
? `Generate speech for "${currentStory.name}"...`
|
||||||
|
: selectedProfile
|
||||||
|
? `Generate speech using ${selectedProfile.name}...`
|
||||||
|
: 'Select a voice profile above...'
|
||||||
|
}
|
||||||
|
className="resize-none bg-transparent border-none focus-visible:ring-0 focus-visible:ring-offset-0 focus:outline-none focus:ring-0 outline-none ring-0 rounded-2xl text-sm placeholder:text-muted-foreground/60 w-full"
|
||||||
|
style={{
|
||||||
|
minHeight: isExpanded ? '100px' : '32px',
|
||||||
|
maxHeight: '300px',
|
||||||
|
}}
|
||||||
|
disabled={!selectedProfileId}
|
||||||
|
onClick={() => setIsExpanded(true)}
|
||||||
|
onFocus={() => setIsExpanded(true)}
|
||||||
|
/>
|
||||||
|
</motion.div>
|
||||||
|
</FormControl>
|
||||||
|
<FormMessage className="text-xs" />
|
||||||
|
</FormItem>
|
||||||
|
)}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
{/* Instruct field - hidden when in text mode */}
|
||||||
|
<div style={{ display: isInstructMode ? 'block' : 'none' }}>
|
||||||
|
<FormField
|
||||||
|
control={form.control}
|
||||||
|
name="instruct"
|
||||||
|
render={({ field }) => (
|
||||||
|
<FormItem>
|
||||||
|
<FormControl>
|
||||||
|
<motion.div
|
||||||
|
animate={{
|
||||||
|
height: isExpanded ? 'auto' : '32px',
|
||||||
|
}}
|
||||||
|
transition={{ duration: 0.15, ease: 'easeOut' }}
|
||||||
|
style={{ overflow: 'hidden' }}
|
||||||
|
>
|
||||||
|
<Textarea
|
||||||
|
{...field}
|
||||||
|
ref={(node: HTMLTextAreaElement | null) => {
|
||||||
|
// Store ref for auto-resize (only for active field)
|
||||||
|
if (isInstructMode) {
|
||||||
|
textareaRef.current = node;
|
||||||
|
}
|
||||||
|
// Forward ref to react-hook-form
|
||||||
|
if (typeof field.ref === 'function') {
|
||||||
|
field.ref(node);
|
||||||
|
}
|
||||||
|
}}
|
||||||
|
placeholder="Add delivery instructions..."
|
||||||
|
className="resize-none bg-transparent border-none focus-visible:ring-0 focus-visible:ring-offset-0 focus:outline-none focus:ring-0 outline-none ring-0 rounded-2xl text-sm placeholder:text-muted-foreground/60 w-full"
|
||||||
|
style={{
|
||||||
|
minHeight: isExpanded ? '100px' : '32px',
|
||||||
|
maxHeight: '300px',
|
||||||
|
}}
|
||||||
|
disabled={!selectedProfileId}
|
||||||
|
onClick={() => setIsExpanded(true)}
|
||||||
|
onFocus={() => setIsExpanded(true)}
|
||||||
|
/>
|
||||||
|
</motion.div>
|
||||||
|
</FormControl>
|
||||||
|
<FormMessage className="text-xs" />
|
||||||
|
</FormItem>
|
||||||
|
)}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
</motion.div>
|
</motion.div>
|
||||||
|
|
||||||
<Button
|
<div className="relative shrink-0">
|
||||||
type="submit"
|
<Button
|
||||||
disabled={generation.isPending || !selectedProfileId}
|
type="submit"
|
||||||
className="h-10 w-10 rounded-full bg-accent hover:bg-accent/90 hover:scale-105 text-accent-foreground shadow-lg hover:shadow-accent/50 shrink-0 transition-all duration-200"
|
disabled={isPending || !selectedProfileId}
|
||||||
size="icon"
|
className="h-10 w-10 rounded-full bg-accent hover:bg-accent/90 hover:scale-105 text-accent-foreground shadow-lg hover:shadow-accent/50 transition-all duration-200"
|
||||||
>
|
size="icon"
|
||||||
{generation.isPending ? (
|
>
|
||||||
<Loader2 className="h-4 w-4 animate-spin" />
|
{isPending ? (
|
||||||
) : (
|
<Loader2 className="h-4 w-4 animate-spin" />
|
||||||
<Sparkles className="h-4 w-4" />
|
) : (
|
||||||
)}
|
<Sparkles className="h-4 w-4" />
|
||||||
</Button>
|
)}
|
||||||
|
</Button>
|
||||||
|
<AnimatePresence>
|
||||||
|
{isExpanded && (
|
||||||
|
<motion.div
|
||||||
|
initial={{ opacity: 0, scale: 0.8 }}
|
||||||
|
animate={{ opacity: 1, scale: 1 }}
|
||||||
|
exit={{ opacity: 0, scale: 0.8 }}
|
||||||
|
transition={{ duration: 0.2 }}
|
||||||
|
className="absolute top-0 right-[calc(100%+0.5rem)]"
|
||||||
|
>
|
||||||
|
<Button
|
||||||
|
type="button"
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
onClick={() => setIsInstructMode(!isInstructMode)}
|
||||||
|
className={cn(
|
||||||
|
'h-10 w-10 rounded-full transition-all duration-200',
|
||||||
|
isInstructMode
|
||||||
|
? 'bg-accent text-accent-foreground border border-accent hover:bg-accent/90'
|
||||||
|
: 'bg-card border border-border hover:bg-background/50',
|
||||||
|
)}
|
||||||
|
>
|
||||||
|
<MessageSquare className="h-4 w-4" />
|
||||||
|
</Button>
|
||||||
|
</motion.div>
|
||||||
|
)}
|
||||||
|
</AnimatePresence>
|
||||||
|
</div>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
<AnimatePresence>
|
<AnimatePresence>
|
||||||
@@ -223,11 +344,31 @@ export function FloatingGenerateBox({ isPlayerOpen }: FloatingGenerateBoxProps)
|
|||||||
className=" mt-3"
|
className=" mt-3"
|
||||||
>
|
>
|
||||||
<div className="flex items-center gap-2">
|
<div className="flex items-center gap-2">
|
||||||
|
{showVoiceSelector && (
|
||||||
|
<div className="flex-1">
|
||||||
|
<Select
|
||||||
|
value={selectedProfileId || ''}
|
||||||
|
onValueChange={(value) => setSelectedProfileId(value || null)}
|
||||||
|
>
|
||||||
|
<SelectTrigger className="h-8 text-xs bg-card border-border rounded-full hover:bg-background/50 transition-all w-full">
|
||||||
|
<SelectValue placeholder="Select a voice..." />
|
||||||
|
</SelectTrigger>
|
||||||
|
<SelectContent>
|
||||||
|
{profiles?.map((profile) => (
|
||||||
|
<SelectItem key={profile.id} value={profile.id} className="text-xs">
|
||||||
|
{profile.name}
|
||||||
|
</SelectItem>
|
||||||
|
))}
|
||||||
|
</SelectContent>
|
||||||
|
</Select>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
<FormField
|
<FormField
|
||||||
control={form.control}
|
control={form.control}
|
||||||
name="language"
|
name="language"
|
||||||
render={({ field }) => (
|
render={({ field }) => (
|
||||||
<FormItem className="flex-1">
|
<FormItem className="flex-1 space-y-0">
|
||||||
<Select onValueChange={field.onChange} defaultValue={field.value}>
|
<Select onValueChange={field.onChange} defaultValue={field.value}>
|
||||||
<FormControl>
|
<FormControl>
|
||||||
<SelectTrigger className="h-8 text-xs bg-card border-border rounded-full hover:bg-background/50 transition-all">
|
<SelectTrigger className="h-8 text-xs bg-card border-border rounded-full hover:bg-background/50 transition-all">
|
||||||
@@ -251,7 +392,7 @@ export function FloatingGenerateBox({ isPlayerOpen }: FloatingGenerateBoxProps)
|
|||||||
control={form.control}
|
control={form.control}
|
||||||
name="modelSize"
|
name="modelSize"
|
||||||
render={({ field }) => (
|
render={({ field }) => (
|
||||||
<FormItem className="flex-1">
|
<FormItem className="flex-1 space-y-0">
|
||||||
<Select onValueChange={field.onChange} defaultValue={field.value}>
|
<Select onValueChange={field.onChange} defaultValue={field.value}>
|
||||||
<FormControl>
|
<FormControl>
|
||||||
<SelectTrigger className="h-8 text-xs bg-card border-border rounded-full hover:bg-background/50 transition-all">
|
<SelectTrigger className="h-8 text-xs bg-card border-border rounded-full hover:bg-background/50 transition-all">
|
||||||
|
|||||||
@@ -1,8 +1,4 @@
|
|||||||
import { zodResolver } from '@hookform/resolvers/zod';
|
|
||||||
import { Loader2, Mic } from 'lucide-react';
|
import { Loader2, Mic } from 'lucide-react';
|
||||||
import { useState } from 'react';
|
|
||||||
import { useForm } from 'react-hook-form';
|
|
||||||
import * as z from 'zod';
|
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { Card, CardContent, CardHeader, CardTitle } from '@/components/ui/card';
|
import { Card, CardContent, CardHeader, CardTitle } from '@/components/ui/card';
|
||||||
import {
|
import {
|
||||||
@@ -23,118 +19,19 @@ import {
|
|||||||
SelectValue,
|
SelectValue,
|
||||||
} from '@/components/ui/select';
|
} from '@/components/ui/select';
|
||||||
import { Textarea } from '@/components/ui/textarea';
|
import { Textarea } from '@/components/ui/textarea';
|
||||||
import { useToast } from '@/components/ui/use-toast';
|
import { LANGUAGE_OPTIONS } from '@/lib/constants/languages';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { useGenerationForm } from '@/lib/hooks/useGenerationForm';
|
||||||
import { LANGUAGE_CODES, LANGUAGE_OPTIONS, type LanguageCode } from '@/lib/constants/languages';
|
|
||||||
import { useGeneration } from '@/lib/hooks/useGeneration';
|
|
||||||
import { useModelDownloadToast } from '@/lib/hooks/useModelDownloadToast';
|
|
||||||
import { useProfile } from '@/lib/hooks/useProfiles';
|
import { useProfile } from '@/lib/hooks/useProfiles';
|
||||||
import { useGenerationStore } from '@/stores/generationStore';
|
|
||||||
import { usePlayerStore } from '@/stores/playerStore';
|
|
||||||
import { useUIStore } from '@/stores/uiStore';
|
import { useUIStore } from '@/stores/uiStore';
|
||||||
|
|
||||||
const generationSchema = z.object({
|
|
||||||
text: z.string().min(1, 'Text is required').max(5000),
|
|
||||||
language: z.enum(LANGUAGE_CODES as [LanguageCode, ...LanguageCode[]]),
|
|
||||||
seed: z.number().int().optional(),
|
|
||||||
modelSize: z.enum(['1.7B', '0.6B']).optional(),
|
|
||||||
instruct: z.string().max(500).optional(),
|
|
||||||
});
|
|
||||||
|
|
||||||
type GenerationFormValues = z.infer<typeof generationSchema>;
|
|
||||||
|
|
||||||
export function GenerationForm() {
|
export function GenerationForm() {
|
||||||
const selectedProfileId = useUIStore((state) => state.selectedProfileId);
|
const selectedProfileId = useUIStore((state) => state.selectedProfileId);
|
||||||
const { data: selectedProfile } = useProfile(selectedProfileId || '');
|
const { data: selectedProfile } = useProfile(selectedProfileId || '');
|
||||||
const generation = useGeneration();
|
|
||||||
const { toast } = useToast();
|
|
||||||
const setAudio = usePlayerStore((state) => state.setAudio);
|
|
||||||
const setIsGenerating = useGenerationStore((state) => state.setIsGenerating);
|
|
||||||
const [downloadingModelName, setDownloadingModelName] = useState<string | null>(null);
|
|
||||||
const [downloadingDisplayName, setDownloadingDisplayName] = useState<string | null>(null);
|
|
||||||
|
|
||||||
// Use the download toast hook to show progress when model is downloading
|
const { form, handleSubmit, isPending } = useGenerationForm();
|
||||||
useModelDownloadToast({
|
|
||||||
modelName: downloadingModelName || '',
|
|
||||||
displayName: downloadingDisplayName || '',
|
|
||||||
enabled: !!downloadingModelName,
|
|
||||||
});
|
|
||||||
|
|
||||||
const form = useForm<GenerationFormValues>({
|
async function onSubmit(data: Parameters<typeof handleSubmit>[0]) {
|
||||||
resolver: zodResolver(generationSchema),
|
await handleSubmit(data, selectedProfileId);
|
||||||
defaultValues: {
|
|
||||||
text: '',
|
|
||||||
language: 'en',
|
|
||||||
seed: undefined,
|
|
||||||
modelSize: '1.7B',
|
|
||||||
instruct: '',
|
|
||||||
},
|
|
||||||
});
|
|
||||||
|
|
||||||
async function onSubmit(data: GenerationFormValues) {
|
|
||||||
if (!selectedProfileId) {
|
|
||||||
toast({
|
|
||||||
title: 'No profile selected',
|
|
||||||
description: 'Please select a voice profile from the cards above.',
|
|
||||||
variant: 'destructive',
|
|
||||||
});
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
|
|
||||||
try {
|
|
||||||
setIsGenerating(true);
|
|
||||||
|
|
||||||
// Determine model name and display name
|
|
||||||
const modelName = `qwen-tts-${data.modelSize}`;
|
|
||||||
const displayName = data.modelSize === '1.7B' ? 'Qwen TTS 1.7B' : 'Qwen TTS 0.6B';
|
|
||||||
|
|
||||||
// Check if model is downloaded before starting generation
|
|
||||||
try {
|
|
||||||
const modelStatus = await apiClient.getModelStatus();
|
|
||||||
const model = modelStatus.models.find((m) => m.model_name === modelName);
|
|
||||||
|
|
||||||
if (model && !model.downloaded) {
|
|
||||||
// Model is not downloaded, enable download toast
|
|
||||||
setDownloadingModelName(modelName);
|
|
||||||
setDownloadingDisplayName(displayName);
|
|
||||||
}
|
|
||||||
} catch (error) {
|
|
||||||
// If status check fails, continue anyway - generation will handle it
|
|
||||||
console.error('Failed to check model status:', error);
|
|
||||||
}
|
|
||||||
|
|
||||||
// Proceed with generation (which will trigger download if needed)
|
|
||||||
const result = await generation.mutateAsync({
|
|
||||||
profile_id: selectedProfileId,
|
|
||||||
text: data.text,
|
|
||||||
language: data.language,
|
|
||||||
seed: data.seed,
|
|
||||||
model_size: data.modelSize,
|
|
||||||
instruct: data.instruct || undefined,
|
|
||||||
});
|
|
||||||
|
|
||||||
toast({
|
|
||||||
title: 'Generation complete!',
|
|
||||||
description: `Audio generated (${result.duration.toFixed(2)}s)`,
|
|
||||||
});
|
|
||||||
|
|
||||||
// Autoplay the generated audio
|
|
||||||
const audioUrl = apiClient.getAudioUrl(result.id);
|
|
||||||
setAudio(audioUrl, result.id, selectedProfileId, data.text.substring(0, 50));
|
|
||||||
|
|
||||||
form.reset();
|
|
||||||
} catch (error) {
|
|
||||||
toast({
|
|
||||||
title: 'Generation failed',
|
|
||||||
description: error instanceof Error ? error.message : 'Failed to generate audio',
|
|
||||||
variant: 'destructive',
|
|
||||||
});
|
|
||||||
} finally {
|
|
||||||
setIsGenerating(false);
|
|
||||||
// Clear download state after generation completes
|
|
||||||
setDownloadingModelName(null);
|
|
||||||
setDownloadingDisplayName(null);
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
return (
|
return (
|
||||||
@@ -276,9 +173,9 @@ export function GenerationForm() {
|
|||||||
<Button
|
<Button
|
||||||
type="submit"
|
type="submit"
|
||||||
className="w-full"
|
className="w-full"
|
||||||
disabled={generation.isPending || !selectedProfileId}
|
disabled={isPending || !selectedProfileId}
|
||||||
>
|
>
|
||||||
{generation.isPending ? (
|
{isPending ? (
|
||||||
<>
|
<>
|
||||||
<Loader2 className="mr-2 h-4 w-4 animate-spin" />
|
<Loader2 className="mr-2 h-4 w-4 animate-spin" />
|
||||||
Generating...
|
Generating...
|
||||||
|
|||||||
@@ -1,5 +1,6 @@
|
|||||||
import { AudioWaveform, Download, FileArchive, MoreHorizontal, Play, Trash2 } from 'lucide-react';
|
import { AudioWaveform, Download, FileArchive, Loader2, MoreHorizontal, Play, Trash2 } from 'lucide-react';
|
||||||
import { useEffect, useRef, useState } from 'react';
|
import { useEffect, useRef, useState } from 'react';
|
||||||
|
import type { HistoryResponse } from '@/lib/api/types';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import {
|
import {
|
||||||
Dialog,
|
Dialog,
|
||||||
@@ -16,6 +17,7 @@ import {
|
|||||||
DropdownMenuTrigger,
|
DropdownMenuTrigger,
|
||||||
} from '@/components/ui/dropdown-menu';
|
} from '@/components/ui/dropdown-menu';
|
||||||
import { Textarea } from '@/components/ui/textarea';
|
import { Textarea } from '@/components/ui/textarea';
|
||||||
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { apiClient } from '@/lib/api/client';
|
||||||
import { BOTTOM_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
import { BOTTOM_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
||||||
import {
|
import {
|
||||||
@@ -32,17 +34,23 @@ import { usePlayerStore } from '@/stores/playerStore';
|
|||||||
// OLD TABLE-BASED COMPONENT - REMOVED (can be found in git history)
|
// OLD TABLE-BASED COMPONENT - REMOVED (can be found in git history)
|
||||||
// This is the new alternate history view with fixed height rows
|
// This is the new alternate history view with fixed height rows
|
||||||
|
|
||||||
// NEW ALTERNATE HISTORY VIEW - FIXED HEIGHT ROWS
|
// NEW ALTERNATE HISTORY VIEW - FIXED HEIGHT ROWS WITH INFINITE SCROLL
|
||||||
export function HistoryTable() {
|
export function HistoryTable() {
|
||||||
const [page, _setPage] = useState(0);
|
const [page, setPage] = useState(0);
|
||||||
|
const [allHistory, setAllHistory] = useState<HistoryResponse[]>([]);
|
||||||
|
const [total, setTotal] = useState(0);
|
||||||
const [isScrolled, setIsScrolled] = useState(false);
|
const [isScrolled, setIsScrolled] = useState(false);
|
||||||
const scrollRef = useRef<HTMLDivElement>(null);
|
const scrollRef = useRef<HTMLDivElement>(null);
|
||||||
|
const loadMoreRef = useRef<HTMLDivElement>(null);
|
||||||
const fileInputRef = useRef<HTMLInputElement>(null);
|
const fileInputRef = useRef<HTMLInputElement>(null);
|
||||||
const [importDialogOpen, setImportDialogOpen] = useState(false);
|
const [importDialogOpen, setImportDialogOpen] = useState(false);
|
||||||
const [selectedFile, setSelectedFile] = useState<File | null>(null);
|
const [selectedFile, setSelectedFile] = useState<File | null>(null);
|
||||||
|
const [deleteDialogOpen, setDeleteDialogOpen] = useState(false);
|
||||||
|
const [generationToDelete, setGenerationToDelete] = useState<{ id: string; name: string } | null>(null);
|
||||||
const limit = 20;
|
const limit = 20;
|
||||||
|
const { toast } = useToast();
|
||||||
|
|
||||||
const { data: historyData, isLoading } = useHistory({
|
const { data: historyData, isLoading, isFetching } = useHistory({
|
||||||
limit,
|
limit,
|
||||||
offset: page * limit,
|
offset: page * limit,
|
||||||
});
|
});
|
||||||
@@ -51,13 +59,63 @@ export function HistoryTable() {
|
|||||||
const exportGeneration = useExportGeneration();
|
const exportGeneration = useExportGeneration();
|
||||||
const exportGenerationAudio = useExportGenerationAudio();
|
const exportGenerationAudio = useExportGenerationAudio();
|
||||||
const importGeneration = useImportGeneration();
|
const importGeneration = useImportGeneration();
|
||||||
const setAudio = usePlayerStore((state) => state.setAudio);
|
const setAudioWithAutoPlay = usePlayerStore((state) => state.setAudioWithAutoPlay);
|
||||||
const restartCurrentAudio = usePlayerStore((state) => state.restartCurrentAudio);
|
const restartCurrentAudio = usePlayerStore((state) => state.restartCurrentAudio);
|
||||||
const currentAudioId = usePlayerStore((state) => state.audioId);
|
const currentAudioId = usePlayerStore((state) => state.audioId);
|
||||||
const isPlaying = usePlayerStore((state) => state.isPlaying);
|
const isPlaying = usePlayerStore((state) => state.isPlaying);
|
||||||
const audioUrl = usePlayerStore((state) => state.audioUrl);
|
const audioUrl = usePlayerStore((state) => state.audioUrl);
|
||||||
const isPlayerVisible = !!audioUrl;
|
const isPlayerVisible = !!audioUrl;
|
||||||
|
|
||||||
|
// Update accumulated history when new data arrives
|
||||||
|
useEffect(() => {
|
||||||
|
if (historyData?.items) {
|
||||||
|
setTotal(historyData.total);
|
||||||
|
if (page === 0) {
|
||||||
|
// Reset to first page
|
||||||
|
setAllHistory(historyData.items);
|
||||||
|
} else {
|
||||||
|
// Append new items, avoiding duplicates
|
||||||
|
setAllHistory((prev) => {
|
||||||
|
const existingIds = new Set(prev.map((item) => item.id));
|
||||||
|
const newItems = historyData.items.filter((item) => !existingIds.has(item.id));
|
||||||
|
return [...prev, ...newItems];
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}, [historyData, page]);
|
||||||
|
|
||||||
|
// Reset to page 0 when deletions or imports occur
|
||||||
|
useEffect(() => {
|
||||||
|
if (deleteGeneration.isSuccess || importGeneration.isSuccess) {
|
||||||
|
setPage(0);
|
||||||
|
setAllHistory([]);
|
||||||
|
}
|
||||||
|
}, [deleteGeneration.isSuccess, importGeneration.isSuccess]);
|
||||||
|
|
||||||
|
// Intersection Observer for infinite scroll
|
||||||
|
useEffect(() => {
|
||||||
|
const loadMoreEl = loadMoreRef.current;
|
||||||
|
if (!loadMoreEl) return;
|
||||||
|
|
||||||
|
const observer = new IntersectionObserver(
|
||||||
|
(entries) => {
|
||||||
|
const target = entries[0];
|
||||||
|
if (target.isIntersecting && !isFetching && allHistory.length < total) {
|
||||||
|
setPage((prev) => prev + 1);
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
root: scrollRef.current,
|
||||||
|
rootMargin: '100px',
|
||||||
|
threshold: 0.1,
|
||||||
|
},
|
||||||
|
);
|
||||||
|
|
||||||
|
observer.observe(loadMoreEl);
|
||||||
|
return () => observer.disconnect();
|
||||||
|
}, [isFetching, allHistory.length, total]);
|
||||||
|
|
||||||
|
// Track scroll position for gradient effect
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
const scrollEl = scrollRef.current;
|
const scrollEl = scrollRef.current;
|
||||||
if (!scrollEl) return;
|
if (!scrollEl) return;
|
||||||
@@ -75,9 +133,9 @@ export function HistoryTable() {
|
|||||||
if (currentAudioId === audioId) {
|
if (currentAudioId === audioId) {
|
||||||
restartCurrentAudio();
|
restartCurrentAudio();
|
||||||
} else {
|
} else {
|
||||||
// Otherwise, load the new audio
|
// Otherwise, load the new audio and auto-play it
|
||||||
const audioUrl = apiClient.getAudioUrl(audioId);
|
const audioUrl = apiClient.getAudioUrl(audioId);
|
||||||
setAudio(audioUrl, audioId, profileId, text.substring(0, 50));
|
setAudioWithAutoPlay(audioUrl, audioId, profileId, text.substring(0, 50));
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -86,7 +144,11 @@ export function HistoryTable() {
|
|||||||
{ generationId, text },
|
{ generationId, text },
|
||||||
{
|
{
|
||||||
onError: (error) => {
|
onError: (error) => {
|
||||||
alert(`Failed to download audio: ${error.message}`);
|
toast({
|
||||||
|
title: 'Failed to download audio',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
);
|
);
|
||||||
@@ -97,26 +159,26 @@ export function HistoryTable() {
|
|||||||
{ generationId, text },
|
{ generationId, text },
|
||||||
{
|
{
|
||||||
onError: (error) => {
|
onError: (error) => {
|
||||||
alert(`Failed to export generation: ${error.message}`);
|
toast({
|
||||||
|
title: 'Failed to export generation',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
);
|
);
|
||||||
};
|
};
|
||||||
|
|
||||||
const _handleImportClick = () => {
|
const handleDeleteClick = (generationId: string, profileName: string) => {
|
||||||
file_handleImportClickk.click();
|
setGenerationToDelete({ id: generationId, name: profileName });
|
||||||
|
setDeleteDialogOpen(true);
|
||||||
};
|
};
|
||||||
|
|
||||||
const _handleFileChange = (_e: React.ChangeEvent<HTMLInputElement>) => {
|
const handleDeleteConfirm = () => {
|
||||||
cons_handleFileChangeet.files?.[0];
|
if (generationToDelete) {
|
||||||
if (file) {
|
deleteGeneration.mutate(generationToDelete.id);
|
||||||
// Validate file extension
|
setDeleteDialogOpen(false);
|
||||||
if (!file.name.endsWith('.voicebox.zip')) {
|
setGenerationToDelete(null);
|
||||||
alert('Please select a valid .voicebox.zip file');
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
setSelectedFile(file);
|
|
||||||
setImportDialogOpen(true);
|
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -129,42 +191,35 @@ export function HistoryTable() {
|
|||||||
if (fileInputRef.current) {
|
if (fileInputRef.current) {
|
||||||
fileInputRef.current.value = '';
|
fileInputRef.current.value = '';
|
||||||
}
|
}
|
||||||
alert(data.message || 'Generation imported successfully');
|
toast({
|
||||||
|
title: 'Generation imported',
|
||||||
|
description: data.message || 'Generation imported successfully',
|
||||||
|
});
|
||||||
},
|
},
|
||||||
onError: (error) => {
|
onError: (error) => {
|
||||||
alert(`Failed to import generation: ${error.message}`);
|
toast({
|
||||||
|
title: 'Failed to import generation',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
if (isLoading) {
|
if (isLoading && page === 0) {
|
||||||
return null;
|
return (
|
||||||
|
<div className="flex items-center justify-center h-full">
|
||||||
|
<Loader2 className="h-8 w-8 animate-spin text-muted-foreground" />
|
||||||
|
</div>
|
||||||
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
const history = historyData?.items || [];
|
const history = allHistory;
|
||||||
const total = historyData?.total || 0;
|
const hasMore = allHistory.length < total;
|
||||||
const _hasMore = history.length === limit && (page + 1) * limit < total;
|
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<div className="flex flex-col h-full min-h-0 relative">
|
<div className="flex flex-col h-full min-h-0 relative">
|
||||||
{/* <div className="flex justify-between items-center mb-4 shrink-0">
|
|
||||||
<h2 className="text-2xl font-bold">History</h2>
|
|
||||||
<div className="flex gap-2">
|
|
||||||
<Button variant="outline" onClick={handleImportClick}>
|
|
||||||
<Upload className="mr-2 h-4 w-4" />
|
|
||||||
Import Generation
|
|
||||||
</Button>
|
|
||||||
<input
|
|
||||||
ref={fileInputRef}
|
|
||||||
type="file"
|
|
||||||
accept=".voicebox.zip"
|
|
||||||
onChange={handleFileChange}
|
|
||||||
className="hidden"
|
|
||||||
/>
|
|
||||||
</div>
|
|
||||||
</div> */}
|
|
||||||
|
|
||||||
{history.length === 0 ? (
|
{history.length === 0 ? (
|
||||||
<div className="text-center py-12 px-5 border-2 border-dashed mb-5 border-muted rounded-md text-muted-foreground flex-1 flex items-center justify-center">
|
<div className="text-center py-12 px-5 border-2 border-dashed mb-5 border-muted rounded-md text-muted-foreground flex-1 flex items-center justify-center">
|
||||||
No voice generations, yet...
|
No voice generations, yet...
|
||||||
@@ -229,7 +284,11 @@ export function HistoryTable() {
|
|||||||
</div>
|
</div>
|
||||||
|
|
||||||
{/* Far right - Ellipsis actions */}
|
{/* Far right - Ellipsis actions */}
|
||||||
<div className="w-10 shrink-0 flex justify-end">
|
<div
|
||||||
|
className="w-10 shrink-0 flex justify-end"
|
||||||
|
onMouseDown={(e) => e.stopPropagation()}
|
||||||
|
onClick={(e) => e.stopPropagation()}
|
||||||
|
>
|
||||||
<DropdownMenu>
|
<DropdownMenu>
|
||||||
<DropdownMenuTrigger asChild>
|
<DropdownMenuTrigger asChild>
|
||||||
<Button
|
<Button
|
||||||
@@ -237,7 +296,6 @@ export function HistoryTable() {
|
|||||||
size="icon"
|
size="icon"
|
||||||
className="h-8 w-8"
|
className="h-8 w-8"
|
||||||
aria-label="Actions"
|
aria-label="Actions"
|
||||||
onClick={(e) => e.stopPropagation()}
|
|
||||||
>
|
>
|
||||||
<MoreHorizontal className="h-4 w-4" />
|
<MoreHorizontal className="h-4 w-4" />
|
||||||
</Button>
|
</Button>
|
||||||
@@ -264,7 +322,7 @@ export function HistoryTable() {
|
|||||||
Export Package
|
Export Package
|
||||||
</DropdownMenuItem>
|
</DropdownMenuItem>
|
||||||
<DropdownMenuItem
|
<DropdownMenuItem
|
||||||
onClick={() => deleteGeneration.mutate(gen.id)}
|
onClick={() => handleDeleteClick(gen.id, gen.profile_name)}
|
||||||
disabled={deleteGeneration.isPending}
|
disabled={deleteGeneration.isPending}
|
||||||
className="text-destructive focus:text-destructive"
|
className="text-destructive focus:text-destructive"
|
||||||
>
|
>
|
||||||
@@ -277,10 +335,53 @@ export function HistoryTable() {
|
|||||||
</div>
|
</div>
|
||||||
);
|
);
|
||||||
})}
|
})}
|
||||||
|
|
||||||
|
{/* Load more trigger element */}
|
||||||
|
{hasMore && (
|
||||||
|
<div ref={loadMoreRef} className="flex items-center justify-center py-4">
|
||||||
|
{isFetching && <Loader2 className="h-6 w-6 animate-spin text-muted-foreground" />}
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{/* End of list indicator */}
|
||||||
|
{!hasMore && history.length > 0 && (
|
||||||
|
<div className="text-center py-4 text-xs text-muted-foreground">
|
||||||
|
You've reached the end
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
</div>
|
</div>
|
||||||
</>
|
</>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
|
<Dialog open={deleteDialogOpen} onOpenChange={setDeleteDialogOpen}>
|
||||||
|
<DialogContent>
|
||||||
|
<DialogHeader>
|
||||||
|
<DialogTitle>Delete Generation</DialogTitle>
|
||||||
|
<DialogDescription>
|
||||||
|
Are you sure you want to delete this generation from "{generationToDelete?.name}"? This action cannot be undone.
|
||||||
|
</DialogDescription>
|
||||||
|
</DialogHeader>
|
||||||
|
<DialogFooter>
|
||||||
|
<Button
|
||||||
|
variant="outline"
|
||||||
|
onClick={() => {
|
||||||
|
setDeleteDialogOpen(false);
|
||||||
|
setGenerationToDelete(null);
|
||||||
|
}}
|
||||||
|
>
|
||||||
|
Cancel
|
||||||
|
</Button>
|
||||||
|
<Button
|
||||||
|
variant="destructive"
|
||||||
|
onClick={handleDeleteConfirm}
|
||||||
|
disabled={deleteGeneration.isPending}
|
||||||
|
>
|
||||||
|
{deleteGeneration.isPending ? 'Deleting...' : 'Delete'}
|
||||||
|
</Button>
|
||||||
|
</DialogFooter>
|
||||||
|
</DialogContent>
|
||||||
|
</Dialog>
|
||||||
|
|
||||||
<Dialog open={importDialogOpen} onOpenChange={setImportDialogOpen}>
|
<Dialog open={importDialogOpen} onOpenChange={setImportDialogOpen}>
|
||||||
<DialogContent>
|
<DialogContent>
|
||||||
<DialogHeader>
|
<DialogHeader>
|
||||||
|
|||||||
@@ -11,6 +11,7 @@ import {
|
|||||||
DialogHeader,
|
DialogHeader,
|
||||||
DialogTitle,
|
DialogTitle,
|
||||||
} from '@/components/ui/dialog';
|
} from '@/components/ui/dialog';
|
||||||
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
import { ProfileList } from '@/components/VoiceProfiles/ProfileList';
|
import { ProfileList } from '@/components/VoiceProfiles/ProfileList';
|
||||||
import { BOTTOM_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
import { BOTTOM_SAFE_AREA_PADDING } from '@/lib/constants/ui';
|
||||||
import { useImportProfile } from '@/lib/hooks/useProfiles';
|
import { useImportProfile } from '@/lib/hooks/useProfiles';
|
||||||
@@ -27,6 +28,7 @@ export function MainEditor() {
|
|||||||
const fileInputRef = useRef<HTMLInputElement>(null);
|
const fileInputRef = useRef<HTMLInputElement>(null);
|
||||||
const [importDialogOpen, setImportDialogOpen] = useState(false);
|
const [importDialogOpen, setImportDialogOpen] = useState(false);
|
||||||
const [selectedFile, setSelectedFile] = useState<File | null>(null);
|
const [selectedFile, setSelectedFile] = useState<File | null>(null);
|
||||||
|
const { toast } = useToast();
|
||||||
|
|
||||||
const handleImportClick = () => {
|
const handleImportClick = () => {
|
||||||
fileInputRef.current?.click();
|
fileInputRef.current?.click();
|
||||||
@@ -36,7 +38,11 @@ export function MainEditor() {
|
|||||||
const file = e.target.files?.[0];
|
const file = e.target.files?.[0];
|
||||||
if (file) {
|
if (file) {
|
||||||
if (!file.name.endsWith('.voicebox.zip')) {
|
if (!file.name.endsWith('.voicebox.zip')) {
|
||||||
alert('Please select a valid .voicebox.zip file');
|
toast({
|
||||||
|
title: 'Invalid file type',
|
||||||
|
description: 'Please select a valid .voicebox.zip file',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
setSelectedFile(file);
|
setSelectedFile(file);
|
||||||
@@ -53,9 +59,17 @@ export function MainEditor() {
|
|||||||
if (fileInputRef.current) {
|
if (fileInputRef.current) {
|
||||||
fileInputRef.current.value = '';
|
fileInputRef.current.value = '';
|
||||||
}
|
}
|
||||||
|
toast({
|
||||||
|
title: 'Profile imported',
|
||||||
|
description: 'Voice profile imported successfully',
|
||||||
|
});
|
||||||
},
|
},
|
||||||
onError: (error) => {
|
onError: (error) => {
|
||||||
alert(`Failed to import profile: ${error.message}`);
|
toast({
|
||||||
|
title: 'Failed to import profile',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
@@ -102,15 +116,9 @@ export function MainEditor() {
|
|||||||
)}
|
)}
|
||||||
>
|
>
|
||||||
<div className="flex flex-col gap-6">
|
<div className="flex flex-col gap-6">
|
||||||
{/* Profiles - Top Left */}
|
|
||||||
<div className="shrink-0 flex flex-col">
|
<div className="shrink-0 flex flex-col">
|
||||||
<ProfileList />
|
<ProfileList />
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
{/* Generator - Bottom Left */}
|
|
||||||
{/* <div className="shrink-0">
|
|
||||||
<GenerationForm />
|
|
||||||
</div> */}
|
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ import { Input } from '@/components/ui/input';
|
|||||||
import { Checkbox } from '@/components/ui/checkbox';
|
import { Checkbox } from '@/components/ui/checkbox';
|
||||||
import { useToast } from '@/components/ui/use-toast';
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
import { useServerStore } from '@/stores/serverStore';
|
import { useServerStore } from '@/stores/serverStore';
|
||||||
import { setKeepServerRunning } from '@/lib/tauri';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
const connectionSchema = z.object({
|
const connectionSchema = z.object({
|
||||||
serverUrl: z.string().url('Please enter a valid URL'),
|
serverUrl: z.string().url('Please enter a valid URL'),
|
||||||
@@ -26,6 +26,7 @@ const connectionSchema = z.object({
|
|||||||
type ConnectionFormValues = z.infer<typeof connectionSchema>;
|
type ConnectionFormValues = z.infer<typeof connectionSchema>;
|
||||||
|
|
||||||
export function ConnectionForm() {
|
export function ConnectionForm() {
|
||||||
|
const platform = usePlatform();
|
||||||
const serverUrl = useServerStore((state) => state.serverUrl);
|
const serverUrl = useServerStore((state) => state.serverUrl);
|
||||||
const setServerUrl = useServerStore((state) => state.setServerUrl);
|
const setServerUrl = useServerStore((state) => state.setServerUrl);
|
||||||
const keepServerRunningOnClose = useServerStore((state) => state.keepServerRunningOnClose);
|
const keepServerRunningOnClose = useServerStore((state) => state.keepServerRunningOnClose);
|
||||||
@@ -89,7 +90,7 @@ export function ConnectionForm() {
|
|||||||
checked={keepServerRunningOnClose}
|
checked={keepServerRunningOnClose}
|
||||||
onCheckedChange={(checked: boolean) => {
|
onCheckedChange={(checked: boolean) => {
|
||||||
setKeepServerRunningOnClose(checked);
|
setKeepServerRunningOnClose(checked);
|
||||||
setKeepServerRunning(checked).catch((error) => {
|
platform.lifecycle.setKeepServerRunning(checked).catch((error) => {
|
||||||
console.error('Failed to sync setting to Rust:', error);
|
console.error('Failed to sync setting to Rust:', error);
|
||||||
});
|
});
|
||||||
toast({
|
toast({
|
||||||
|
|||||||
@@ -12,11 +12,10 @@ interface ModelProgressProps {
|
|||||||
|
|
||||||
export function ModelProgress({ modelName, displayName }: ModelProgressProps) {
|
export function ModelProgress({ modelName, displayName }: ModelProgressProps) {
|
||||||
const [progress, setProgress] = useState<ModelProgressType | null>(null);
|
const [progress, setProgress] = useState<ModelProgressType | null>(null);
|
||||||
const [isSubscribed, setIsSubscribed] = useState(false);
|
|
||||||
const serverUrl = useServerStore((state) => state.serverUrl);
|
const serverUrl = useServerStore((state) => state.serverUrl);
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!serverUrl || isSubscribed) return;
|
if (!serverUrl) return;
|
||||||
|
|
||||||
// Subscribe to progress updates via Server-Sent Events
|
// Subscribe to progress updates via Server-Sent Events
|
||||||
const eventSource = new EventSource(`${serverUrl}/models/progress/${modelName}`);
|
const eventSource = new EventSource(`${serverUrl}/models/progress/${modelName}`);
|
||||||
@@ -29,7 +28,6 @@ export function ModelProgress({ modelName, displayName }: ModelProgressProps) {
|
|||||||
// Close connection if complete or error
|
// Close connection if complete or error
|
||||||
if (data.status === 'complete' || data.status === 'error') {
|
if (data.status === 'complete' || data.status === 'error') {
|
||||||
eventSource.close();
|
eventSource.close();
|
||||||
setIsSubscribed(false);
|
|
||||||
}
|
}
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
console.error('Error parsing progress event:', error);
|
console.error('Error parsing progress event:', error);
|
||||||
@@ -39,16 +37,12 @@ export function ModelProgress({ modelName, displayName }: ModelProgressProps) {
|
|||||||
eventSource.onerror = (error) => {
|
eventSource.onerror = (error) => {
|
||||||
console.error('SSE error:', error);
|
console.error('SSE error:', error);
|
||||||
eventSource.close();
|
eventSource.close();
|
||||||
setIsSubscribed(false);
|
|
||||||
};
|
};
|
||||||
|
|
||||||
setIsSubscribed(true);
|
|
||||||
|
|
||||||
return () => {
|
return () => {
|
||||||
eventSource.close();
|
eventSource.close();
|
||||||
setIsSubscribed(false);
|
|
||||||
};
|
};
|
||||||
}, [serverUrl, modelName, isSubscribed]);
|
}, [serverUrl, modelName]);
|
||||||
|
|
||||||
// Don't render if no progress or if complete/error and some time has passed
|
// Don't render if no progress or if complete/error and some time has passed
|
||||||
if (
|
if (
|
||||||
|
|||||||
@@ -1,21 +1,22 @@
|
|||||||
import { getVersion } from '@tauri-apps/api/app';
|
import { AlertCircle, Download, RefreshCw } from 'lucide-react';
|
||||||
import { RefreshCw, Download, AlertCircle } from 'lucide-react';
|
|
||||||
import { useEffect, useState } from 'react';
|
import { useEffect, useState } from 'react';
|
||||||
import { Badge } from '@/components/ui/badge';
|
import { Badge } from '@/components/ui/badge';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { Card, CardContent, CardHeader, CardTitle } from '@/components/ui/card';
|
import { Card, CardContent, CardHeader, CardTitle } from '@/components/ui/card';
|
||||||
import { Progress } from '@/components/ui/progress';
|
import { Progress } from '@/components/ui/progress';
|
||||||
import { useAutoUpdater } from '@/hooks/useAutoUpdater';
|
import { useAutoUpdater } from '@/hooks/useAutoUpdater';
|
||||||
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
export function UpdateStatus() {
|
export function UpdateStatus() {
|
||||||
|
const platform = usePlatform();
|
||||||
const { status, checkForUpdates, downloadAndInstall, restartAndInstall } = useAutoUpdater(false);
|
const { status, checkForUpdates, downloadAndInstall, restartAndInstall } = useAutoUpdater(false);
|
||||||
const [currentVersion, setCurrentVersion] = useState<string>('');
|
const [currentVersion, setCurrentVersion] = useState<string>('');
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
getVersion()
|
platform.metadata.getVersion()
|
||||||
.then(setCurrentVersion)
|
.then(setCurrentVersion)
|
||||||
.catch(() => setCurrentVersion('0.1.0'));
|
.catch(() => setCurrentVersion('0.1.0'));
|
||||||
}, []);
|
}, [platform]);
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<Card>
|
<Card>
|
||||||
@@ -93,7 +94,7 @@ export function UpdateStatus() {
|
|||||||
)}
|
)}
|
||||||
|
|
||||||
{status.readyToInstall && (
|
{status.readyToInstall && (
|
||||||
<div className="space-y-3 p-4 border rounded-lg bg-green-500/10 border-green-500/20">
|
<div className="space-y-3 p-4 border rounded-lg bg-accent/30 border-accent/50">
|
||||||
<div className="flex items-center gap-2">
|
<div className="flex items-center gap-2">
|
||||||
<div>
|
<div>
|
||||||
<div className="font-semibold">Update Ready to Install</div>
|
<div className="font-semibold">Update Ready to Install</div>
|
||||||
|
|||||||
@@ -1,16 +1,17 @@
|
|||||||
import { ConnectionForm } from '@/components/ServerSettings/ConnectionForm';
|
import { ConnectionForm } from '@/components/ServerSettings/ConnectionForm';
|
||||||
import { ServerStatus } from '@/components/ServerSettings/ServerStatus';
|
import { ServerStatus } from '@/components/ServerSettings/ServerStatus';
|
||||||
import { UpdateStatus } from '@/components/ServerSettings/UpdateStatus';
|
import { UpdateStatus } from '@/components/ServerSettings/UpdateStatus';
|
||||||
import { isTauri } from '@/lib/tauri';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
export function ServerTab() {
|
export function ServerTab() {
|
||||||
|
const platform = usePlatform();
|
||||||
return (
|
return (
|
||||||
<div className="space-y-4 overflow-y-auto flex flex-col">
|
<div className="space-y-4 overflow-y-auto flex flex-col">
|
||||||
<div className="grid gap-4 md:grid-cols-2">
|
<div className="grid gap-4 md:grid-cols-2">
|
||||||
<ConnectionForm />
|
<ConnectionForm />
|
||||||
<ServerStatus />
|
<ServerStatus />
|
||||||
</div>
|
</div>
|
||||||
{isTauri() && <UpdateStatus />}
|
{platform.metadata.isTauri && <UpdateStatus />}
|
||||||
<div className="py-8 text-center text-sm text-muted-foreground">
|
<div className="py-8 text-center text-sm text-muted-foreground">
|
||||||
Created by{' '}
|
Created by{' '}
|
||||||
<a
|
<a
|
||||||
|
|||||||
@@ -1,27 +1,28 @@
|
|||||||
import { Box, Loader2, Mic, Server, Speaker, Volume2 } from 'lucide-react';
|
import { Link, useMatchRoute } from '@tanstack/react-router';
|
||||||
|
import { Box, BookOpen, Loader2, Mic, Server, Speaker, Volume2 } from 'lucide-react';
|
||||||
import voiceboxLogo from '@/assets/voicebox-logo.png';
|
import voiceboxLogo from '@/assets/voicebox-logo.png';
|
||||||
import { cn } from '@/lib/utils/cn';
|
import { cn } from '@/lib/utils/cn';
|
||||||
import { useGenerationStore } from '@/stores/generationStore';
|
import { useGenerationStore } from '@/stores/generationStore';
|
||||||
import { usePlayerStore } from '@/stores/playerStore';
|
import { usePlayerStore } from '@/stores/playerStore';
|
||||||
|
|
||||||
interface SidebarProps {
|
interface SidebarProps {
|
||||||
activeTab: string;
|
|
||||||
onTabChange: (tab: string) => void;
|
|
||||||
isMacOS?: boolean;
|
isMacOS?: boolean;
|
||||||
}
|
}
|
||||||
|
|
||||||
const tabs = [
|
const tabs = [
|
||||||
{ id: 'main', icon: Volume2, label: 'Generate' },
|
{ id: 'main', path: '/', icon: Volume2, label: 'Generate' },
|
||||||
{ id: 'voices', icon: Mic, label: 'Voices' },
|
{ id: 'stories', path: '/stories', icon: BookOpen, label: 'Stories' },
|
||||||
{ id: 'audio', icon: Speaker, label: 'Audio' },
|
{ id: 'voices', path: '/voices', icon: Mic, label: 'Voices' },
|
||||||
{ id: 'models', icon: Box, label: 'Models' },
|
{ id: 'audio', path: '/audio', icon: Speaker, label: 'Audio' },
|
||||||
{ id: 'server', icon: Server, label: 'Server' },
|
{ id: 'models', path: '/models', icon: Box, label: 'Models' },
|
||||||
|
{ id: 'server', path: '/server', icon: Server, label: 'Server' },
|
||||||
];
|
];
|
||||||
|
|
||||||
export function Sidebar({ activeTab, onTabChange, isMacOS }: SidebarProps) {
|
export function Sidebar({ isMacOS }: SidebarProps) {
|
||||||
const isGenerating = useGenerationStore((state) => state.isGenerating);
|
const isGenerating = useGenerationStore((state) => state.isGenerating);
|
||||||
const audioUrl = usePlayerStore((state) => state.audioUrl);
|
const audioUrl = usePlayerStore((state) => state.audioUrl);
|
||||||
const isPlayerVisible = !!audioUrl;
|
const isPlayerVisible = !!audioUrl;
|
||||||
|
const matchRoute = useMatchRoute();
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<div
|
<div
|
||||||
@@ -39,13 +40,16 @@ export function Sidebar({ activeTab, onTabChange, isMacOS }: SidebarProps) {
|
|||||||
<div className="flex flex-col gap-3">
|
<div className="flex flex-col gap-3">
|
||||||
{tabs.map((tab) => {
|
{tabs.map((tab) => {
|
||||||
const Icon = tab.icon;
|
const Icon = tab.icon;
|
||||||
const isActive = activeTab === tab.id;
|
// For index route, use exact match; for others, use default matching
|
||||||
|
const isActive =
|
||||||
|
tab.path === '/'
|
||||||
|
? matchRoute({ to: '/', exact: true })
|
||||||
|
: matchRoute({ to: tab.path });
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<button
|
<Link
|
||||||
key={tab.id}
|
key={tab.id}
|
||||||
type="button"
|
to={tab.path}
|
||||||
onClick={() => onTabChange(tab.id)}
|
|
||||||
className={cn(
|
className={cn(
|
||||||
'w-12 h-12 rounded-full flex items-center justify-center transition-all duration-200',
|
'w-12 h-12 rounded-full flex items-center justify-center transition-all duration-200',
|
||||||
'hover:bg-muted/50',
|
'hover:bg-muted/50',
|
||||||
@@ -55,7 +59,7 @@ export function Sidebar({ activeTab, onTabChange, isMacOS }: SidebarProps) {
|
|||||||
aria-label={tab.label}
|
aria-label={tab.label}
|
||||||
>
|
>
|
||||||
<Icon className="h-5 w-5" />
|
<Icon className="h-5 w-5" />
|
||||||
</button>
|
</Link>
|
||||||
);
|
);
|
||||||
})}
|
})}
|
||||||
</div>
|
</div>
|
||||||
|
|||||||
@@ -0,0 +1,25 @@
|
|||||||
|
import { FloatingGenerateBox } from '@/components/Generation/FloatingGenerateBox';
|
||||||
|
import { StoryContent } from './StoryContent';
|
||||||
|
import { StoryList } from './StoryList';
|
||||||
|
|
||||||
|
export function StoriesTab() {
|
||||||
|
return (
|
||||||
|
<div className="flex flex-col h-full min-h-0 overflow-hidden">
|
||||||
|
{/* Main content area */}
|
||||||
|
<div className="flex-1 min-h-0 flex gap-6 overflow-hidden relative">
|
||||||
|
{/* Left Column - Story List */}
|
||||||
|
<div className="flex flex-col min-h-0 overflow-hidden w-full max-w-[360px] shrink-0">
|
||||||
|
<StoryList />
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Right Column - Story Content */}
|
||||||
|
<div className="flex flex-col min-h-0 overflow-hidden flex-1">
|
||||||
|
<StoryContent />
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Floating Generate Box - position is managed via storyStore.trackEditorHeight */}
|
||||||
|
<FloatingGenerateBox showVoiceSelector />
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -0,0 +1,166 @@
|
|||||||
|
import { useSortable } from '@dnd-kit/sortable';
|
||||||
|
import { CSS } from '@dnd-kit/utilities';
|
||||||
|
import { GripVertical, Mic, MoreHorizontal, Play, Trash2 } from 'lucide-react';
|
||||||
|
import { useState } from 'react';
|
||||||
|
import { Button } from '@/components/ui/button';
|
||||||
|
import {
|
||||||
|
DropdownMenu,
|
||||||
|
DropdownMenuContent,
|
||||||
|
DropdownMenuItem,
|
||||||
|
DropdownMenuTrigger,
|
||||||
|
} from '@/components/ui/dropdown-menu';
|
||||||
|
import { Textarea } from '@/components/ui/textarea';
|
||||||
|
import type { StoryItemDetail } from '@/lib/api/types';
|
||||||
|
import { cn } from '@/lib/utils/cn';
|
||||||
|
import { useStoryStore } from '@/stores/storyStore';
|
||||||
|
import { useServerStore } from '@/stores/serverStore';
|
||||||
|
|
||||||
|
interface StoryChatItemProps {
|
||||||
|
item: StoryItemDetail;
|
||||||
|
storyId: string;
|
||||||
|
index: number;
|
||||||
|
onRemove: () => void;
|
||||||
|
currentTimeMs: number;
|
||||||
|
isPlaying: boolean;
|
||||||
|
dragHandleProps?: React.HTMLAttributes<HTMLButtonElement>;
|
||||||
|
isDragging?: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function StoryChatItem({
|
||||||
|
item,
|
||||||
|
onRemove,
|
||||||
|
currentTimeMs,
|
||||||
|
isPlaying,
|
||||||
|
dragHandleProps,
|
||||||
|
isDragging,
|
||||||
|
}: StoryChatItemProps) {
|
||||||
|
const seek = useStoryStore((state) => state.seek);
|
||||||
|
const serverUrl = useServerStore((state) => state.serverUrl);
|
||||||
|
const [avatarError, setAvatarError] = useState(false);
|
||||||
|
|
||||||
|
const avatarUrl = `${serverUrl}/profiles/${item.profile_id}/avatar`;
|
||||||
|
|
||||||
|
// Check if this item is currently playing based on timecode
|
||||||
|
const itemStartMs = item.start_time_ms;
|
||||||
|
const itemEndMs = item.start_time_ms + item.duration * 1000;
|
||||||
|
const isCurrentlyPlaying = isPlaying && currentTimeMs >= itemStartMs && currentTimeMs < itemEndMs;
|
||||||
|
|
||||||
|
const handlePlay = () => {
|
||||||
|
// Seek to the start of this item
|
||||||
|
seek(itemStartMs);
|
||||||
|
};
|
||||||
|
|
||||||
|
const formatTime = (ms: number): string => {
|
||||||
|
const totalSeconds = Math.floor(ms / 1000);
|
||||||
|
const minutes = Math.floor(totalSeconds / 60);
|
||||||
|
const seconds = totalSeconds % 60;
|
||||||
|
const milliseconds = Math.floor((ms % 1000) / 100);
|
||||||
|
return `${minutes}:${seconds.toString().padStart(2, '0')}.${milliseconds}`;
|
||||||
|
};
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div
|
||||||
|
className={cn(
|
||||||
|
'flex items-start gap-3 p-4 rounded-lg border transition-colors',
|
||||||
|
isCurrentlyPlaying && 'bg-muted/70 border-primary',
|
||||||
|
!isCurrentlyPlaying && 'hover:bg-muted/50',
|
||||||
|
isDragging && 'opacity-50 shadow-lg',
|
||||||
|
)}
|
||||||
|
>
|
||||||
|
{/* Drag Handle */}
|
||||||
|
{dragHandleProps && (
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="shrink-0 cursor-grab active:cursor-grabbing touch-none text-muted-foreground hover:text-foreground transition-colors"
|
||||||
|
{...dragHandleProps}
|
||||||
|
>
|
||||||
|
<GripVertical className="h-5 w-5" />
|
||||||
|
</button>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{/* Voice Avatar */}
|
||||||
|
<div className="shrink-0">
|
||||||
|
<div className="h-10 w-10 rounded-full bg-muted flex items-center justify-center overflow-hidden">
|
||||||
|
{!avatarError ? (
|
||||||
|
<img
|
||||||
|
src={avatarUrl}
|
||||||
|
alt={`${item.profile_name} avatar`}
|
||||||
|
className={cn(
|
||||||
|
'h-full w-full object-cover transition-all duration-200',
|
||||||
|
!isCurrentlyPlaying && 'grayscale'
|
||||||
|
)}
|
||||||
|
onError={() => setAvatarError(true)}
|
||||||
|
/>
|
||||||
|
) : (
|
||||||
|
<Mic className="h-5 w-5 text-muted-foreground" />
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Content */}
|
||||||
|
<div className="flex-1 min-w-0">
|
||||||
|
<div className="flex items-center gap-2 mb-2">
|
||||||
|
<span className="font-medium text-sm">{item.profile_name}</span>
|
||||||
|
<span className="text-xs text-muted-foreground">{item.language}</span>
|
||||||
|
<span className="text-xs text-muted-foreground tabular-nums ml-auto">
|
||||||
|
{formatTime(itemStartMs)}
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
<Textarea
|
||||||
|
value={item.text}
|
||||||
|
className="flex-1 resize-none text-sm text-muted-foreground select-text bg-card cursor-text"
|
||||||
|
readOnly
|
||||||
|
onDoubleClick={handlePlay}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Actions */}
|
||||||
|
<div className="shrink-0">
|
||||||
|
<DropdownMenu>
|
||||||
|
<DropdownMenuTrigger asChild>
|
||||||
|
<Button variant="ghost" size="icon" className="h-8 w-8" aria-label="Actions">
|
||||||
|
<MoreHorizontal className="h-4 w-4" />
|
||||||
|
</Button>
|
||||||
|
</DropdownMenuTrigger>
|
||||||
|
<DropdownMenuContent align="end">
|
||||||
|
<DropdownMenuItem onClick={handlePlay}>
|
||||||
|
<Play className="mr-2 h-4 w-4" />
|
||||||
|
Play from here
|
||||||
|
</DropdownMenuItem>
|
||||||
|
<DropdownMenuItem onClick={onRemove} className="text-destructive focus:text-destructive">
|
||||||
|
<Trash2 className="mr-2 h-4 w-4" />
|
||||||
|
Remove from Story
|
||||||
|
</DropdownMenuItem>
|
||||||
|
</DropdownMenuContent>
|
||||||
|
</DropdownMenu>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Sortable wrapper component
|
||||||
|
export function SortableStoryChatItem(props: Omit<StoryChatItemProps, 'dragHandleProps' | 'isDragging'>) {
|
||||||
|
const {
|
||||||
|
attributes,
|
||||||
|
listeners,
|
||||||
|
setNodeRef,
|
||||||
|
transform,
|
||||||
|
transition,
|
||||||
|
isDragging,
|
||||||
|
} = useSortable({ id: props.item.generation_id });
|
||||||
|
|
||||||
|
const style = {
|
||||||
|
transform: CSS.Transform.toString(transform),
|
||||||
|
transition,
|
||||||
|
};
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div ref={setNodeRef} style={style} {...attributes}>
|
||||||
|
<StoryChatItem
|
||||||
|
{...props}
|
||||||
|
dragHandleProps={listeners}
|
||||||
|
isDragging={isDragging}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -0,0 +1,376 @@
|
|||||||
|
import {
|
||||||
|
closestCenter,
|
||||||
|
DndContext,
|
||||||
|
type DragEndEvent,
|
||||||
|
KeyboardSensor,
|
||||||
|
PointerSensor,
|
||||||
|
useSensor,
|
||||||
|
useSensors,
|
||||||
|
} from '@dnd-kit/core';
|
||||||
|
import {
|
||||||
|
arrayMove,
|
||||||
|
SortableContext,
|
||||||
|
sortableKeyboardCoordinates,
|
||||||
|
verticalListSortingStrategy,
|
||||||
|
} from '@dnd-kit/sortable';
|
||||||
|
import { Download, Plus } from 'lucide-react';
|
||||||
|
import { useEffect, useMemo, useRef, useState } from 'react';
|
||||||
|
import { Button } from '@/components/ui/button';
|
||||||
|
import { Input } from '@/components/ui/input';
|
||||||
|
import { Popover, PopoverContent, PopoverTrigger } from '@/components/ui/popover';
|
||||||
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
|
import { useHistory } from '@/lib/hooks/useHistory';
|
||||||
|
import {
|
||||||
|
useAddStoryItem,
|
||||||
|
useExportStoryAudio,
|
||||||
|
useRemoveStoryItem,
|
||||||
|
useReorderStoryItems,
|
||||||
|
useStory,
|
||||||
|
} from '@/lib/hooks/useStories';
|
||||||
|
import { useStoryPlayback } from '@/lib/hooks/useStoryPlayback';
|
||||||
|
import { useStoryStore } from '@/stores/storyStore';
|
||||||
|
import { SortableStoryChatItem } from './StoryChatItem';
|
||||||
|
|
||||||
|
export function StoryContent() {
|
||||||
|
const selectedStoryId = useStoryStore((state) => state.selectedStoryId);
|
||||||
|
const { data: story, isLoading } = useStory(selectedStoryId);
|
||||||
|
const removeItem = useRemoveStoryItem();
|
||||||
|
const reorderItems = useReorderStoryItems();
|
||||||
|
const exportAudio = useExportStoryAudio();
|
||||||
|
const addStoryItem = useAddStoryItem();
|
||||||
|
const { toast } = useToast();
|
||||||
|
const scrollRef = useRef<HTMLDivElement>(null);
|
||||||
|
|
||||||
|
// Add generation popover state
|
||||||
|
const [searchQuery, setSearchQuery] = useState('');
|
||||||
|
const [isAddOpen, setIsAddOpen] = useState(false);
|
||||||
|
const { data: historyData } = useHistory();
|
||||||
|
|
||||||
|
// Filter generations not in story and matching search
|
||||||
|
const availableGenerations = useMemo(() => {
|
||||||
|
if (!historyData?.items || !story) return [];
|
||||||
|
const storyGenerationIds = new Set(story.items.map((i) => i.generation_id));
|
||||||
|
const query = searchQuery.toLowerCase();
|
||||||
|
return historyData.items.filter(
|
||||||
|
(gen) =>
|
||||||
|
!storyGenerationIds.has(gen.id) &&
|
||||||
|
(gen.text.toLowerCase().includes(query) ||
|
||||||
|
gen.profile_name.toLowerCase().includes(query)),
|
||||||
|
);
|
||||||
|
}, [historyData, story, searchQuery]);
|
||||||
|
|
||||||
|
// Get track editor height from store for dynamic padding
|
||||||
|
const trackEditorHeight = useStoryStore((state) => state.trackEditorHeight);
|
||||||
|
|
||||||
|
// Track editor is shown when story has items
|
||||||
|
const hasBottomBar = story && story.items.length > 0;
|
||||||
|
|
||||||
|
// Calculate dynamic bottom padding: track editor + gap
|
||||||
|
const bottomPadding = hasBottomBar ? trackEditorHeight + 24 : 0;
|
||||||
|
|
||||||
|
// Drag and drop sensors
|
||||||
|
const sensors = useSensors(
|
||||||
|
useSensor(PointerSensor, {
|
||||||
|
activationConstraint: {
|
||||||
|
distance: 8,
|
||||||
|
},
|
||||||
|
}),
|
||||||
|
useSensor(KeyboardSensor, {
|
||||||
|
coordinateGetter: sortableKeyboardCoordinates,
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
|
||||||
|
// Playback state (for auto-scroll and item highlighting)
|
||||||
|
const isPlaying = useStoryStore((state) => state.isPlaying);
|
||||||
|
const currentTimeMs = useStoryStore((state) => state.currentTimeMs);
|
||||||
|
const playbackStoryId = useStoryStore((state) => state.playbackStoryId);
|
||||||
|
|
||||||
|
// Refs for auto-scrolling to playing item
|
||||||
|
const itemRefsMap = useRef<Map<string, HTMLDivElement>>(new Map());
|
||||||
|
const lastScrolledItemRef = useRef<string | null>(null);
|
||||||
|
|
||||||
|
// Use playback hook
|
||||||
|
useStoryPlayback(story?.items);
|
||||||
|
|
||||||
|
// Sort items by start_time_ms
|
||||||
|
const sortedItems = useMemo(() => {
|
||||||
|
if (!story?.items) return [];
|
||||||
|
return [...story.items].sort((a, b) => a.start_time_ms - b.start_time_ms);
|
||||||
|
}, [story?.items]);
|
||||||
|
|
||||||
|
// Find the currently playing item based on timecode
|
||||||
|
const currentlyPlayingItemId = useMemo(() => {
|
||||||
|
if (!isPlaying || playbackStoryId !== story?.id || !sortedItems.length) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
const playingItem = sortedItems.find((item) => {
|
||||||
|
const itemStart = item.start_time_ms;
|
||||||
|
const itemEnd = item.start_time_ms + item.duration * 1000;
|
||||||
|
return currentTimeMs >= itemStart && currentTimeMs < itemEnd;
|
||||||
|
});
|
||||||
|
return playingItem?.generation_id ?? null;
|
||||||
|
}, [isPlaying, playbackStoryId, story?.id, sortedItems, currentTimeMs]);
|
||||||
|
|
||||||
|
// Auto-scroll to the currently playing item
|
||||||
|
useEffect(() => {
|
||||||
|
if (!currentlyPlayingItemId || currentlyPlayingItemId === lastScrolledItemRef.current) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const element = itemRefsMap.current.get(currentlyPlayingItemId);
|
||||||
|
if (element && scrollRef.current) {
|
||||||
|
element.scrollIntoView({ behavior: 'smooth', block: 'start' });
|
||||||
|
lastScrolledItemRef.current = currentlyPlayingItemId;
|
||||||
|
}
|
||||||
|
}, [currentlyPlayingItemId]);
|
||||||
|
|
||||||
|
// Reset last scrolled item when playback stops
|
||||||
|
useEffect(() => {
|
||||||
|
if (!isPlaying) {
|
||||||
|
lastScrolledItemRef.current = null;
|
||||||
|
}
|
||||||
|
}, [isPlaying]);
|
||||||
|
|
||||||
|
const handleRemoveItem = (itemId: string) => {
|
||||||
|
if (!story) return;
|
||||||
|
|
||||||
|
removeItem.mutate(
|
||||||
|
{
|
||||||
|
storyId: story.id,
|
||||||
|
itemId,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to remove item',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleDragEnd = (event: DragEndEvent) => {
|
||||||
|
const { active, over } = event;
|
||||||
|
|
||||||
|
if (!story || !over || active.id === over.id) return;
|
||||||
|
|
||||||
|
const oldIndex = sortedItems.findIndex((item) => item.generation_id === active.id);
|
||||||
|
const newIndex = sortedItems.findIndex((item) => item.generation_id === over.id);
|
||||||
|
|
||||||
|
if (oldIndex === -1 || newIndex === -1) return;
|
||||||
|
|
||||||
|
// Calculate the new order
|
||||||
|
const newOrder = arrayMove(sortedItems, oldIndex, newIndex);
|
||||||
|
const generationIds = newOrder.map((item) => item.generation_id);
|
||||||
|
|
||||||
|
// Send reorder request to backend
|
||||||
|
reorderItems.mutate(
|
||||||
|
{
|
||||||
|
storyId: story.id,
|
||||||
|
data: { generation_ids: generationIds },
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to reorder items',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleExportAudio = () => {
|
||||||
|
if (!story) return;
|
||||||
|
|
||||||
|
exportAudio.mutate(
|
||||||
|
{
|
||||||
|
storyId: story.id,
|
||||||
|
storyName: story.name,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to export audio',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleAddGeneration = (generationId: string) => {
|
||||||
|
if (!story) return;
|
||||||
|
|
||||||
|
addStoryItem.mutate(
|
||||||
|
{
|
||||||
|
storyId: story.id,
|
||||||
|
data: { generation_id: generationId },
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onSuccess: () => {
|
||||||
|
setIsAddOpen(false);
|
||||||
|
setSearchQuery('');
|
||||||
|
},
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to add generation',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
};
|
||||||
|
|
||||||
|
if (!selectedStoryId) {
|
||||||
|
return (
|
||||||
|
<div className="flex items-center justify-center h-full text-muted-foreground">
|
||||||
|
<div className="text-center">
|
||||||
|
<p className="text-lg font-medium mb-2">Select a story</p>
|
||||||
|
<p className="text-sm">Choose a story from the list to view its content</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (isLoading) {
|
||||||
|
return (
|
||||||
|
<div className="flex items-center justify-center h-full">
|
||||||
|
<div className="text-muted-foreground">Loading story...</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!story) {
|
||||||
|
return (
|
||||||
|
<div className="flex items-center justify-center h-full text-muted-foreground">
|
||||||
|
<div className="text-center">
|
||||||
|
<p className="text-lg font-medium mb-2">Story not found</p>
|
||||||
|
<p className="text-sm">The selected story could not be loaded</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="flex flex-col h-full min-h-0">
|
||||||
|
{/* Header */}
|
||||||
|
<div className="flex items-center justify-between mb-4 px-1">
|
||||||
|
<div>
|
||||||
|
<h2 className="text-2xl font-bold">{story.name}</h2>
|
||||||
|
{story.description && (
|
||||||
|
<p className="text-sm text-muted-foreground mt-1">{story.description}</p>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
<div className="flex gap-2">
|
||||||
|
<Popover open={isAddOpen} onOpenChange={setIsAddOpen}>
|
||||||
|
<PopoverTrigger asChild>
|
||||||
|
<Button variant="outline" size="sm">
|
||||||
|
<Plus className="mr-2 h-4 w-4" />
|
||||||
|
Add
|
||||||
|
</Button>
|
||||||
|
</PopoverTrigger>
|
||||||
|
<PopoverContent className="w-80 p-0" align="end">
|
||||||
|
<div className="p-2 border-b">
|
||||||
|
<Input
|
||||||
|
placeholder="Search by name or transcript..."
|
||||||
|
value={searchQuery}
|
||||||
|
onChange={(e) => setSearchQuery(e.target.value)}
|
||||||
|
autoFocus
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
<div className="max-h-60 overflow-y-auto">
|
||||||
|
{availableGenerations.length === 0 ? (
|
||||||
|
<div className="p-4 text-center text-sm text-muted-foreground">
|
||||||
|
{searchQuery
|
||||||
|
? 'No matching generations found'
|
||||||
|
: 'No available generations'}
|
||||||
|
</div>
|
||||||
|
) : (
|
||||||
|
availableGenerations.map((gen) => (
|
||||||
|
<button
|
||||||
|
key={gen.id}
|
||||||
|
type="button"
|
||||||
|
className="w-full text-left px-3 py-2 hover:bg-muted transition-colors border-b last:border-b-0"
|
||||||
|
onClick={() => handleAddGeneration(gen.id)}
|
||||||
|
>
|
||||||
|
<div className="font-medium text-sm">{gen.profile_name}</div>
|
||||||
|
<div className="text-xs text-muted-foreground truncate">
|
||||||
|
{gen.text.length > 50 ? `${gen.text.substring(0, 50)}...` : gen.text}
|
||||||
|
</div>
|
||||||
|
</button>
|
||||||
|
))
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
</PopoverContent>
|
||||||
|
</Popover>
|
||||||
|
{story.items.length > 0 && (
|
||||||
|
<Button
|
||||||
|
variant="outline"
|
||||||
|
size="sm"
|
||||||
|
onClick={handleExportAudio}
|
||||||
|
disabled={exportAudio.isPending}
|
||||||
|
>
|
||||||
|
<Download className="mr-2 h-4 w-4" />
|
||||||
|
Export Audio
|
||||||
|
</Button>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Content */}
|
||||||
|
<div
|
||||||
|
ref={scrollRef}
|
||||||
|
className="flex-1 min-h-0 overflow-y-auto space-y-3"
|
||||||
|
style={{ paddingBottom: bottomPadding > 0 ? `${bottomPadding}px` : undefined }}
|
||||||
|
>
|
||||||
|
{sortedItems.length === 0 ? (
|
||||||
|
<div className="text-center py-12 px-5 border-2 border-dashed border-muted rounded-md text-muted-foreground">
|
||||||
|
<p className="text-sm">No items in this story</p>
|
||||||
|
<p className="text-xs mt-2">Generate speech using the box below to add items</p>
|
||||||
|
</div>
|
||||||
|
) : (
|
||||||
|
<DndContext
|
||||||
|
sensors={sensors}
|
||||||
|
collisionDetection={closestCenter}
|
||||||
|
onDragEnd={handleDragEnd}
|
||||||
|
>
|
||||||
|
<SortableContext
|
||||||
|
items={sortedItems.map((item) => item.generation_id)}
|
||||||
|
strategy={verticalListSortingStrategy}
|
||||||
|
>
|
||||||
|
<div className="space-y-3">
|
||||||
|
{sortedItems.map((item, index) => (
|
||||||
|
<div
|
||||||
|
key={item.id}
|
||||||
|
ref={(el) => {
|
||||||
|
if (el) {
|
||||||
|
itemRefsMap.current.set(item.generation_id, el);
|
||||||
|
} else {
|
||||||
|
itemRefsMap.current.delete(item.generation_id);
|
||||||
|
}
|
||||||
|
}}
|
||||||
|
>
|
||||||
|
<SortableStoryChatItem
|
||||||
|
item={item}
|
||||||
|
storyId={story.id}
|
||||||
|
index={index}
|
||||||
|
onRemove={() => handleRemoveItem(item.id)}
|
||||||
|
currentTimeMs={currentTimeMs}
|
||||||
|
isPlaying={isPlaying && playbackStoryId === story.id}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
</SortableContext>
|
||||||
|
</DndContext>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -0,0 +1,369 @@
|
|||||||
|
import { Plus, BookOpen, MoreHorizontal, Pencil, Trash2 } from 'lucide-react';
|
||||||
|
import { useState } from 'react';
|
||||||
|
import {
|
||||||
|
AlertDialog,
|
||||||
|
AlertDialogAction,
|
||||||
|
AlertDialogCancel,
|
||||||
|
AlertDialogContent,
|
||||||
|
AlertDialogDescription,
|
||||||
|
AlertDialogFooter,
|
||||||
|
AlertDialogHeader,
|
||||||
|
AlertDialogTitle,
|
||||||
|
} from '@/components/ui/alert-dialog';
|
||||||
|
import { Button } from '@/components/ui/button';
|
||||||
|
import {
|
||||||
|
Dialog,
|
||||||
|
DialogContent,
|
||||||
|
DialogDescription,
|
||||||
|
DialogFooter,
|
||||||
|
DialogHeader,
|
||||||
|
DialogTitle,
|
||||||
|
} from '@/components/ui/dialog';
|
||||||
|
import {
|
||||||
|
DropdownMenu,
|
||||||
|
DropdownMenuContent,
|
||||||
|
DropdownMenuItem,
|
||||||
|
DropdownMenuTrigger,
|
||||||
|
} from '@/components/ui/dropdown-menu';
|
||||||
|
import { Input } from '@/components/ui/input';
|
||||||
|
import { Label } from '@/components/ui/label';
|
||||||
|
import { Textarea } from '@/components/ui/textarea';
|
||||||
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
|
import { useStories, useCreateStory, useUpdateStory, useDeleteStory } from '@/lib/hooks/useStories';
|
||||||
|
import { cn } from '@/lib/utils/cn';
|
||||||
|
import { formatDate } from '@/lib/utils/format';
|
||||||
|
import { useStoryStore } from '@/stores/storyStore';
|
||||||
|
|
||||||
|
export function StoryList() {
|
||||||
|
const { data: stories, isLoading } = useStories();
|
||||||
|
const selectedStoryId = useStoryStore((state) => state.selectedStoryId);
|
||||||
|
const setSelectedStoryId = useStoryStore((state) => state.setSelectedStoryId);
|
||||||
|
const createStory = useCreateStory();
|
||||||
|
const updateStory = useUpdateStory();
|
||||||
|
const deleteStory = useDeleteStory();
|
||||||
|
const [createDialogOpen, setCreateDialogOpen] = useState(false);
|
||||||
|
const [editDialogOpen, setEditDialogOpen] = useState(false);
|
||||||
|
const [deleteDialogOpen, setDeleteDialogOpen] = useState(false);
|
||||||
|
const [editingStory, setEditingStory] = useState<{
|
||||||
|
id: string;
|
||||||
|
name: string;
|
||||||
|
description?: string;
|
||||||
|
} | null>(null);
|
||||||
|
const [deletingStoryId, setDeletingStoryId] = useState<string | null>(null);
|
||||||
|
const [newStoryName, setNewStoryName] = useState('');
|
||||||
|
const [newStoryDescription, setNewStoryDescription] = useState('');
|
||||||
|
const { toast } = useToast();
|
||||||
|
|
||||||
|
const handleCreateStory = () => {
|
||||||
|
if (!newStoryName.trim()) {
|
||||||
|
toast({
|
||||||
|
title: 'Name required',
|
||||||
|
description: 'Please enter a story name',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
createStory.mutate(
|
||||||
|
{
|
||||||
|
name: newStoryName.trim(),
|
||||||
|
description: newStoryDescription.trim() || undefined,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onSuccess: (story) => {
|
||||||
|
setSelectedStoryId(story.id);
|
||||||
|
setCreateDialogOpen(false);
|
||||||
|
setNewStoryName('');
|
||||||
|
setNewStoryDescription('');
|
||||||
|
toast({
|
||||||
|
title: 'Story created',
|
||||||
|
description: `"${story.name}" has been created`,
|
||||||
|
});
|
||||||
|
},
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to create story',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleEditClick = (story: { id: string; name: string; description?: string }) => {
|
||||||
|
setEditingStory(story);
|
||||||
|
setNewStoryName(story.name);
|
||||||
|
setNewStoryDescription(story.description || '');
|
||||||
|
setEditDialogOpen(true);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleUpdateStory = () => {
|
||||||
|
if (!editingStory || !newStoryName.trim()) {
|
||||||
|
toast({
|
||||||
|
title: 'Name required',
|
||||||
|
description: 'Please enter a story name',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
updateStory.mutate(
|
||||||
|
{
|
||||||
|
storyId: editingStory.id,
|
||||||
|
data: {
|
||||||
|
name: newStoryName.trim(),
|
||||||
|
description: newStoryDescription.trim() || undefined,
|
||||||
|
},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onSuccess: () => {
|
||||||
|
setEditDialogOpen(false);
|
||||||
|
setEditingStory(null);
|
||||||
|
setNewStoryName('');
|
||||||
|
setNewStoryDescription('');
|
||||||
|
},
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to update story',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleDeleteClick = (storyId: string) => {
|
||||||
|
setDeletingStoryId(storyId);
|
||||||
|
setDeleteDialogOpen(true);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleDeleteConfirm = () => {
|
||||||
|
if (!deletingStoryId) return;
|
||||||
|
|
||||||
|
deleteStory.mutate(deletingStoryId, {
|
||||||
|
onSuccess: () => {
|
||||||
|
// Clear selection if deleting the currently selected story
|
||||||
|
if (selectedStoryId === deletingStoryId) {
|
||||||
|
setSelectedStoryId(null);
|
||||||
|
}
|
||||||
|
setDeleteDialogOpen(false);
|
||||||
|
setDeletingStoryId(null);
|
||||||
|
},
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to delete story',
|
||||||
|
description: error.message,
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
});
|
||||||
|
};
|
||||||
|
|
||||||
|
if (isLoading) {
|
||||||
|
return (
|
||||||
|
<div className="flex items-center justify-center h-full">
|
||||||
|
<div className="text-muted-foreground">Loading stories...</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const storyList = stories || [];
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="flex flex-col h-full min-h-0">
|
||||||
|
{/* Header */}
|
||||||
|
<div className="flex items-center justify-between mb-4 px-1">
|
||||||
|
<h2 className="text-2xl font-bold">Stories</h2>
|
||||||
|
<Button onClick={() => setCreateDialogOpen(true)} size="sm">
|
||||||
|
<Plus className="mr-2 h-4 w-4" />
|
||||||
|
New Story
|
||||||
|
</Button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Story List */}
|
||||||
|
<div className="flex-1 min-h-0 overflow-y-auto space-y-2">
|
||||||
|
{storyList.length === 0 ? (
|
||||||
|
<div className="text-center py-12 px-5 border-2 border-dashed border-muted rounded-2xl text-muted-foreground">
|
||||||
|
<BookOpen className="h-12 w-12 mx-auto mb-4 opacity-50" />
|
||||||
|
<p className="text-sm">No stories yet</p>
|
||||||
|
<p className="text-xs mt-2">Create your first story to get started</p>
|
||||||
|
</div>
|
||||||
|
) : (
|
||||||
|
storyList.map((story) => (
|
||||||
|
<div
|
||||||
|
key={story.id}
|
||||||
|
className={cn(
|
||||||
|
'h-24 p-4 border rounded-2xl transition-colors group flex items-center',
|
||||||
|
selectedStoryId === story.id && 'bg-muted border-primary',
|
||||||
|
)}
|
||||||
|
>
|
||||||
|
<div className="flex items-start justify-between gap-2 w-full min-w-0">
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="flex-1 min-w-0 text-left cursor-pointer overflow-hidden"
|
||||||
|
onClick={() => setSelectedStoryId(story.id)}
|
||||||
|
>
|
||||||
|
<h3 className="font-medium truncate">{story.name}</h3>
|
||||||
|
{story.description && (
|
||||||
|
<p className="text-sm text-muted-foreground mt-1 truncate">
|
||||||
|
{story.description}
|
||||||
|
</p>
|
||||||
|
)}
|
||||||
|
<div className="flex items-center gap-3 mt-2 text-xs text-muted-foreground">
|
||||||
|
<span>
|
||||||
|
{story.item_count} {story.item_count === 1 ? 'item' : 'items'}
|
||||||
|
</span>
|
||||||
|
<span>•</span>
|
||||||
|
<span>{formatDate(story.updated_at)}</span>
|
||||||
|
</div>
|
||||||
|
</button>
|
||||||
|
<DropdownMenu>
|
||||||
|
<DropdownMenuTrigger asChild>
|
||||||
|
<Button
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-8 w-8 opacity-0 group-hover:opacity-100 transition-opacity"
|
||||||
|
onClick={(e) => e.stopPropagation()}
|
||||||
|
>
|
||||||
|
<MoreHorizontal className="h-4 w-4" />
|
||||||
|
</Button>
|
||||||
|
</DropdownMenuTrigger>
|
||||||
|
<DropdownMenuContent align="end">
|
||||||
|
<DropdownMenuItem onClick={() => handleEditClick(story)}>
|
||||||
|
<Pencil className="mr-2 h-4 w-4" />
|
||||||
|
Edit
|
||||||
|
</DropdownMenuItem>
|
||||||
|
<DropdownMenuItem
|
||||||
|
onClick={() => handleDeleteClick(story.id)}
|
||||||
|
className="text-destructive focus:text-destructive"
|
||||||
|
>
|
||||||
|
<Trash2 className="mr-2 h-4 w-4" />
|
||||||
|
Delete
|
||||||
|
</DropdownMenuItem>
|
||||||
|
</DropdownMenuContent>
|
||||||
|
</DropdownMenu>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
))
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Create Story Dialog */}
|
||||||
|
<Dialog open={createDialogOpen} onOpenChange={setCreateDialogOpen}>
|
||||||
|
<DialogContent>
|
||||||
|
<DialogHeader>
|
||||||
|
<DialogTitle>Create New Story</DialogTitle>
|
||||||
|
<DialogDescription>
|
||||||
|
Create a new story to organize your voice generations into conversations.
|
||||||
|
</DialogDescription>
|
||||||
|
</DialogHeader>
|
||||||
|
<div className="space-y-4 py-4">
|
||||||
|
<div className="space-y-2">
|
||||||
|
<Label htmlFor="story-name">Name</Label>
|
||||||
|
<Input
|
||||||
|
id="story-name"
|
||||||
|
placeholder="My Story"
|
||||||
|
value={newStoryName}
|
||||||
|
onChange={(e) => setNewStoryName(e.target.value)}
|
||||||
|
onKeyDown={(e) => {
|
||||||
|
if (e.key === 'Enter') {
|
||||||
|
handleCreateStory();
|
||||||
|
}
|
||||||
|
}}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
<div className="space-y-2">
|
||||||
|
<Label htmlFor="story-description">Description (optional)</Label>
|
||||||
|
<Textarea
|
||||||
|
id="story-description"
|
||||||
|
placeholder="A conversation between..."
|
||||||
|
value={newStoryDescription}
|
||||||
|
onChange={(e) => setNewStoryDescription(e.target.value)}
|
||||||
|
rows={3}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<DialogFooter>
|
||||||
|
<Button variant="outline" onClick={() => setCreateDialogOpen(false)}>
|
||||||
|
Cancel
|
||||||
|
</Button>
|
||||||
|
<Button onClick={handleCreateStory} disabled={createStory.isPending}>
|
||||||
|
{createStory.isPending ? 'Creating...' : 'Create'}
|
||||||
|
</Button>
|
||||||
|
</DialogFooter>
|
||||||
|
</DialogContent>
|
||||||
|
</Dialog>
|
||||||
|
|
||||||
|
{/* Edit Story Dialog */}
|
||||||
|
<Dialog open={editDialogOpen} onOpenChange={setEditDialogOpen}>
|
||||||
|
<DialogContent>
|
||||||
|
<DialogHeader>
|
||||||
|
<DialogTitle>Edit Story</DialogTitle>
|
||||||
|
<DialogDescription>Update the story name and description.</DialogDescription>
|
||||||
|
</DialogHeader>
|
||||||
|
<div className="space-y-4 py-4">
|
||||||
|
<div className="space-y-2">
|
||||||
|
<Label htmlFor="edit-story-name">Name</Label>
|
||||||
|
<Input
|
||||||
|
id="edit-story-name"
|
||||||
|
placeholder="My Story"
|
||||||
|
value={newStoryName}
|
||||||
|
onChange={(e) => setNewStoryName(e.target.value)}
|
||||||
|
onKeyDown={(e) => {
|
||||||
|
if (e.key === 'Enter') {
|
||||||
|
handleUpdateStory();
|
||||||
|
}
|
||||||
|
}}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
<div className="space-y-2">
|
||||||
|
<Label htmlFor="edit-story-description">Description (optional)</Label>
|
||||||
|
<Textarea
|
||||||
|
id="edit-story-description"
|
||||||
|
placeholder="A conversation between..."
|
||||||
|
value={newStoryDescription}
|
||||||
|
onChange={(e) => setNewStoryDescription(e.target.value)}
|
||||||
|
rows={3}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<DialogFooter>
|
||||||
|
<Button variant="outline" onClick={() => setEditDialogOpen(false)}>
|
||||||
|
Cancel
|
||||||
|
</Button>
|
||||||
|
<Button onClick={handleUpdateStory} disabled={updateStory.isPending}>
|
||||||
|
{updateStory.isPending ? 'Saving...' : 'Save'}
|
||||||
|
</Button>
|
||||||
|
</DialogFooter>
|
||||||
|
</DialogContent>
|
||||||
|
</Dialog>
|
||||||
|
|
||||||
|
{/* Delete Story Confirmation Dialog */}
|
||||||
|
<AlertDialog open={deleteDialogOpen} onOpenChange={setDeleteDialogOpen}>
|
||||||
|
<AlertDialogContent>
|
||||||
|
<AlertDialogHeader>
|
||||||
|
<AlertDialogTitle>Are you sure?</AlertDialogTitle>
|
||||||
|
<AlertDialogDescription>
|
||||||
|
This will permanently delete the story and all its items. This action cannot be
|
||||||
|
undone.
|
||||||
|
</AlertDialogDescription>
|
||||||
|
</AlertDialogHeader>
|
||||||
|
<AlertDialogFooter>
|
||||||
|
<AlertDialogCancel>Cancel</AlertDialogCancel>
|
||||||
|
<AlertDialogAction asChild>
|
||||||
|
<Button
|
||||||
|
onClick={handleDeleteConfirm}
|
||||||
|
disabled={deleteStory.isPending}
|
||||||
|
className="bg-destructive text-destructive-foreground hover:bg-destructive/90"
|
||||||
|
>
|
||||||
|
{deleteStory.isPending ? 'Deleting...' : 'Delete'}
|
||||||
|
</Button>
|
||||||
|
</AlertDialogAction>
|
||||||
|
</AlertDialogFooter>
|
||||||
|
</AlertDialogContent>
|
||||||
|
</AlertDialog>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -0,0 +1,988 @@
|
|||||||
|
import {
|
||||||
|
Copy,
|
||||||
|
GripHorizontal,
|
||||||
|
Minus,
|
||||||
|
Pause,
|
||||||
|
Play,
|
||||||
|
Plus,
|
||||||
|
Scissors,
|
||||||
|
Square,
|
||||||
|
Trash2,
|
||||||
|
} from 'lucide-react';
|
||||||
|
import { useCallback, useEffect, useMemo, useRef, useState } from 'react';
|
||||||
|
import WaveSurfer from 'wavesurfer.js';
|
||||||
|
import { Button } from '@/components/ui/button';
|
||||||
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
|
import { apiClient } from '@/lib/api/client';
|
||||||
|
import type { StoryItemDetail } from '@/lib/api/types';
|
||||||
|
import {
|
||||||
|
useDuplicateStoryItem,
|
||||||
|
useMoveStoryItem,
|
||||||
|
useRemoveStoryItem,
|
||||||
|
useSplitStoryItem,
|
||||||
|
useTrimStoryItem,
|
||||||
|
} from '@/lib/hooks/useStories';
|
||||||
|
import { cn } from '@/lib/utils/cn';
|
||||||
|
import { useStoryStore } from '@/stores/storyStore';
|
||||||
|
|
||||||
|
// Clip waveform component with trim support
|
||||||
|
function ClipWaveform({
|
||||||
|
generationId,
|
||||||
|
width,
|
||||||
|
trimStartMs,
|
||||||
|
trimEndMs,
|
||||||
|
duration,
|
||||||
|
}: {
|
||||||
|
generationId: string;
|
||||||
|
width: number;
|
||||||
|
trimStartMs: number;
|
||||||
|
trimEndMs: number;
|
||||||
|
duration: number;
|
||||||
|
}) {
|
||||||
|
const waveformRef = useRef<HTMLDivElement>(null);
|
||||||
|
const wavesurferRef = useRef<WaveSurfer | null>(null);
|
||||||
|
|
||||||
|
// Calculate the full waveform width based on the original duration
|
||||||
|
// The visible portion (width) represents the effective duration after trimming
|
||||||
|
const effectiveDurationMs = duration * 1000 - trimStartMs - trimEndMs;
|
||||||
|
const fullWaveformWidth =
|
||||||
|
effectiveDurationMs > 0 ? (width / effectiveDurationMs) * (duration * 1000) : width;
|
||||||
|
|
||||||
|
// Calculate how much to offset the waveform to hide the trimmed start
|
||||||
|
const offsetX =
|
||||||
|
effectiveDurationMs > 0 ? (trimStartMs / (duration * 1000)) * fullWaveformWidth : 0;
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
if (!waveformRef.current || fullWaveformWidth < 20) return;
|
||||||
|
|
||||||
|
// Get CSS colors
|
||||||
|
const root = document.documentElement;
|
||||||
|
const getCSSVar = (varName: string) => {
|
||||||
|
const value = getComputedStyle(root).getPropertyValue(varName).trim();
|
||||||
|
return value ? `hsl(${value})` : '';
|
||||||
|
};
|
||||||
|
|
||||||
|
const waveColor = getCSSVar('--accent-foreground');
|
||||||
|
|
||||||
|
const wavesurfer = WaveSurfer.create({
|
||||||
|
container: waveformRef.current,
|
||||||
|
waveColor,
|
||||||
|
progressColor: waveColor,
|
||||||
|
cursorWidth: 0,
|
||||||
|
barWidth: 1,
|
||||||
|
barRadius: 1,
|
||||||
|
barGap: 1,
|
||||||
|
height: 28,
|
||||||
|
normalize: true,
|
||||||
|
interact: false,
|
||||||
|
});
|
||||||
|
|
||||||
|
wavesurferRef.current = wavesurfer;
|
||||||
|
|
||||||
|
const audioUrl = apiClient.getAudioUrl(generationId);
|
||||||
|
wavesurfer.load(audioUrl).catch(() => {
|
||||||
|
// Ignore load errors
|
||||||
|
});
|
||||||
|
|
||||||
|
return () => {
|
||||||
|
wavesurfer.destroy();
|
||||||
|
wavesurferRef.current = null;
|
||||||
|
};
|
||||||
|
}, [generationId, fullWaveformWidth]);
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="w-full h-full opacity-60 overflow-hidden">
|
||||||
|
{/* Inner container that holds the full waveform, offset to show only visible portion */}
|
||||||
|
<div
|
||||||
|
ref={waveformRef}
|
||||||
|
style={{
|
||||||
|
width: `${fullWaveformWidth}px`,
|
||||||
|
transform: `translateX(-${offsetX}px)`,
|
||||||
|
}}
|
||||||
|
className="h-full"
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
interface StoryTrackEditorProps {
|
||||||
|
storyId: string;
|
||||||
|
items: StoryItemDetail[];
|
||||||
|
}
|
||||||
|
|
||||||
|
const TRACK_HEIGHT = 48;
|
||||||
|
const TIME_RULER_HEIGHT = 24; // h-6 = 1.5rem = 24px
|
||||||
|
const MIN_PIXELS_PER_SECOND = 10;
|
||||||
|
const MAX_PIXELS_PER_SECOND = 200;
|
||||||
|
const DEFAULT_PIXELS_PER_SECOND = 50;
|
||||||
|
const DEFAULT_TRACKS = [1, 0, -1]; // Default 3 tracks
|
||||||
|
const MIN_EDITOR_HEIGHT = 120;
|
||||||
|
const MAX_EDITOR_HEIGHT = 500;
|
||||||
|
|
||||||
|
export function StoryTrackEditor({ storyId, items }: StoryTrackEditorProps) {
|
||||||
|
const [pixelsPerSecond, setPixelsPerSecond] = useState(DEFAULT_PIXELS_PER_SECOND);
|
||||||
|
const [draggingItem, setDraggingItem] = useState<string | null>(null);
|
||||||
|
const [dragOffset, setDragOffset] = useState({ x: 0, y: 0 });
|
||||||
|
const [dragPosition, setDragPosition] = useState({ x: 0, y: 0 });
|
||||||
|
const [isResizing, setIsResizing] = useState(false);
|
||||||
|
const [containerWidth, setContainerWidth] = useState(0);
|
||||||
|
const containerRef = useRef<HTMLDivElement>(null);
|
||||||
|
const tracksRef = useRef<HTMLDivElement>(null);
|
||||||
|
const resizeStartY = useRef(0);
|
||||||
|
const resizeStartHeight = useRef(0);
|
||||||
|
const moveItem = useMoveStoryItem();
|
||||||
|
const trimItem = useTrimStoryItem();
|
||||||
|
const splitItem = useSplitStoryItem();
|
||||||
|
const duplicateItem = useDuplicateStoryItem();
|
||||||
|
const removeItem = useRemoveStoryItem();
|
||||||
|
const { toast } = useToast();
|
||||||
|
|
||||||
|
// Selection state
|
||||||
|
const selectedClipId = useStoryStore((state) => state.selectedClipId);
|
||||||
|
const setSelectedClipId = useStoryStore((state) => state.setSelectedClipId);
|
||||||
|
|
||||||
|
// Trim state
|
||||||
|
const [trimmingItem, setTrimmingItem] = useState<string | null>(null);
|
||||||
|
const [trimSide, setTrimSide] = useState<'start' | 'end' | null>(null);
|
||||||
|
const [trimStartX, setTrimStartX] = useState(0);
|
||||||
|
const [tempTrimValues, setTempTrimValues] = useState<{
|
||||||
|
trim_start_ms: number;
|
||||||
|
trim_end_ms: number;
|
||||||
|
} | null>(null);
|
||||||
|
|
||||||
|
// Track editor height from store (shared with FloatingGenerateBox)
|
||||||
|
const editorHeight = useStoryStore((state) => state.trackEditorHeight);
|
||||||
|
const setEditorHeight = useStoryStore((state) => state.setTrackEditorHeight);
|
||||||
|
|
||||||
|
// Playback state
|
||||||
|
const isPlaying = useStoryStore((state) => state.isPlaying);
|
||||||
|
const currentTimeMs = useStoryStore((state) => state.currentTimeMs);
|
||||||
|
const playbackStoryId = useStoryStore((state) => state.playbackStoryId);
|
||||||
|
const play = useStoryStore((state) => state.play);
|
||||||
|
const pause = useStoryStore((state) => state.pause);
|
||||||
|
const stop = useStoryStore((state) => state.stop);
|
||||||
|
const seek = useStoryStore((state) => state.seek);
|
||||||
|
const setActiveStory = useStoryStore((state) => state.setActiveStory);
|
||||||
|
|
||||||
|
const isActiveStory = playbackStoryId === storyId;
|
||||||
|
const isCurrentlyPlaying = isPlaying && isActiveStory;
|
||||||
|
|
||||||
|
// Auto-activate this story when the editor is shown so playhead is visible
|
||||||
|
useEffect(() => {
|
||||||
|
if (items.length > 0 && !isActiveStory) {
|
||||||
|
const totalDuration = Math.max(
|
||||||
|
...items.map((item) => {
|
||||||
|
const trimStart = item.trim_start_ms || 0;
|
||||||
|
const trimEnd = item.trim_end_ms || 0;
|
||||||
|
const effectiveDuration = item.duration * 1000 - trimStart - trimEnd;
|
||||||
|
return item.start_time_ms + effectiveDuration;
|
||||||
|
}),
|
||||||
|
0,
|
||||||
|
);
|
||||||
|
setActiveStory(storyId, items, totalDuration);
|
||||||
|
}
|
||||||
|
}, [storyId, items, isActiveStory, setActiveStory]);
|
||||||
|
|
||||||
|
// Sort items by start time for play
|
||||||
|
const sortedItems = useMemo(() => {
|
||||||
|
return [...items].sort((a, b) => a.start_time_ms - b.start_time_ms);
|
||||||
|
}, [items]);
|
||||||
|
|
||||||
|
const handlePlayPause = () => {
|
||||||
|
if (isCurrentlyPlaying) {
|
||||||
|
pause();
|
||||||
|
} else {
|
||||||
|
play(storyId, sortedItems);
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleStop = () => {
|
||||||
|
stop();
|
||||||
|
};
|
||||||
|
|
||||||
|
// Calculate unique tracks from items, always showing at least 3 default tracks
|
||||||
|
const tracks = useMemo(() => {
|
||||||
|
const trackSet = new Set([...DEFAULT_TRACKS, ...items.map((item) => item.track)]);
|
||||||
|
return Array.from(trackSet).sort((a, b) => b - a); // Higher tracks on top
|
||||||
|
}, [items]);
|
||||||
|
|
||||||
|
// Track container width for full-width minimum
|
||||||
|
useEffect(() => {
|
||||||
|
const container = tracksRef.current;
|
||||||
|
if (!container) return;
|
||||||
|
|
||||||
|
const observer = new ResizeObserver((entries) => {
|
||||||
|
for (const entry of entries) {
|
||||||
|
setContainerWidth(entry.contentRect.width);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
observer.observe(container);
|
||||||
|
// Set initial width
|
||||||
|
setContainerWidth(container.clientWidth);
|
||||||
|
|
||||||
|
return () => observer.disconnect();
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
// Calculate effective duration (accounting for trims)
|
||||||
|
const getEffectiveDuration = (item: StoryItemDetail) => {
|
||||||
|
return item.duration * 1000 - (item.trim_start_ms || 0) - (item.trim_end_ms || 0);
|
||||||
|
};
|
||||||
|
|
||||||
|
// Calculate total duration (using effective durations)
|
||||||
|
const totalDurationMs = useMemo(() => {
|
||||||
|
if (items.length === 0) return 10000; // Default 10 seconds
|
||||||
|
return Math.max(...items.map((item) => item.start_time_ms + getEffectiveDuration(item)), 10000);
|
||||||
|
}, [items, getEffectiveDuration]);
|
||||||
|
|
||||||
|
// Calculate timeline width - at least full container width
|
||||||
|
const contentWidth = (totalDurationMs / 1000) * pixelsPerSecond + 200; // Content width with padding
|
||||||
|
const timelineWidth = Math.max(contentWidth, containerWidth);
|
||||||
|
|
||||||
|
// Generate time markers
|
||||||
|
const timeMarkers = useMemo(() => {
|
||||||
|
const markers: number[] = [];
|
||||||
|
// Determine interval based on zoom level
|
||||||
|
let intervalMs = 5000; // 5 seconds
|
||||||
|
if (pixelsPerSecond > 100) intervalMs = 1000;
|
||||||
|
else if (pixelsPerSecond > 50) intervalMs = 2000;
|
||||||
|
else if (pixelsPerSecond < 20) intervalMs = 10000;
|
||||||
|
|
||||||
|
for (let ms = 0; ms <= totalDurationMs + intervalMs; ms += intervalMs) {
|
||||||
|
markers.push(ms);
|
||||||
|
}
|
||||||
|
return markers;
|
||||||
|
}, [totalDurationMs, pixelsPerSecond]);
|
||||||
|
|
||||||
|
const formatTime = (ms: number): string => {
|
||||||
|
const totalSeconds = Math.floor(ms / 1000);
|
||||||
|
const minutes = Math.floor(totalSeconds / 60);
|
||||||
|
const seconds = totalSeconds % 60;
|
||||||
|
return `${minutes}:${seconds.toString().padStart(2, '0')}`;
|
||||||
|
};
|
||||||
|
|
||||||
|
const msToPixels = useCallback((ms: number) => (ms / 1000) * pixelsPerSecond, [pixelsPerSecond]);
|
||||||
|
|
||||||
|
const pixelsToMs = useCallback((px: number) => (px / pixelsPerSecond) * 1000, [pixelsPerSecond]);
|
||||||
|
|
||||||
|
const handleZoomIn = () => {
|
||||||
|
setPixelsPerSecond((prev) => Math.min(prev * 1.5, MAX_PIXELS_PER_SECOND));
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleZoomOut = () => {
|
||||||
|
setPixelsPerSecond((prev) => Math.max(prev / 1.5, MIN_PIXELS_PER_SECOND));
|
||||||
|
};
|
||||||
|
|
||||||
|
// Resize handlers
|
||||||
|
const handleResizeStart = useCallback(
|
||||||
|
(e: React.MouseEvent) => {
|
||||||
|
e.preventDefault();
|
||||||
|
setIsResizing(true);
|
||||||
|
resizeStartY.current = e.clientY;
|
||||||
|
resizeStartHeight.current = editorHeight;
|
||||||
|
},
|
||||||
|
[editorHeight],
|
||||||
|
);
|
||||||
|
|
||||||
|
const handleResizeMove = useCallback(
|
||||||
|
(e: MouseEvent) => {
|
||||||
|
if (!isResizing) return;
|
||||||
|
const deltaY = resizeStartY.current - e.clientY;
|
||||||
|
const newHeight = Math.min(
|
||||||
|
MAX_EDITOR_HEIGHT,
|
||||||
|
Math.max(MIN_EDITOR_HEIGHT, resizeStartHeight.current + deltaY),
|
||||||
|
);
|
||||||
|
setEditorHeight(newHeight);
|
||||||
|
},
|
||||||
|
[isResizing, setEditorHeight],
|
||||||
|
);
|
||||||
|
|
||||||
|
const handleResizeEnd = useCallback(() => {
|
||||||
|
setIsResizing(false);
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
// Add global mouse listeners for resizing
|
||||||
|
useEffect(() => {
|
||||||
|
if (isResizing) {
|
||||||
|
window.addEventListener('mousemove', handleResizeMove);
|
||||||
|
window.addEventListener('mouseup', handleResizeEnd);
|
||||||
|
return () => {
|
||||||
|
window.removeEventListener('mousemove', handleResizeMove);
|
||||||
|
window.removeEventListener('mouseup', handleResizeEnd);
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}, [isResizing, handleResizeMove, handleResizeEnd]);
|
||||||
|
|
||||||
|
const handleTimelineClick = (e: React.MouseEvent<HTMLDivElement>) => {
|
||||||
|
if (!tracksRef.current || draggingItem || trimmingItem) return;
|
||||||
|
const rect = tracksRef.current.getBoundingClientRect();
|
||||||
|
const x = e.clientX - rect.left + tracksRef.current.scrollLeft;
|
||||||
|
const timeMs = Math.max(0, pixelsToMs(x));
|
||||||
|
seek(timeMs);
|
||||||
|
// Deselect clip when clicking on timeline
|
||||||
|
setSelectedClipId(null);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleClipClick = (e: React.MouseEvent, item: StoryItemDetail) => {
|
||||||
|
e.stopPropagation();
|
||||||
|
if (draggingItem || trimmingItem) return;
|
||||||
|
setSelectedClipId(item.id);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleTrimStart = (e: React.MouseEvent, item: StoryItemDetail, side: 'start' | 'end') => {
|
||||||
|
e.stopPropagation();
|
||||||
|
if (!tracksRef.current) return;
|
||||||
|
setTrimmingItem(item.id);
|
||||||
|
setTrimSide(side);
|
||||||
|
setSelectedClipId(item.id);
|
||||||
|
setTrimStartX(e.clientX);
|
||||||
|
trimStartItemRef.current = {
|
||||||
|
item,
|
||||||
|
initialTrimStart: item.trim_start_ms || 0,
|
||||||
|
initialTrimEnd: item.trim_end_ms || 0,
|
||||||
|
};
|
||||||
|
};
|
||||||
|
|
||||||
|
const trimStartItemRef = useRef<{
|
||||||
|
item: StoryItemDetail;
|
||||||
|
initialTrimStart: number;
|
||||||
|
initialTrimEnd: number;
|
||||||
|
} | null>(null);
|
||||||
|
|
||||||
|
const handleTrimMove = useCallback(
|
||||||
|
(e: MouseEvent) => {
|
||||||
|
if (!trimmingItem || !trimSide || !trimStartItemRef.current) return;
|
||||||
|
|
||||||
|
const deltaX = e.clientX - trimStartX;
|
||||||
|
const deltaMs = pixelsToMs(deltaX); // Signed delta in milliseconds
|
||||||
|
|
||||||
|
const { item, initialTrimStart, initialTrimEnd } = trimStartItemRef.current;
|
||||||
|
const originalDurationMs = item.duration * 1000;
|
||||||
|
|
||||||
|
let newTrimStart = initialTrimStart;
|
||||||
|
let newTrimEnd = initialTrimEnd;
|
||||||
|
|
||||||
|
if (trimSide === 'start') {
|
||||||
|
// Moving right increases trim_start (trims more from start)
|
||||||
|
// Moving left decreases trim_start (restores from start)
|
||||||
|
newTrimStart = Math.round(
|
||||||
|
Math.max(
|
||||||
|
0,
|
||||||
|
Math.min(initialTrimStart + deltaMs, originalDurationMs - initialTrimEnd - 100),
|
||||||
|
),
|
||||||
|
);
|
||||||
|
} else {
|
||||||
|
// Moving right decreases trim_end (restores from end)
|
||||||
|
// Moving left increases trim_end (trims more from end)
|
||||||
|
newTrimEnd = Math.round(
|
||||||
|
Math.max(
|
||||||
|
0,
|
||||||
|
Math.min(initialTrimEnd - deltaMs, originalDurationMs - initialTrimStart - 100),
|
||||||
|
),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Validate that we don't exceed duration
|
||||||
|
if (newTrimStart + newTrimEnd >= originalDurationMs - 100) {
|
||||||
|
return; // Don't allow trimming to less than 100ms
|
||||||
|
}
|
||||||
|
|
||||||
|
// Update temporary trim values for visual feedback
|
||||||
|
setTempTrimValues({
|
||||||
|
trim_start_ms: newTrimStart,
|
||||||
|
trim_end_ms: newTrimEnd,
|
||||||
|
});
|
||||||
|
},
|
||||||
|
[trimmingItem, trimSide, trimStartX, pixelsToMs],
|
||||||
|
);
|
||||||
|
|
||||||
|
const handleTrimEnd = useCallback(() => {
|
||||||
|
if (!trimmingItem || !trimSide || !trimStartItemRef.current) {
|
||||||
|
setTrimmingItem(null);
|
||||||
|
setTrimSide(null);
|
||||||
|
setTempTrimValues(null);
|
||||||
|
trimStartItemRef.current = null;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const { initialTrimStart, initialTrimEnd } = trimStartItemRef.current;
|
||||||
|
|
||||||
|
// Use temporary trim values if available, otherwise use initial values
|
||||||
|
// Ensure values are integers for the backend
|
||||||
|
const finalTrimStart = Math.round(tempTrimValues?.trim_start_ms ?? initialTrimStart);
|
||||||
|
const finalTrimEnd = Math.round(tempTrimValues?.trim_end_ms ?? initialTrimEnd);
|
||||||
|
|
||||||
|
// Only update if values changed
|
||||||
|
if (finalTrimStart !== initialTrimStart || finalTrimEnd !== initialTrimEnd) {
|
||||||
|
trimItem.mutate(
|
||||||
|
{
|
||||||
|
storyId,
|
||||||
|
itemId: trimmingItem,
|
||||||
|
data: {
|
||||||
|
trim_start_ms: finalTrimStart,
|
||||||
|
trim_end_ms: finalTrimEnd,
|
||||||
|
},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to trim clip',
|
||||||
|
description: error instanceof Error ? error.message : String(error),
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
setTrimmingItem(null);
|
||||||
|
setTrimSide(null);
|
||||||
|
setTempTrimValues(null);
|
||||||
|
trimStartItemRef.current = null;
|
||||||
|
}, [trimmingItem, trimSide, tempTrimValues, storyId, trimItem, toast]);
|
||||||
|
|
||||||
|
const handleSplit = useCallback(() => {
|
||||||
|
if (!selectedClipId) return;
|
||||||
|
|
||||||
|
const item = items.find((i) => i.id === selectedClipId);
|
||||||
|
if (!item) return;
|
||||||
|
|
||||||
|
const splitTimeMs = currentTimeMs - item.start_time_ms;
|
||||||
|
const effectiveDuration = getEffectiveDuration(item);
|
||||||
|
|
||||||
|
if (splitTimeMs <= 0 || splitTimeMs >= effectiveDuration) {
|
||||||
|
toast({
|
||||||
|
title: 'Invalid split point',
|
||||||
|
description: 'Playhead must be within the selected clip',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
splitItem.mutate(
|
||||||
|
{
|
||||||
|
storyId,
|
||||||
|
itemId: selectedClipId,
|
||||||
|
data: { split_time_ms: splitTimeMs },
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onSuccess: () => {
|
||||||
|
setSelectedClipId(null);
|
||||||
|
},
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to split clip',
|
||||||
|
description: error instanceof Error ? error.message : String(error),
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}, [
|
||||||
|
selectedClipId,
|
||||||
|
items,
|
||||||
|
currentTimeMs,
|
||||||
|
getEffectiveDuration,
|
||||||
|
storyId,
|
||||||
|
splitItem,
|
||||||
|
toast,
|
||||||
|
setSelectedClipId,
|
||||||
|
]);
|
||||||
|
|
||||||
|
const handleDuplicate = useCallback(() => {
|
||||||
|
if (!selectedClipId) return;
|
||||||
|
|
||||||
|
duplicateItem.mutate(
|
||||||
|
{
|
||||||
|
storyId,
|
||||||
|
itemId: selectedClipId,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to duplicate clip',
|
||||||
|
description: error instanceof Error ? error.message : String(error),
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}, [selectedClipId, storyId, duplicateItem, toast]);
|
||||||
|
|
||||||
|
const handleDelete = useCallback(() => {
|
||||||
|
if (!selectedClipId) return;
|
||||||
|
|
||||||
|
removeItem.mutate(
|
||||||
|
{
|
||||||
|
storyId,
|
||||||
|
itemId: selectedClipId,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onSuccess: () => {
|
||||||
|
setSelectedClipId(null);
|
||||||
|
},
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to delete clip',
|
||||||
|
description: error instanceof Error ? error.message : String(error),
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}, [selectedClipId, storyId, removeItem, toast, setSelectedClipId]);
|
||||||
|
|
||||||
|
// Keyboard shortcuts
|
||||||
|
useEffect(() => {
|
||||||
|
const handleKeyDown = (e: KeyboardEvent) => {
|
||||||
|
// Only handle shortcuts when editor is focused or no input is focused
|
||||||
|
if (e.target instanceof HTMLInputElement || e.target instanceof HTMLTextAreaElement) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (e.key === ' ') {
|
||||||
|
e.preventDefault();
|
||||||
|
handlePlayPause();
|
||||||
|
} else if (e.key === 'Escape') {
|
||||||
|
setSelectedClipId(null);
|
||||||
|
} else if (e.key === 's' || e.key === 'S') {
|
||||||
|
if (selectedClipId) {
|
||||||
|
e.preventDefault();
|
||||||
|
handleSplit();
|
||||||
|
}
|
||||||
|
} else if (e.key === 'd' || e.key === 'D') {
|
||||||
|
if (selectedClipId && (e.metaKey || e.ctrlKey)) {
|
||||||
|
e.preventDefault();
|
||||||
|
handleDuplicate();
|
||||||
|
}
|
||||||
|
} else if (e.key === 'Delete' || e.key === 'Backspace') {
|
||||||
|
if (selectedClipId) {
|
||||||
|
e.preventDefault();
|
||||||
|
handleDelete();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
window.addEventListener('keydown', handleKeyDown);
|
||||||
|
return () => window.removeEventListener('keydown', handleKeyDown);
|
||||||
|
}, [
|
||||||
|
selectedClipId,
|
||||||
|
handleSplit,
|
||||||
|
handleDuplicate,
|
||||||
|
handleDelete,
|
||||||
|
setSelectedClipId,
|
||||||
|
handlePlayPause,
|
||||||
|
]);
|
||||||
|
|
||||||
|
// Add global mouse listeners for trimming
|
||||||
|
useEffect(() => {
|
||||||
|
if (trimmingItem) {
|
||||||
|
window.addEventListener('mousemove', handleTrimMove);
|
||||||
|
window.addEventListener('mouseup', handleTrimEnd);
|
||||||
|
return () => {
|
||||||
|
window.removeEventListener('mousemove', handleTrimMove);
|
||||||
|
window.removeEventListener('mouseup', handleTrimEnd);
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}, [trimmingItem, handleTrimMove, handleTrimEnd]);
|
||||||
|
|
||||||
|
const handleDragStart = (e: React.MouseEvent, item: StoryItemDetail) => {
|
||||||
|
e.stopPropagation();
|
||||||
|
if (!tracksRef.current) return;
|
||||||
|
|
||||||
|
const rect = e.currentTarget.getBoundingClientRect();
|
||||||
|
setDragOffset({
|
||||||
|
x: e.clientX - rect.left,
|
||||||
|
y: e.clientY - rect.top,
|
||||||
|
});
|
||||||
|
setDragPosition({
|
||||||
|
x: rect.left - tracksRef.current.getBoundingClientRect().left + tracksRef.current.scrollLeft,
|
||||||
|
// Subtract ruler height since clips are positioned relative to tracks area, not the scrollable container
|
||||||
|
y: rect.top - tracksRef.current.getBoundingClientRect().top - TIME_RULER_HEIGHT,
|
||||||
|
});
|
||||||
|
setDraggingItem(item.id);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleDragMove = useCallback(
|
||||||
|
(e: React.MouseEvent) => {
|
||||||
|
if (!draggingItem || !tracksRef.current) return;
|
||||||
|
|
||||||
|
const rect = tracksRef.current.getBoundingClientRect();
|
||||||
|
const x = e.clientX - rect.left + tracksRef.current.scrollLeft - dragOffset.x;
|
||||||
|
// Subtract ruler height since clips are positioned relative to tracks area
|
||||||
|
const y = e.clientY - rect.top - dragOffset.y - TIME_RULER_HEIGHT;
|
||||||
|
|
||||||
|
setDragPosition({ x: Math.max(0, x), y });
|
||||||
|
},
|
||||||
|
[draggingItem, dragOffset],
|
||||||
|
);
|
||||||
|
|
||||||
|
const handleDragEnd = useCallback(() => {
|
||||||
|
if (!draggingItem || !tracksRef.current) {
|
||||||
|
setDraggingItem(null);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const item = items.find((i) => i.id === draggingItem);
|
||||||
|
if (!item) {
|
||||||
|
setDraggingItem(null);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Calculate new time from x position
|
||||||
|
const newTimeMs = Math.max(0, Math.round(pixelsToMs(dragPosition.x)));
|
||||||
|
|
||||||
|
// Calculate new track from y position
|
||||||
|
const trackIndex = Math.floor(dragPosition.y / TRACK_HEIGHT);
|
||||||
|
const clampedTrackIndex = Math.max(0, Math.min(trackIndex, tracks.length - 1));
|
||||||
|
const newTrack = tracks[clampedTrackIndex] ?? 0;
|
||||||
|
|
||||||
|
// Check if position changed
|
||||||
|
if (newTimeMs !== item.start_time_ms || newTrack !== item.track) {
|
||||||
|
moveItem.mutate(
|
||||||
|
{
|
||||||
|
storyId,
|
||||||
|
itemId: item.id,
|
||||||
|
data: {
|
||||||
|
start_time_ms: newTimeMs,
|
||||||
|
track: newTrack,
|
||||||
|
},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
onError: (error) => {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to move item',
|
||||||
|
description: error instanceof Error ? error.message : String(error),
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
setDraggingItem(null);
|
||||||
|
}, [draggingItem, dragPosition, items, tracks, pixelsToMs, storyId, moveItem, toast]);
|
||||||
|
|
||||||
|
// Get track index for rendering
|
||||||
|
const getTrackIndex = (trackNumber: number) => tracks.indexOf(trackNumber);
|
||||||
|
|
||||||
|
// Calculate clip position and dimensions
|
||||||
|
const getClipStyle = (item: StoryItemDetail) => {
|
||||||
|
const isDragging = draggingItem === item.id;
|
||||||
|
const trackIndex = getTrackIndex(item.track);
|
||||||
|
const effectiveDuration = getEffectiveDuration(item);
|
||||||
|
const width = msToPixels(effectiveDuration);
|
||||||
|
const left = isDragging ? dragPosition.x : msToPixels(item.start_time_ms);
|
||||||
|
const top = isDragging ? dragPosition.y : trackIndex * TRACK_HEIGHT;
|
||||||
|
|
||||||
|
return {
|
||||||
|
width: `${width}px`,
|
||||||
|
left: `${left}px`,
|
||||||
|
top: `${top}px`,
|
||||||
|
height: `${TRACK_HEIGHT - 4}px`,
|
||||||
|
};
|
||||||
|
};
|
||||||
|
|
||||||
|
// Playhead position
|
||||||
|
const playheadLeft = msToPixels(currentTimeMs);
|
||||||
|
|
||||||
|
// Auto-scroll timeline to follow playhead during playback
|
||||||
|
useEffect(() => {
|
||||||
|
if (!isCurrentlyPlaying || !tracksRef.current) return;
|
||||||
|
|
||||||
|
const container = tracksRef.current;
|
||||||
|
const containerWidth = container.clientWidth;
|
||||||
|
const scrollLeft = container.scrollLeft;
|
||||||
|
const halfwayPoint = scrollLeft + containerWidth / 2;
|
||||||
|
|
||||||
|
// If playhead is past the halfway point, scroll to keep it centered
|
||||||
|
if (playheadLeft > halfwayPoint) {
|
||||||
|
const targetScroll = playheadLeft - containerWidth / 2;
|
||||||
|
container.scrollLeft = targetScroll;
|
||||||
|
}
|
||||||
|
}, [isCurrentlyPlaying, playheadLeft]);
|
||||||
|
|
||||||
|
// Calculate tracks area height
|
||||||
|
const tracksAreaHeight = tracks.length * TRACK_HEIGHT;
|
||||||
|
const timelineContainerHeight = editorHeight - 40; // Subtract toolbar height
|
||||||
|
|
||||||
|
if (items.length === 0) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="fixed bottom-0 left-0 right-0 border-t bg-background/95 backdrop-blur supports-backdrop-filter:bg-background/60 z-50">
|
||||||
|
<div
|
||||||
|
className="border-t bg-background/30 backdrop-blur-2xl overflow-hidden relative"
|
||||||
|
ref={containerRef}
|
||||||
|
>
|
||||||
|
{/* Resize handle at top */}
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="absolute top-0 left-0 right-0 h-2 cursor-ns-resize flex items-center justify-center hover:bg-muted/50 transition-colors z-20 group"
|
||||||
|
onMouseDown={handleResizeStart}
|
||||||
|
aria-label="Resize track editor"
|
||||||
|
>
|
||||||
|
<GripHorizontal className="h-3 w-3 text-muted-foreground/50 group-hover:text-muted-foreground" />
|
||||||
|
</button>
|
||||||
|
|
||||||
|
{/* Toolbar */}
|
||||||
|
<div className="flex items-center justify-between px-3 py-2 border-b bg-muted/30 mt-2">
|
||||||
|
{/* Play controls - left side */}
|
||||||
|
<div className="flex items-center gap-2">
|
||||||
|
<Button
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-7 w-7"
|
||||||
|
onClick={handlePlayPause}
|
||||||
|
title="Play/Pause (Space)"
|
||||||
|
>
|
||||||
|
{isCurrentlyPlaying ? <Pause className="h-4 w-4" /> : <Play className="h-4 w-4" />}
|
||||||
|
</Button>
|
||||||
|
<Button
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-7 w-7"
|
||||||
|
onClick={handleStop}
|
||||||
|
disabled={!isCurrentlyPlaying}
|
||||||
|
>
|
||||||
|
<Square className="h-3 w-3" />
|
||||||
|
</Button>
|
||||||
|
<span className="text-xs text-muted-foreground tabular-nums ml-2">
|
||||||
|
{formatTime(currentTimeMs)} / {formatTime(totalDurationMs)}
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Clip editing controls - center */}
|
||||||
|
{selectedClipId && (
|
||||||
|
<div className="flex items-center gap-1">
|
||||||
|
<Button
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-7 w-7"
|
||||||
|
onClick={handleSplit}
|
||||||
|
title="Split at playhead (S)"
|
||||||
|
>
|
||||||
|
<Scissors className="h-4 w-4" />
|
||||||
|
</Button>
|
||||||
|
<Button
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-7 w-7"
|
||||||
|
onClick={handleDuplicate}
|
||||||
|
title="Duplicate (Cmd/Ctrl+D)"
|
||||||
|
>
|
||||||
|
<Copy className="h-4 w-4" />
|
||||||
|
</Button>
|
||||||
|
<Button
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-7 w-7"
|
||||||
|
onClick={handleDelete}
|
||||||
|
title="Delete (Delete/Backspace)"
|
||||||
|
>
|
||||||
|
<Trash2 className="h-4 w-4" />
|
||||||
|
</Button>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{/* Zoom controls - right side */}
|
||||||
|
<div className="flex items-center gap-2">
|
||||||
|
<span className="text-xs text-muted-foreground">Zoom:</span>
|
||||||
|
<Button variant="ghost" size="icon" className="h-6 w-6" onClick={handleZoomOut}>
|
||||||
|
<Minus className="h-3 w-3" />
|
||||||
|
</Button>
|
||||||
|
<Button variant="ghost" size="icon" className="h-6 w-6" onClick={handleZoomIn}>
|
||||||
|
<Plus className="h-3 w-3" />
|
||||||
|
</Button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Timeline container with track labels sidebar */}
|
||||||
|
<div className="flex" style={{ height: `${timelineContainerHeight}px` }}>
|
||||||
|
{/* Track labels sidebar - fixed width */}
|
||||||
|
<div className="w-16 shrink-0 border-r bg-muted/20 overflow-hidden">
|
||||||
|
{/* Spacer for time ruler */}
|
||||||
|
<div className="h-6 border-b bg-muted/30" />
|
||||||
|
{/* Track labels */}
|
||||||
|
<div style={{ height: `${tracksAreaHeight}px` }}>
|
||||||
|
{tracks.map((trackNumber, index) => (
|
||||||
|
<div
|
||||||
|
key={trackNumber}
|
||||||
|
className={cn(
|
||||||
|
'border-b flex items-center justify-center',
|
||||||
|
index % 2 === 0 ? 'bg-background' : 'bg-muted/10',
|
||||||
|
)}
|
||||||
|
style={{ height: `${TRACK_HEIGHT}px` }}
|
||||||
|
>
|
||||||
|
<span className="text-[10px] text-muted-foreground select-none">
|
||||||
|
{trackNumber}
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Scrollable timeline area */}
|
||||||
|
{/* biome-ignore lint/a11y/noStaticElementInteractions: Container handles drag events for child clips */}
|
||||||
|
<div
|
||||||
|
ref={tracksRef}
|
||||||
|
className="overflow-auto relative flex-1"
|
||||||
|
onMouseMove={draggingItem ? handleDragMove : undefined}
|
||||||
|
onMouseUp={draggingItem ? handleDragEnd : undefined}
|
||||||
|
onMouseLeave={draggingItem ? handleDragEnd : undefined}
|
||||||
|
>
|
||||||
|
{/* Time ruler - clickable to seek */}
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="h-6 border-b bg-muted/20 sticky top-0 z-10 cursor-pointer text-left"
|
||||||
|
style={{ width: `${timelineWidth}px` }}
|
||||||
|
onClick={handleTimelineClick}
|
||||||
|
aria-label="Seek timeline"
|
||||||
|
>
|
||||||
|
{timeMarkers.map((ms) => (
|
||||||
|
<div
|
||||||
|
key={ms}
|
||||||
|
className="absolute top-0 h-full flex flex-col justify-end pointer-events-none"
|
||||||
|
style={{ left: `${msToPixels(ms)}px` }}
|
||||||
|
>
|
||||||
|
<div className="h-2 w-px bg-border" />
|
||||||
|
<span className="text-[10px] text-muted-foreground ml-1 select-none">
|
||||||
|
{formatTime(ms)}
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
</button>
|
||||||
|
|
||||||
|
{/* Tracks area */}
|
||||||
|
<div
|
||||||
|
className="relative"
|
||||||
|
style={{ width: `${timelineWidth}px`, height: `${tracksAreaHeight}px` }}
|
||||||
|
>
|
||||||
|
{/* Track backgrounds - pointer-events-none to allow clicks to pass through */}
|
||||||
|
{tracks.map((trackNumber, index) => (
|
||||||
|
<div
|
||||||
|
key={trackNumber}
|
||||||
|
className={cn(
|
||||||
|
'absolute left-0 right-0 border-b pointer-events-none',
|
||||||
|
index % 2 === 0 ? 'bg-background' : 'bg-muted/10',
|
||||||
|
)}
|
||||||
|
style={{
|
||||||
|
top: `${index * TRACK_HEIGHT}px`,
|
||||||
|
height: `${TRACK_HEIGHT}px`,
|
||||||
|
}}
|
||||||
|
/>
|
||||||
|
))}
|
||||||
|
|
||||||
|
{/* Click area for seeking - z-index lower than clips */}
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="absolute inset-0 z-0 cursor-pointer"
|
||||||
|
onClick={handleTimelineClick}
|
||||||
|
aria-label="Seek timeline"
|
||||||
|
/>
|
||||||
|
|
||||||
|
{/* Audio clips */}
|
||||||
|
{items.map((item) => {
|
||||||
|
const isDragging = draggingItem === item.id;
|
||||||
|
const isSelected = selectedClipId === item.id;
|
||||||
|
const isTrimming = trimmingItem === item.id;
|
||||||
|
|
||||||
|
// Use temporary trim values during trimming for visual feedback
|
||||||
|
const displayTrimStart =
|
||||||
|
isTrimming && tempTrimValues
|
||||||
|
? tempTrimValues.trim_start_ms
|
||||||
|
: item.trim_start_ms || 0;
|
||||||
|
const displayTrimEnd =
|
||||||
|
isTrimming && tempTrimValues ? tempTrimValues.trim_end_ms : item.trim_end_ms || 0;
|
||||||
|
const effectiveDuration = item.duration * 1000 - displayTrimStart - displayTrimEnd;
|
||||||
|
|
||||||
|
const style = getClipStyle({
|
||||||
|
...item,
|
||||||
|
trim_start_ms: displayTrimStart,
|
||||||
|
trim_end_ms: displayTrimEnd,
|
||||||
|
});
|
||||||
|
const clipWidth = msToPixels(effectiveDuration);
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div
|
||||||
|
key={item.id}
|
||||||
|
className={cn(
|
||||||
|
'absolute rounded select-none overflow-visible z-10',
|
||||||
|
isSelected && 'ring-2 ring-primary ring-offset-1',
|
||||||
|
isTrimming && 'ring-2 ring-accent',
|
||||||
|
)}
|
||||||
|
style={style}
|
||||||
|
>
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className={cn(
|
||||||
|
'w-full h-full rounded cursor-move overflow-hidden',
|
||||||
|
'bg-accent/80 hover:bg-accent border border-accent-foreground/20',
|
||||||
|
'flex flex-col justify-center',
|
||||||
|
isDragging && 'opacity-80 shadow-lg z-20',
|
||||||
|
!isDragging && 'transition-all duration-100',
|
||||||
|
)}
|
||||||
|
onClick={(e) => handleClipClick(e, item)}
|
||||||
|
onMouseDown={(e) => {
|
||||||
|
// Only start drag if not clicking on trim handles
|
||||||
|
if (!(e.target as HTMLElement).closest('.trim-handle')) {
|
||||||
|
handleDragStart(e, item);
|
||||||
|
}
|
||||||
|
}}
|
||||||
|
>
|
||||||
|
{/* Clip label */}
|
||||||
|
<div className="absolute top-0 left-1 right-1 z-10">
|
||||||
|
<p className="text-[9px] font-medium text-accent-foreground truncate">
|
||||||
|
{item.profile_name}
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
{/* Waveform */}
|
||||||
|
<div className="absolute inset-0 top-3">
|
||||||
|
<ClipWaveform
|
||||||
|
generationId={item.generation_id}
|
||||||
|
width={clipWidth}
|
||||||
|
trimStartMs={displayTrimStart}
|
||||||
|
trimEndMs={displayTrimEnd}
|
||||||
|
duration={item.duration}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
</button>
|
||||||
|
|
||||||
|
{/* Trim handles */}
|
||||||
|
{isSelected && (
|
||||||
|
<>
|
||||||
|
{/* Left trim handle */}
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="trim-handle absolute left-0 top-0 bottom-0 w-2 cursor-ew-resize hover:bg-primary/30 bg-primary/20 z-30 rounded-l"
|
||||||
|
onMouseDown={(e) => handleTrimStart(e, item, 'start')}
|
||||||
|
aria-label="Trim start"
|
||||||
|
/>
|
||||||
|
{/* Right trim handle */}
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="trim-handle absolute right-0 top-0 bottom-0 w-2 cursor-ew-resize hover:bg-primary/30 bg-primary/20 z-30 rounded-r"
|
||||||
|
onMouseDown={(e) => handleTrimStart(e, item, 'end')}
|
||||||
|
aria-label="Trim end"
|
||||||
|
/>
|
||||||
|
</>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
})}
|
||||||
|
|
||||||
|
{/* Playhead - always visible */}
|
||||||
|
<div
|
||||||
|
className="absolute top-0 bottom-0 w-1 bg-accent z-30 pointer-events-none rounded-full"
|
||||||
|
style={{ left: `${playheadLeft}px` }}
|
||||||
|
>
|
||||||
|
<div className="absolute -top-1 left-1/2 -translate-x-1/2 w-3 h-3 bg-accent rounded-full" />
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -1,8 +1,31 @@
|
|||||||
import { Mic, Pause, Play, Square } from 'lucide-react';
|
import { Mic, Pause, Play, Square } from 'lucide-react';
|
||||||
|
import { memo, useEffect, useState } from 'react';
|
||||||
|
import { Visualizer } from 'react-sound-visualizer';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { FormControl, FormItem, FormLabel, FormMessage } from '@/components/ui/form';
|
import { FormControl, FormItem, FormMessage } from '@/components/ui/form';
|
||||||
import { formatAudioDuration } from '@/lib/utils/audio';
|
import { formatAudioDuration } from '@/lib/utils/audio';
|
||||||
|
|
||||||
|
const MemoizedWaveform = memo(function MemoizedWaveform({
|
||||||
|
audioStream,
|
||||||
|
}: {
|
||||||
|
audioStream: MediaStream;
|
||||||
|
}) {
|
||||||
|
return (
|
||||||
|
<div className="absolute inset-0 pointer-events-none flex items-center justify-center opacity-30">
|
||||||
|
<Visualizer audio={audioStream} autoStart strokeColor="#b39a3d">
|
||||||
|
{({ canvasRef }) => (
|
||||||
|
<canvas
|
||||||
|
ref={canvasRef}
|
||||||
|
width={500}
|
||||||
|
height={150}
|
||||||
|
className="w-full h-full"
|
||||||
|
/>
|
||||||
|
)}
|
||||||
|
</Visualizer>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
interface AudioSampleRecordingProps {
|
interface AudioSampleRecordingProps {
|
||||||
file: File | null | undefined;
|
file: File | null | undefined;
|
||||||
isRecording: boolean;
|
isRecording: boolean;
|
||||||
@@ -14,6 +37,7 @@ interface AudioSampleRecordingProps {
|
|||||||
onPlayPause: () => void;
|
onPlayPause: () => void;
|
||||||
isPlaying: boolean;
|
isPlaying: boolean;
|
||||||
isTranscribing?: boolean;
|
isTranscribing?: boolean;
|
||||||
|
showWaveform?: boolean;
|
||||||
}
|
}
|
||||||
|
|
||||||
export function AudioSampleRecording({
|
export function AudioSampleRecording({
|
||||||
@@ -27,29 +51,67 @@ export function AudioSampleRecording({
|
|||||||
onPlayPause,
|
onPlayPause,
|
||||||
isPlaying,
|
isPlaying,
|
||||||
isTranscribing = false,
|
isTranscribing = false,
|
||||||
|
showWaveform = true,
|
||||||
}: AudioSampleRecordingProps) {
|
}: AudioSampleRecordingProps) {
|
||||||
|
const [audioStream, setAudioStream] = useState<MediaStream | null>(null);
|
||||||
|
|
||||||
|
// Request microphone access when component mounts
|
||||||
|
useEffect(() => {
|
||||||
|
if (!showWaveform) return;
|
||||||
|
|
||||||
|
let stream: MediaStream | null = null;
|
||||||
|
|
||||||
|
navigator.mediaDevices
|
||||||
|
.getUserMedia({ audio: true, video: false })
|
||||||
|
.then((s) => {
|
||||||
|
stream = s;
|
||||||
|
setAudioStream(s);
|
||||||
|
})
|
||||||
|
.catch((err) => {
|
||||||
|
console.warn('Could not access microphone for visualization:', err);
|
||||||
|
});
|
||||||
|
|
||||||
|
return () => {
|
||||||
|
if (stream) {
|
||||||
|
stream.getTracks().forEach((track) => {
|
||||||
|
track.stop();
|
||||||
|
});
|
||||||
|
}
|
||||||
|
};
|
||||||
|
}, [showWaveform]);
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<FormItem>
|
<FormItem>
|
||||||
<FormLabel>Record Audio</FormLabel>
|
|
||||||
<FormControl>
|
<FormControl>
|
||||||
<div className="space-y-4">
|
<div className="space-y-4">
|
||||||
{!isRecording && !file && (
|
{!isRecording && !file && (
|
||||||
<div className="flex flex-col items-center justify-center gap-4 p-4 border-2 border-dashed rounded-lg min-h-[180px]">
|
<div className="relative flex flex-col items-center justify-center gap-4 p-4 border-2 border-dashed rounded-lg min-h-[180px] overflow-hidden">
|
||||||
<Button type="button" onClick={onStart} size="lg" className="flex items-center gap-2">
|
{showWaveform && audioStream && (
|
||||||
|
<MemoizedWaveform audioStream={audioStream} />
|
||||||
|
)}
|
||||||
|
<Button
|
||||||
|
type="button"
|
||||||
|
onClick={onStart}
|
||||||
|
size="lg"
|
||||||
|
className="relative z-10 flex items-center gap-2"
|
||||||
|
>
|
||||||
<Mic className="h-5 w-5" />
|
<Mic className="h-5 w-5" />
|
||||||
Start Recording
|
Start Recording
|
||||||
</Button>
|
</Button>
|
||||||
<p className="text-sm text-muted-foreground text-center">
|
<p className="relative z-10 text-sm text-muted-foreground text-center">
|
||||||
Click to start recording. Maximum duration: 30 seconds.
|
Click to start recording. Maximum duration: 30 seconds.
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
{isRecording && (
|
{isRecording && (
|
||||||
<div className="flex flex-col items-center justify-center gap-4 p-4 border-2 border-destructive rounded-lg bg-destructive/5 min-h-[180px]">
|
<div className="relative flex flex-col items-center justify-center gap-4 p-4 border-2 border-accent rounded-lg bg-accent/5 min-h-[180px] overflow-hidden">
|
||||||
<div className="flex items-center gap-4">
|
{showWaveform && audioStream && (
|
||||||
|
<MemoizedWaveform audioStream={audioStream} />
|
||||||
|
)}
|
||||||
|
<div className="relative z-10 flex items-center gap-4">
|
||||||
<div className="flex items-center gap-2">
|
<div className="flex items-center gap-2">
|
||||||
<div className="h-3 w-3 rounded-full bg-destructive animate-pulse" />
|
<div className="h-3 w-3 rounded-full bg-accent animate-pulse" />
|
||||||
<span className="text-lg font-mono font-semibold">
|
<span className="text-lg font-mono font-semibold">
|
||||||
{formatAudioDuration(duration)}
|
{formatAudioDuration(duration)}
|
||||||
</span>
|
</span>
|
||||||
@@ -58,13 +120,12 @@ export function AudioSampleRecording({
|
|||||||
<Button
|
<Button
|
||||||
type="button"
|
type="button"
|
||||||
onClick={onStop}
|
onClick={onStop}
|
||||||
variant="destructive"
|
className="relative z-10 flex items-center gap-2 bg-accent text-accent-foreground hover:bg-accent/90"
|
||||||
className="flex items-center gap-2"
|
|
||||||
>
|
>
|
||||||
<Square className="h-4 w-4" />
|
<Square className="h-4 w-4" />
|
||||||
Stop Recording
|
Stop Recording
|
||||||
</Button>
|
</Button>
|
||||||
<p className="text-sm text-muted-foreground text-center">
|
<p className="relative z-10 text-sm text-muted-foreground text-center">
|
||||||
{formatAudioDuration(30 - duration)} remaining
|
{formatAudioDuration(30 - duration)} remaining
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
import { Mic, Monitor, Pause, Play, Square } from 'lucide-react';
|
import { Mic, Monitor, Pause, Play, Square } from 'lucide-react';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { FormControl, FormItem, FormLabel, FormMessage } from '@/components/ui/form';
|
import { FormControl, FormItem, FormMessage } from '@/components/ui/form';
|
||||||
import { formatAudioDuration } from '@/lib/utils/audio';
|
import { formatAudioDuration } from '@/lib/utils/audio';
|
||||||
|
|
||||||
interface AudioSampleSystemProps {
|
interface AudioSampleSystemProps {
|
||||||
@@ -30,7 +30,6 @@ export function AudioSampleSystem({
|
|||||||
}: AudioSampleSystemProps) {
|
}: AudioSampleSystemProps) {
|
||||||
return (
|
return (
|
||||||
<FormItem>
|
<FormItem>
|
||||||
<FormLabel>Capture System Audio</FormLabel>
|
|
||||||
<FormControl>
|
<FormControl>
|
||||||
<div className="space-y-4">
|
<div className="space-y-4">
|
||||||
{!isRecording && !file && (
|
{!isRecording && !file && (
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
import { Mic, Pause, Play, Upload } from 'lucide-react';
|
import { Mic, Pause, Play, Upload } from 'lucide-react';
|
||||||
import { useRef, useState } from 'react';
|
import { useRef, useState } from 'react';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { FormControl, FormItem, FormLabel, FormMessage } from '@/components/ui/form';
|
import { FormControl, FormItem, FormMessage } from '@/components/ui/form';
|
||||||
|
|
||||||
interface AudioSampleUploadProps {
|
interface AudioSampleUploadProps {
|
||||||
file: File | null | undefined;
|
file: File | null | undefined;
|
||||||
@@ -31,7 +31,6 @@ export function AudioSampleUpload({
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<FormItem>
|
<FormItem>
|
||||||
<FormLabel>Audio File</FormLabel>
|
|
||||||
<FormControl>
|
<FormControl>
|
||||||
<div className="flex flex-col gap-2">
|
<div className="flex flex-col gap-2">
|
||||||
<input
|
<input
|
||||||
|
|||||||
@@ -15,6 +15,7 @@ import {
|
|||||||
import type { VoiceProfileResponse } from '@/lib/api/types';
|
import type { VoiceProfileResponse } from '@/lib/api/types';
|
||||||
import { useDeleteProfile, useExportProfile } from '@/lib/hooks/useProfiles';
|
import { useDeleteProfile, useExportProfile } from '@/lib/hooks/useProfiles';
|
||||||
import { cn } from '@/lib/utils/cn';
|
import { cn } from '@/lib/utils/cn';
|
||||||
|
import { useServerStore } from '@/stores/serverStore';
|
||||||
import { useUIStore } from '@/stores/uiStore';
|
import { useUIStore } from '@/stores/uiStore';
|
||||||
|
|
||||||
interface ProfileCardProps {
|
interface ProfileCardProps {
|
||||||
@@ -23,15 +24,19 @@ interface ProfileCardProps {
|
|||||||
|
|
||||||
export function ProfileCard({ profile }: ProfileCardProps) {
|
export function ProfileCard({ profile }: ProfileCardProps) {
|
||||||
const [deleteDialogOpen, setDeleteDialogOpen] = useState(false);
|
const [deleteDialogOpen, setDeleteDialogOpen] = useState(false);
|
||||||
|
const [avatarError, setAvatarError] = useState(false);
|
||||||
const deleteProfile = useDeleteProfile();
|
const deleteProfile = useDeleteProfile();
|
||||||
const exportProfile = useExportProfile();
|
const exportProfile = useExportProfile();
|
||||||
const setEditingProfileId = useUIStore((state) => state.setEditingProfileId);
|
const setEditingProfileId = useUIStore((state) => state.setEditingProfileId);
|
||||||
const setProfileDialogOpen = useUIStore((state) => state.setProfileDialogOpen);
|
const setProfileDialogOpen = useUIStore((state) => state.setProfileDialogOpen);
|
||||||
const selectedProfileId = useUIStore((state) => state.selectedProfileId);
|
const selectedProfileId = useUIStore((state) => state.selectedProfileId);
|
||||||
const setSelectedProfileId = useUIStore((state) => state.setSelectedProfileId);
|
const setSelectedProfileId = useUIStore((state) => state.setSelectedProfileId);
|
||||||
|
const serverUrl = useServerStore((state) => state.serverUrl);
|
||||||
|
|
||||||
const isSelected = selectedProfileId === profile.id;
|
const isSelected = selectedProfileId === profile.id;
|
||||||
|
|
||||||
|
const avatarUrl = profile.avatar_path ? `${serverUrl}/profiles/${profile.id}/avatar` : null;
|
||||||
|
|
||||||
const handleSelect = () => {
|
const handleSelect = () => {
|
||||||
setSelectedProfileId(isSelected ? null : profile.id);
|
setSelectedProfileId(isSelected ? null : profile.id);
|
||||||
};
|
};
|
||||||
@@ -67,8 +72,20 @@ export function ProfileCard({ profile }: ProfileCardProps) {
|
|||||||
>
|
>
|
||||||
<CardHeader className="p-3 pb-2">
|
<CardHeader className="p-3 pb-2">
|
||||||
<CardTitle className="flex items-center gap-1.5 text-base font-medium">
|
<CardTitle className="flex items-center gap-1.5 text-base font-medium">
|
||||||
<div className="h-6 w-6 rounded-full bg-muted flex items-center justify-center shrink-0">
|
<div className="h-6 w-6 rounded-full bg-muted flex items-center justify-center shrink-0 overflow-hidden">
|
||||||
<Mic className="h-3.5 w-3.5 text-muted-foreground" />
|
{avatarUrl && !avatarError ? (
|
||||||
|
<img
|
||||||
|
src={avatarUrl}
|
||||||
|
alt={`${profile.name} avatar`}
|
||||||
|
className={cn(
|
||||||
|
'h-full w-full object-cover transition-all duration-200',
|
||||||
|
!isSelected && 'grayscale',
|
||||||
|
)}
|
||||||
|
onError={() => setAvatarError(true)}
|
||||||
|
/>
|
||||||
|
) : (
|
||||||
|
<Mic className="h-3.5 w-3.5 text-muted-foreground" />
|
||||||
|
)}
|
||||||
</div>
|
</div>
|
||||||
<span className="break-words">{profile.name}</span>
|
<span className="break-words">{profile.name}</span>
|
||||||
</CardTitle>
|
</CardTitle>
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
import { zodResolver } from '@hookform/resolvers/zod';
|
import { zodResolver } from '@hookform/resolvers/zod';
|
||||||
import { Mic, Monitor, Upload } from 'lucide-react';
|
import { Edit2, Mic, Monitor, Upload, X } from 'lucide-react';
|
||||||
import { useEffect, useState } from 'react';
|
import { useEffect, useRef, useState } from 'react';
|
||||||
import { useForm } from 'react-hook-form';
|
import { useForm } from 'react-hook-form';
|
||||||
import * as z from 'zod';
|
import * as z from 'zod';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
@@ -36,51 +36,22 @@ import { useAudioRecording } from '@/lib/hooks/useAudioRecording';
|
|||||||
import {
|
import {
|
||||||
useAddSample,
|
useAddSample,
|
||||||
useCreateProfile,
|
useCreateProfile,
|
||||||
|
useDeleteAvatar,
|
||||||
useProfile,
|
useProfile,
|
||||||
useUpdateProfile,
|
useUpdateProfile,
|
||||||
|
useUploadAvatar,
|
||||||
} from '@/lib/hooks/useProfiles';
|
} from '@/lib/hooks/useProfiles';
|
||||||
import { useSystemAudioCapture } from '@/lib/hooks/useSystemAudioCapture';
|
import { useSystemAudioCapture } from '@/lib/hooks/useSystemAudioCapture';
|
||||||
import { useTranscription } from '@/lib/hooks/useTranscription';
|
import { useTranscription } from '@/lib/hooks/useTranscription';
|
||||||
import { isTauri } from '@/lib/tauri';
|
import { formatAudioDuration, getAudioDuration } from '@/lib/utils/audio';
|
||||||
import { formatAudioDuration } from '@/lib/utils/audio';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
import { useUIStore } from '@/stores/uiStore';
|
import { useServerStore } from '@/stores/serverStore';
|
||||||
|
import { type ProfileFormDraft, useUIStore } from '@/stores/uiStore';
|
||||||
import { AudioSampleRecording } from './AudioSampleRecording';
|
import { AudioSampleRecording } from './AudioSampleRecording';
|
||||||
import { AudioSampleSystem } from './AudioSampleSystem';
|
import { AudioSampleSystem } from './AudioSampleSystem';
|
||||||
import { AudioSampleUpload } from './AudioSampleUpload';
|
import { AudioSampleUpload } from './AudioSampleUpload';
|
||||||
import { SampleList } from './SampleList';
|
import { SampleList } from './SampleList';
|
||||||
|
|
||||||
// Helper function to get audio duration from File
|
|
||||||
async function getAudioDuration(file: File & { recordedDuration?: number }): Promise<number> {
|
|
||||||
// If the file has a recordedDuration property (from our recording hooks),
|
|
||||||
// use that instead of trying to read metadata. This fixes issues on Windows
|
|
||||||
// where WebM files from MediaRecorder don't have proper duration metadata.
|
|
||||||
if (file.recordedDuration !== undefined && Number.isFinite(file.recordedDuration)) {
|
|
||||||
return file.recordedDuration;
|
|
||||||
}
|
|
||||||
|
|
||||||
return new Promise((resolve, reject) => {
|
|
||||||
const audio = new Audio();
|
|
||||||
const url = URL.createObjectURL(file);
|
|
||||||
|
|
||||||
audio.addEventListener('loadedmetadata', () => {
|
|
||||||
URL.revokeObjectURL(url);
|
|
||||||
// Check if duration is valid (not Infinity or NaN)
|
|
||||||
if (Number.isFinite(audio.duration) && audio.duration > 0) {
|
|
||||||
resolve(audio.duration);
|
|
||||||
} else {
|
|
||||||
reject(new Error('Audio file has invalid duration metadata'));
|
|
||||||
}
|
|
||||||
});
|
|
||||||
|
|
||||||
audio.addEventListener('error', () => {
|
|
||||||
URL.revokeObjectURL(url);
|
|
||||||
reject(new Error('Failed to load audio file'));
|
|
||||||
});
|
|
||||||
|
|
||||||
audio.src = url;
|
|
||||||
});
|
|
||||||
}
|
|
||||||
|
|
||||||
const MAX_AUDIO_DURATION_SECONDS = 30;
|
const MAX_AUDIO_DURATION_SECONDS = 30;
|
||||||
|
|
||||||
const baseProfileSchema = z.object({
|
const baseProfileSchema = z.object({
|
||||||
@@ -89,6 +60,7 @@ const baseProfileSchema = z.object({
|
|||||||
language: z.enum(LANGUAGE_CODES as [LanguageCode, ...LanguageCode[]]),
|
language: z.enum(LANGUAGE_CODES as [LanguageCode, ...LanguageCode[]]),
|
||||||
sampleFile: z.instanceof(File).optional(),
|
sampleFile: z.instanceof(File).optional(),
|
||||||
referenceText: z.string().max(1000).optional(),
|
referenceText: z.string().max(1000).optional(),
|
||||||
|
avatarFile: z.instanceof(File).optional(),
|
||||||
});
|
});
|
||||||
|
|
||||||
const profileSchema = baseProfileSchema.refine(
|
const profileSchema = baseProfileSchema.refine(
|
||||||
@@ -107,22 +79,52 @@ const profileSchema = baseProfileSchema.refine(
|
|||||||
|
|
||||||
type ProfileFormValues = z.infer<typeof profileSchema>;
|
type ProfileFormValues = z.infer<typeof profileSchema>;
|
||||||
|
|
||||||
|
// Helper to convert File to base64
|
||||||
|
async function fileToBase64(file: File): Promise<string> {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const reader = new FileReader();
|
||||||
|
reader.onload = () => resolve(reader.result as string);
|
||||||
|
reader.onerror = reject;
|
||||||
|
reader.readAsDataURL(file);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
// Helper to convert base64 to File
|
||||||
|
function base64ToFile(base64: string, fileName: string, fileType: string): File {
|
||||||
|
const arr = base64.split(',');
|
||||||
|
const bstr = atob(arr[1]);
|
||||||
|
let n = bstr.length;
|
||||||
|
const u8arr = new Uint8Array(n);
|
||||||
|
while (n--) {
|
||||||
|
u8arr[n] = bstr.charCodeAt(n);
|
||||||
|
}
|
||||||
|
return new File([u8arr], fileName, { type: fileType });
|
||||||
|
}
|
||||||
|
|
||||||
export function ProfileForm() {
|
export function ProfileForm() {
|
||||||
|
const platform = usePlatform();
|
||||||
const open = useUIStore((state) => state.profileDialogOpen);
|
const open = useUIStore((state) => state.profileDialogOpen);
|
||||||
const setOpen = useUIStore((state) => state.setProfileDialogOpen);
|
const setOpen = useUIStore((state) => state.setProfileDialogOpen);
|
||||||
const editingProfileId = useUIStore((state) => state.editingProfileId);
|
const editingProfileId = useUIStore((state) => state.editingProfileId);
|
||||||
const setEditingProfileId = useUIStore((state) => state.setEditingProfileId);
|
const setEditingProfileId = useUIStore((state) => state.setEditingProfileId);
|
||||||
|
const profileFormDraft = useUIStore((state) => state.profileFormDraft);
|
||||||
|
const setProfileFormDraft = useUIStore((state) => state.setProfileFormDraft);
|
||||||
const { data: editingProfile } = useProfile(editingProfileId || '');
|
const { data: editingProfile } = useProfile(editingProfileId || '');
|
||||||
const createProfile = useCreateProfile();
|
const createProfile = useCreateProfile();
|
||||||
const updateProfile = useUpdateProfile();
|
const updateProfile = useUpdateProfile();
|
||||||
const addSample = useAddSample();
|
const addSample = useAddSample();
|
||||||
|
const uploadAvatar = useUploadAvatar();
|
||||||
|
const deleteAvatar = useDeleteAvatar();
|
||||||
const transcribe = useTranscription();
|
const transcribe = useTranscription();
|
||||||
const { toast } = useToast();
|
const { toast } = useToast();
|
||||||
const [sampleMode, setSampleMode] = useState<'upload' | 'record' | 'system'>('upload');
|
const [sampleMode, setSampleMode] = useState<'upload' | 'record' | 'system'>('record');
|
||||||
const [audioDuration, setAudioDuration] = useState<number | null>(null);
|
const [audioDuration, setAudioDuration] = useState<number | null>(null);
|
||||||
const [isValidatingAudio, setIsValidatingAudio] = useState(false);
|
const [isValidatingAudio, setIsValidatingAudio] = useState(false);
|
||||||
|
const [avatarPreview, setAvatarPreview] = useState<string | null>(null);
|
||||||
|
const avatarInputRef = useRef<HTMLInputElement>(null);
|
||||||
const { isPlaying, playPause, cleanup: cleanupAudio } = useAudioPlayer();
|
const { isPlaying, playPause, cleanup: cleanupAudio } = useAudioPlayer();
|
||||||
const isCreating = !editingProfileId;
|
const isCreating = !editingProfileId;
|
||||||
|
const serverUrl = useServerStore((state) => state.serverUrl);
|
||||||
|
|
||||||
const form = useForm<ProfileFormValues>({
|
const form = useForm<ProfileFormValues>({
|
||||||
resolver: zodResolver(profileSchema),
|
resolver: zodResolver(profileSchema),
|
||||||
@@ -132,10 +134,12 @@ export function ProfileForm() {
|
|||||||
language: 'en',
|
language: 'en',
|
||||||
sampleFile: undefined,
|
sampleFile: undefined,
|
||||||
referenceText: '',
|
referenceText: '',
|
||||||
|
avatarFile: undefined,
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
const selectedFile = form.watch('sampleFile');
|
const selectedFile = form.watch('sampleFile');
|
||||||
|
const selectedAvatarFile = form.watch('avatarFile');
|
||||||
|
|
||||||
// Validate audio duration when file is selected
|
// Validate audio duration when file is selected
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
@@ -187,7 +191,7 @@ export function ProfileForm() {
|
|||||||
stopRecording,
|
stopRecording,
|
||||||
cancelRecording,
|
cancelRecording,
|
||||||
} = useAudioRecording({
|
} = useAudioRecording({
|
||||||
maxDurationSeconds: 30,
|
maxDurationSeconds: 29,
|
||||||
onRecordingComplete: (blob, recordedDuration) => {
|
onRecordingComplete: (blob, recordedDuration) => {
|
||||||
const file = new File([blob], `recording-${Date.now()}.webm`, {
|
const file = new File([blob], `recording-${Date.now()}.webm`, {
|
||||||
type: blob.type || 'audio/webm',
|
type: blob.type || 'audio/webm',
|
||||||
@@ -213,7 +217,7 @@ export function ProfileForm() {
|
|||||||
stopRecording: stopSystemRecording,
|
stopRecording: stopSystemRecording,
|
||||||
cancelRecording: cancelSystemRecording,
|
cancelRecording: cancelSystemRecording,
|
||||||
} = useSystemAudioCapture({
|
} = useSystemAudioCapture({
|
||||||
maxDurationSeconds: 30,
|
maxDurationSeconds: 29,
|
||||||
onRecordingComplete: (blob, recordedDuration) => {
|
onRecordingComplete: (blob, recordedDuration) => {
|
||||||
const file = new File([blob], `system-audio-${Date.now()}.wav`, {
|
const file = new File([blob], `system-audio-${Date.now()}.wav`, {
|
||||||
type: blob.type || 'audio/wav',
|
type: blob.type || 'audio/wav',
|
||||||
@@ -252,6 +256,20 @@ export function ProfileForm() {
|
|||||||
}
|
}
|
||||||
}, [systemRecordingError, toast]);
|
}, [systemRecordingError, toast]);
|
||||||
|
|
||||||
|
// Handle avatar preview
|
||||||
|
useEffect(() => {
|
||||||
|
if (selectedAvatarFile instanceof File) {
|
||||||
|
const url = URL.createObjectURL(selectedAvatarFile);
|
||||||
|
setAvatarPreview(url);
|
||||||
|
return () => URL.revokeObjectURL(url);
|
||||||
|
} else if (editingProfile?.avatar_path) {
|
||||||
|
setAvatarPreview(`${serverUrl}/profiles/${editingProfile.id}/avatar`);
|
||||||
|
} else {
|
||||||
|
setAvatarPreview(null);
|
||||||
|
}
|
||||||
|
}, [selectedAvatarFile, editingProfile, serverUrl]);
|
||||||
|
|
||||||
|
// Restore form state from draft or editing profile
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (editingProfile) {
|
if (editingProfile) {
|
||||||
form.reset({
|
form.reset({
|
||||||
@@ -260,18 +278,46 @@ export function ProfileForm() {
|
|||||||
language: editingProfile.language as LanguageCode,
|
language: editingProfile.language as LanguageCode,
|
||||||
sampleFile: undefined,
|
sampleFile: undefined,
|
||||||
referenceText: undefined,
|
referenceText: undefined,
|
||||||
|
avatarFile: undefined,
|
||||||
});
|
});
|
||||||
} else {
|
} else if (profileFormDraft && open) {
|
||||||
|
// Restore from draft when opening in create mode
|
||||||
|
form.reset({
|
||||||
|
name: profileFormDraft.name,
|
||||||
|
description: profileFormDraft.description,
|
||||||
|
language: profileFormDraft.language as LanguageCode,
|
||||||
|
referenceText: profileFormDraft.referenceText,
|
||||||
|
sampleFile: undefined,
|
||||||
|
avatarFile: undefined,
|
||||||
|
});
|
||||||
|
setSampleMode(profileFormDraft.sampleMode);
|
||||||
|
// Restore the file if we have it saved
|
||||||
|
if (
|
||||||
|
profileFormDraft.sampleFileData &&
|
||||||
|
profileFormDraft.sampleFileName &&
|
||||||
|
profileFormDraft.sampleFileType
|
||||||
|
) {
|
||||||
|
const file = base64ToFile(
|
||||||
|
profileFormDraft.sampleFileData,
|
||||||
|
profileFormDraft.sampleFileName,
|
||||||
|
profileFormDraft.sampleFileType,
|
||||||
|
);
|
||||||
|
form.setValue('sampleFile', file);
|
||||||
|
}
|
||||||
|
} else if (!open) {
|
||||||
|
// Only reset to defaults when modal is closed and no draft
|
||||||
form.reset({
|
form.reset({
|
||||||
name: '',
|
name: '',
|
||||||
description: '',
|
description: '',
|
||||||
language: 'en',
|
language: 'en',
|
||||||
sampleFile: undefined,
|
sampleFile: undefined,
|
||||||
referenceText: undefined,
|
referenceText: undefined,
|
||||||
|
avatarFile: undefined,
|
||||||
});
|
});
|
||||||
setSampleMode('upload');
|
setSampleMode('record');
|
||||||
|
setAvatarPreview(null);
|
||||||
}
|
}
|
||||||
}, [editingProfile, form]);
|
}, [editingProfile, profileFormDraft, open, form]);
|
||||||
|
|
||||||
async function handleTranscribe() {
|
async function handleTranscribe() {
|
||||||
const file = form.getValues('sampleFile');
|
const file = form.getValues('sampleFile');
|
||||||
@@ -313,6 +359,52 @@ export function ProfileForm() {
|
|||||||
playPause(file);
|
playPause(file);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function handleAvatarFileChange(e: React.ChangeEvent<HTMLInputElement>) {
|
||||||
|
const file = e.target.files?.[0];
|
||||||
|
if (file) {
|
||||||
|
if (!file.type.startsWith('image/')) {
|
||||||
|
toast({
|
||||||
|
title: 'Invalid file type',
|
||||||
|
description: 'Please select an image file (PNG, JPG, or WebP)',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (file.size > 5 * 1024 * 1024) {
|
||||||
|
toast({
|
||||||
|
title: 'File too large',
|
||||||
|
description: 'Image must be less than 5MB',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
form.setValue('avatarFile', file);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async function handleRemoveAvatar() {
|
||||||
|
if (editingProfileId && editingProfile?.avatar_path) {
|
||||||
|
try {
|
||||||
|
await deleteAvatar.mutateAsync(editingProfileId);
|
||||||
|
toast({
|
||||||
|
title: 'Avatar removed',
|
||||||
|
description: 'Avatar image has been removed successfully.',
|
||||||
|
});
|
||||||
|
} catch (error) {
|
||||||
|
toast({
|
||||||
|
title: 'Failed to remove avatar',
|
||||||
|
description: error instanceof Error ? error.message : 'Unknown error',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
form.setValue('avatarFile', undefined);
|
||||||
|
setAvatarPreview(null);
|
||||||
|
if (avatarInputRef.current) {
|
||||||
|
avatarInputRef.current.value = '';
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
async function onSubmit(data: ProfileFormValues) {
|
async function onSubmit(data: ProfileFormValues) {
|
||||||
try {
|
try {
|
||||||
if (editingProfileId) {
|
if (editingProfileId) {
|
||||||
@@ -325,6 +417,24 @@ export function ProfileForm() {
|
|||||||
language: data.language,
|
language: data.language,
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// Handle avatar upload/update if file changed
|
||||||
|
if (data.avatarFile) {
|
||||||
|
try {
|
||||||
|
await uploadAvatar.mutateAsync({
|
||||||
|
profileId: editingProfileId,
|
||||||
|
file: data.avatarFile,
|
||||||
|
});
|
||||||
|
} catch (avatarError) {
|
||||||
|
toast({
|
||||||
|
title: 'Avatar upload failed',
|
||||||
|
description:
|
||||||
|
avatarError instanceof Error ? avatarError.message : 'Failed to upload avatar',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
toast({
|
toast({
|
||||||
title: 'Voice updated',
|
title: 'Voice updated',
|
||||||
description: `"${data.name}" has been updated successfully.`,
|
description: `"${data.name}" has been updated successfully.`,
|
||||||
@@ -401,6 +511,24 @@ export function ProfileForm() {
|
|||||||
file: sampleFile,
|
file: sampleFile,
|
||||||
referenceText: referenceText,
|
referenceText: referenceText,
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// Handle avatar upload if provided
|
||||||
|
if (data.avatarFile) {
|
||||||
|
try {
|
||||||
|
await uploadAvatar.mutateAsync({
|
||||||
|
profileId: profile.id,
|
||||||
|
file: data.avatarFile,
|
||||||
|
});
|
||||||
|
} catch (avatarError) {
|
||||||
|
toast({
|
||||||
|
title: 'Avatar upload failed',
|
||||||
|
description:
|
||||||
|
avatarError instanceof Error ? avatarError.message : 'Failed to upload avatar',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
toast({
|
toast({
|
||||||
title: 'Profile created',
|
title: 'Profile created',
|
||||||
description: `"${data.name}" has been created with a sample.`,
|
description: `"${data.name}" has been created with a sample.`,
|
||||||
@@ -415,6 +543,8 @@ export function ProfileForm() {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Clear draft and reset form on success
|
||||||
|
setProfileFormDraft(null);
|
||||||
form.reset();
|
form.reset();
|
||||||
setEditingProfileId(null);
|
setEditingProfileId(null);
|
||||||
setOpen(false);
|
setOpen(false);
|
||||||
@@ -427,12 +557,41 @@ export function ProfileForm() {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
function handleOpenChange(open: boolean) {
|
async function handleOpenChange(newOpen: boolean) {
|
||||||
setOpen(open);
|
if (!newOpen && isCreating) {
|
||||||
if (!open) {
|
// Save draft when closing the create modal
|
||||||
|
const values = form.getValues();
|
||||||
|
const hasContent =
|
||||||
|
values.name || values.description || values.referenceText || values.sampleFile;
|
||||||
|
|
||||||
|
if (hasContent) {
|
||||||
|
const draft: ProfileFormDraft = {
|
||||||
|
name: values.name || '',
|
||||||
|
description: values.description || '',
|
||||||
|
language: values.language || 'en',
|
||||||
|
referenceText: values.referenceText || '',
|
||||||
|
sampleMode,
|
||||||
|
};
|
||||||
|
|
||||||
|
// Save file as base64 if present
|
||||||
|
if (values.sampleFile) {
|
||||||
|
try {
|
||||||
|
draft.sampleFileName = values.sampleFile.name;
|
||||||
|
draft.sampleFileType = values.sampleFile.type;
|
||||||
|
draft.sampleFileData = await fileToBase64(values.sampleFile);
|
||||||
|
} catch {
|
||||||
|
// If file conversion fails, just don't save the file
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
setProfileFormDraft(draft);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
setOpen(newOpen);
|
||||||
|
if (!newOpen) {
|
||||||
setEditingProfileId(null);
|
setEditingProfileId(null);
|
||||||
form.reset();
|
// Don't reset form here - let the effect handle it based on draft state
|
||||||
setSampleMode('upload');
|
|
||||||
if (isRecording) {
|
if (isRecording) {
|
||||||
cancelRecording();
|
cancelRecording();
|
||||||
}
|
}
|
||||||
@@ -445,174 +604,119 @@ export function ProfileForm() {
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<Dialog open={open} onOpenChange={handleOpenChange}>
|
<Dialog open={open} onOpenChange={handleOpenChange}>
|
||||||
<DialogContent className="max-w-4xl">
|
<DialogContent className="max-w-none w-screen h-screen left-0 top-0 translate-x-0 translate-y-0 rounded-none p-6 overflow-y-auto">
|
||||||
<DialogHeader>
|
<div className="max-w-5xl max-h-[85vh] mx-auto my-auto w-full flex flex-col">
|
||||||
<DialogTitle>{editingProfileId ? 'Edit Voice' : 'Create Voice Profile'}</DialogTitle>
|
<DialogHeader>
|
||||||
<DialogDescription>
|
<DialogTitle className="text-2xl">
|
||||||
{editingProfileId
|
{editingProfileId ? 'Edit Voice' : 'Clone voice'}
|
||||||
? 'Update your voice profile details and manage samples.'
|
</DialogTitle>
|
||||||
: 'Create a new voice profile with an audio sample to clone the voice.'}
|
<DialogDescription>
|
||||||
</DialogDescription>
|
{editingProfileId
|
||||||
</DialogHeader>
|
? 'Update your voice profile details and manage samples.'
|
||||||
|
: 'Create a new voice profile with an audio sample to clone the voice.'}
|
||||||
<Form {...form}>
|
</DialogDescription>
|
||||||
<form onSubmit={form.handleSubmit(onSubmit)}>
|
{isCreating && profileFormDraft && (
|
||||||
<div className="grid gap-6 grid-cols-2">
|
<div className="flex items-center gap-2 pt-2">
|
||||||
{/* Left column: Profile info */}
|
<span className="text-xs text-muted-foreground">Draft restored</span>
|
||||||
<div className="space-y-4">
|
<Button
|
||||||
<FormField
|
type="button"
|
||||||
control={form.control}
|
variant="ghost"
|
||||||
name="name"
|
size="sm"
|
||||||
render={({ field }) => (
|
className="h-6 px-2 text-xs text-muted-foreground"
|
||||||
<FormItem>
|
onClick={() => {
|
||||||
<FormLabel>Name</FormLabel>
|
setProfileFormDraft(null);
|
||||||
<FormControl>
|
form.reset({
|
||||||
<Input placeholder="My Voice" {...field} />
|
name: '',
|
||||||
</FormControl>
|
description: '',
|
||||||
<FormMessage />
|
language: 'en',
|
||||||
</FormItem>
|
sampleFile: undefined,
|
||||||
)}
|
referenceText: '',
|
||||||
/>
|
});
|
||||||
|
setSampleMode('record');
|
||||||
<FormField
|
}}
|
||||||
control={form.control}
|
>
|
||||||
name="description"
|
<X className="h-3 w-3 mr-1" />
|
||||||
render={({ field }) => (
|
Discard
|
||||||
<FormItem>
|
</Button>
|
||||||
<FormLabel>Description (Optional)</FormLabel>
|
|
||||||
<FormControl>
|
|
||||||
<Textarea placeholder="Describe this voice..." {...field} />
|
|
||||||
</FormControl>
|
|
||||||
<FormMessage />
|
|
||||||
</FormItem>
|
|
||||||
)}
|
|
||||||
/>
|
|
||||||
|
|
||||||
<FormField
|
|
||||||
control={form.control}
|
|
||||||
name="language"
|
|
||||||
render={({ field }) => (
|
|
||||||
<FormItem>
|
|
||||||
<FormLabel>Language</FormLabel>
|
|
||||||
<Select onValueChange={field.onChange} defaultValue={field.value}>
|
|
||||||
<FormControl>
|
|
||||||
<SelectTrigger>
|
|
||||||
<SelectValue />
|
|
||||||
</SelectTrigger>
|
|
||||||
</FormControl>
|
|
||||||
<SelectContent>
|
|
||||||
{LANGUAGE_OPTIONS.map((lang) => (
|
|
||||||
<SelectItem key={lang.value} value={lang.value}>
|
|
||||||
{lang.label}
|
|
||||||
</SelectItem>
|
|
||||||
))}
|
|
||||||
</SelectContent>
|
|
||||||
</Select>
|
|
||||||
<FormMessage />
|
|
||||||
</FormItem>
|
|
||||||
)}
|
|
||||||
/>
|
|
||||||
</div>
|
</div>
|
||||||
|
)}
|
||||||
|
</DialogHeader>
|
||||||
|
|
||||||
{/* Right column: Sample management */}
|
<Form {...form}>
|
||||||
<div className="space-y-4 border-l pl-6">
|
<form onSubmit={form.handleSubmit(onSubmit)} className="flex-1 min-h-0 flex flex-col">
|
||||||
{isCreating ? (
|
<div className="grid gap-6 grid-cols-2 flex-1 overflow-y-auto min-h-0">
|
||||||
<>
|
{/* Left column: Sample management */}
|
||||||
<div>
|
<div className="space-y-4 border-r pr-6">
|
||||||
<h3 className="text-sm font-medium mb-2">Add Sample</h3>
|
{isCreating ? (
|
||||||
<p className="text-sm text-muted-foreground mb-4">
|
<>
|
||||||
Provide an audio sample to clone the voice. You can add more samples later.
|
<Tabs
|
||||||
</p>
|
className="pt-4"
|
||||||
</div>
|
value={sampleMode}
|
||||||
|
onValueChange={(v) => {
|
||||||
<Tabs
|
const newMode = v as 'upload' | 'record' | 'system';
|
||||||
value={sampleMode}
|
// Cancel any active recordings when switching modes
|
||||||
onValueChange={(v) => {
|
if (isRecording && newMode !== 'record') {
|
||||||
const newMode = v as 'upload' | 'record' | 'system';
|
cancelRecording();
|
||||||
// Cancel any active recordings when switching modes
|
}
|
||||||
if (isRecording && newMode !== 'record') {
|
if (isSystemRecording && newMode !== 'system') {
|
||||||
cancelRecording();
|
cancelSystemRecording();
|
||||||
}
|
}
|
||||||
if (isSystemRecording && newMode !== 'system') {
|
setSampleMode(newMode);
|
||||||
cancelSystemRecording();
|
}}
|
||||||
}
|
|
||||||
setSampleMode(newMode);
|
|
||||||
}}
|
|
||||||
>
|
|
||||||
<TabsList
|
|
||||||
className={`grid w-full ${isTauri() && isSystemAudioSupported ? 'grid-cols-3' : 'grid-cols-2'}`}
|
|
||||||
>
|
>
|
||||||
<TabsTrigger value="upload" className="flex items-center gap-2">
|
<TabsList
|
||||||
<Upload className="h-4 w-4 shrink-0" />
|
className={`grid w-full ${platform.metadata.isTauri && isSystemAudioSupported ? 'grid-cols-3' : 'grid-cols-2'}`}
|
||||||
Upload
|
>
|
||||||
</TabsTrigger>
|
<TabsTrigger value="upload" className="flex items-center gap-2">
|
||||||
<TabsTrigger value="record" className="flex items-center gap-2">
|
<Upload className="h-4 w-4 shrink-0" />
|
||||||
<Mic className="h-4 w-4 shrink-0" />
|
Upload
|
||||||
Record
|
|
||||||
</TabsTrigger>
|
|
||||||
{isTauri() && isSystemAudioSupported && (
|
|
||||||
<TabsTrigger value="system" className="flex items-center gap-2">
|
|
||||||
<Monitor className="h-4 w-4 shrink-0" />
|
|
||||||
System Audio
|
|
||||||
</TabsTrigger>
|
</TabsTrigger>
|
||||||
)}
|
<TabsTrigger value="record" className="flex items-center gap-2">
|
||||||
</TabsList>
|
<Mic className="h-4 w-4 shrink-0" />
|
||||||
|
Record
|
||||||
<TabsContent value="upload" className="space-y-4">
|
</TabsTrigger>
|
||||||
<FormField
|
{platform.metadata.isTauri && isSystemAudioSupported && (
|
||||||
control={form.control}
|
<TabsTrigger value="system" className="flex items-center gap-2">
|
||||||
name="sampleFile"
|
<Monitor className="h-4 w-4 shrink-0" />
|
||||||
render={({ field: { onChange, name } }) => (
|
System Audio
|
||||||
<AudioSampleUpload
|
</TabsTrigger>
|
||||||
file={selectedFile}
|
|
||||||
onFileChange={onChange}
|
|
||||||
onTranscribe={handleTranscribe}
|
|
||||||
onPlayPause={handlePlayPause}
|
|
||||||
isPlaying={isPlaying}
|
|
||||||
isValidating={isValidatingAudio}
|
|
||||||
isTranscribing={transcribe.isPending}
|
|
||||||
isDisabled={
|
|
||||||
audioDuration !== null && audioDuration > MAX_AUDIO_DURATION_SECONDS
|
|
||||||
}
|
|
||||||
fieldName={name}
|
|
||||||
/>
|
|
||||||
)}
|
)}
|
||||||
/>
|
</TabsList>
|
||||||
</TabsContent>
|
|
||||||
|
|
||||||
<TabsContent value="record" className="space-y-4">
|
<TabsContent value="upload" className="space-y-4">
|
||||||
<FormField
|
<FormField
|
||||||
control={form.control}
|
control={form.control}
|
||||||
name="sampleFile"
|
name="sampleFile"
|
||||||
render={() => (
|
render={({ field: { onChange, name } }) => (
|
||||||
<AudioSampleRecording
|
<AudioSampleUpload
|
||||||
file={selectedFile}
|
file={selectedFile}
|
||||||
isRecording={isRecording}
|
onFileChange={onChange}
|
||||||
duration={duration}
|
onTranscribe={handleTranscribe}
|
||||||
onStart={startRecording}
|
onPlayPause={handlePlayPause}
|
||||||
onStop={stopRecording}
|
isPlaying={isPlaying}
|
||||||
onCancel={handleCancelRecording}
|
isValidating={isValidatingAudio}
|
||||||
onTranscribe={handleTranscribe}
|
isTranscribing={transcribe.isPending}
|
||||||
onPlayPause={handlePlayPause}
|
isDisabled={
|
||||||
isPlaying={isPlaying}
|
audioDuration !== null &&
|
||||||
isTranscribing={transcribe.isPending}
|
audioDuration > MAX_AUDIO_DURATION_SECONDS
|
||||||
/>
|
}
|
||||||
)}
|
fieldName={name}
|
||||||
/>
|
/>
|
||||||
</TabsContent>
|
)}
|
||||||
|
/>
|
||||||
|
</TabsContent>
|
||||||
|
|
||||||
{isTauri() && isSystemAudioSupported && (
|
<TabsContent value="record" className="space-y-4">
|
||||||
<TabsContent value="system" className="space-y-4">
|
|
||||||
<FormField
|
<FormField
|
||||||
control={form.control}
|
control={form.control}
|
||||||
name="sampleFile"
|
name="sampleFile"
|
||||||
render={() => (
|
render={() => (
|
||||||
<AudioSampleSystem
|
<AudioSampleRecording
|
||||||
file={selectedFile}
|
file={selectedFile}
|
||||||
isRecording={isSystemRecording}
|
isRecording={isRecording}
|
||||||
duration={systemDuration}
|
duration={duration}
|
||||||
onStart={startSystemRecording}
|
onStart={startRecording}
|
||||||
onStop={stopSystemRecording}
|
onStop={stopRecording}
|
||||||
onCancel={handleCancelRecording}
|
onCancel={handleCancelRecording}
|
||||||
onTranscribe={handleTranscribe}
|
onTranscribe={handleTranscribe}
|
||||||
onPlayPause={handlePlayPause}
|
onPlayPause={handlePlayPause}
|
||||||
@@ -622,55 +726,188 @@ export function ProfileForm() {
|
|||||||
)}
|
)}
|
||||||
/>
|
/>
|
||||||
</TabsContent>
|
</TabsContent>
|
||||||
)}
|
|
||||||
</Tabs>
|
|
||||||
|
|
||||||
<FormField
|
{platform.metadata.isTauri && isSystemAudioSupported && (
|
||||||
control={form.control}
|
<TabsContent value="system" className="space-y-4">
|
||||||
name="referenceText"
|
<FormField
|
||||||
render={({ field }) => (
|
control={form.control}
|
||||||
<FormItem>
|
name="sampleFile"
|
||||||
<FormLabel>Reference Text</FormLabel>
|
render={() => (
|
||||||
<FormControl>
|
<AudioSampleSystem
|
||||||
<Textarea
|
file={selectedFile}
|
||||||
placeholder="Enter the exact text spoken in the audio..."
|
isRecording={isSystemRecording}
|
||||||
className="min-h-[100px]"
|
duration={systemDuration}
|
||||||
{...field}
|
onStart={startSystemRecording}
|
||||||
|
onStop={stopSystemRecording}
|
||||||
|
onCancel={handleCancelRecording}
|
||||||
|
onTranscribe={handleTranscribe}
|
||||||
|
onPlayPause={handlePlayPause}
|
||||||
|
isPlaying={isPlaying}
|
||||||
|
isTranscribing={transcribe.isPending}
|
||||||
|
/>
|
||||||
|
)}
|
||||||
/>
|
/>
|
||||||
</FormControl>
|
</TabsContent>
|
||||||
<FormMessage />
|
)}
|
||||||
</FormItem>
|
</Tabs>
|
||||||
)}
|
|
||||||
/>
|
|
||||||
</>
|
|
||||||
) : (
|
|
||||||
// Show sample list when editing
|
|
||||||
editingProfileId && (
|
|
||||||
<div>
|
|
||||||
<SampleList profileId={editingProfileId} />
|
|
||||||
</div>
|
|
||||||
)
|
|
||||||
)}
|
|
||||||
</div>
|
|
||||||
</div>
|
|
||||||
|
|
||||||
<div className="flex gap-2 justify-end mt-6 pt-4 border-t">
|
<FormField
|
||||||
<Button type="button" variant="outline" onClick={() => handleOpenChange(false)}>
|
control={form.control}
|
||||||
Cancel
|
name="referenceText"
|
||||||
</Button>
|
render={({ field }) => (
|
||||||
<Button
|
<FormItem>
|
||||||
type="submit"
|
<FormLabel>Reference Text</FormLabel>
|
||||||
disabled={createProfile.isPending || updateProfile.isPending || addSample.isPending}
|
<FormControl>
|
||||||
>
|
<Textarea
|
||||||
{createProfile.isPending || updateProfile.isPending || addSample.isPending
|
placeholder="Enter the exact text spoken in the audio..."
|
||||||
? 'Saving...'
|
className="min-h-[100px]"
|
||||||
: editingProfileId
|
{...field}
|
||||||
? 'Save Changes'
|
/>
|
||||||
: 'Create Profile'}
|
</FormControl>
|
||||||
</Button>
|
<FormMessage />
|
||||||
</div>
|
</FormItem>
|
||||||
</form>
|
)}
|
||||||
</Form>
|
/>
|
||||||
|
</>
|
||||||
|
) : (
|
||||||
|
// Show sample list when editing
|
||||||
|
editingProfileId && (
|
||||||
|
<div>
|
||||||
|
<SampleList profileId={editingProfileId} />
|
||||||
|
</div>
|
||||||
|
)
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Right column: Profile info */}
|
||||||
|
<div className="space-y-4">
|
||||||
|
{/* Avatar Upload */}
|
||||||
|
<FormField
|
||||||
|
control={form.control}
|
||||||
|
name="avatarFile"
|
||||||
|
render={() => (
|
||||||
|
<FormItem>
|
||||||
|
<FormControl>
|
||||||
|
<div className="flex justify-center pt-4 pb-2">
|
||||||
|
<div className="relative group">
|
||||||
|
<div className="h-24 w-24 rounded-full bg-muted flex items-center justify-center shrink-0 overflow-hidden border-2 border-border">
|
||||||
|
{avatarPreview ? (
|
||||||
|
<img
|
||||||
|
src={avatarPreview}
|
||||||
|
alt="Avatar preview"
|
||||||
|
className="h-full w-full object-cover"
|
||||||
|
/>
|
||||||
|
) : (
|
||||||
|
<Mic className="h-10 w-10 text-muted-foreground" />
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
onClick={() => avatarInputRef.current?.click()}
|
||||||
|
className="absolute inset-0 rounded-full bg-accent/60 opacity-0 group-hover:opacity-100 transition-opacity flex items-center justify-center cursor-pointer"
|
||||||
|
>
|
||||||
|
<Edit2 className="h-6 w-6 text-accent-foreground" />
|
||||||
|
</button>
|
||||||
|
{(avatarPreview || editingProfile?.avatar_path) && (
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
onClick={handleRemoveAvatar}
|
||||||
|
disabled={deleteAvatar.isPending}
|
||||||
|
className="absolute bottom-0 right-0 h-6 w-6 rounded-full bg-background/60 backdrop-blur-sm text-muted-foreground flex items-center justify-center hover:bg-background/80 hover:text-foreground transition-colors shadow-sm border border-border/50"
|
||||||
|
>
|
||||||
|
<X className="h-3.5 w-3.5" />
|
||||||
|
</button>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
<input
|
||||||
|
ref={avatarInputRef}
|
||||||
|
type="file"
|
||||||
|
accept="image/png,image/jpeg,image/webp"
|
||||||
|
onChange={handleAvatarFileChange}
|
||||||
|
className="hidden"
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
</FormControl>
|
||||||
|
<FormMessage />
|
||||||
|
</FormItem>
|
||||||
|
)}
|
||||||
|
/>
|
||||||
|
|
||||||
|
<FormField
|
||||||
|
control={form.control}
|
||||||
|
name="name"
|
||||||
|
render={({ field }) => (
|
||||||
|
<FormItem>
|
||||||
|
<FormLabel>Name</FormLabel>
|
||||||
|
<FormControl>
|
||||||
|
<Input placeholder="My Voice" {...field} />
|
||||||
|
</FormControl>
|
||||||
|
<FormMessage />
|
||||||
|
</FormItem>
|
||||||
|
)}
|
||||||
|
/>
|
||||||
|
|
||||||
|
<FormField
|
||||||
|
control={form.control}
|
||||||
|
name="description"
|
||||||
|
render={({ field }) => (
|
||||||
|
<FormItem>
|
||||||
|
<FormLabel>Description (Optional)</FormLabel>
|
||||||
|
<FormControl>
|
||||||
|
<Textarea placeholder="Describe this voice..." {...field} />
|
||||||
|
</FormControl>
|
||||||
|
<FormMessage />
|
||||||
|
</FormItem>
|
||||||
|
)}
|
||||||
|
/>
|
||||||
|
|
||||||
|
<FormField
|
||||||
|
control={form.control}
|
||||||
|
name="language"
|
||||||
|
render={({ field }) => (
|
||||||
|
<FormItem>
|
||||||
|
<FormLabel>Language</FormLabel>
|
||||||
|
<Select onValueChange={field.onChange} defaultValue={field.value}>
|
||||||
|
<FormControl>
|
||||||
|
<SelectTrigger>
|
||||||
|
<SelectValue />
|
||||||
|
</SelectTrigger>
|
||||||
|
</FormControl>
|
||||||
|
<SelectContent>
|
||||||
|
{LANGUAGE_OPTIONS.map((lang) => (
|
||||||
|
<SelectItem key={lang.value} value={lang.value}>
|
||||||
|
{lang.label}
|
||||||
|
</SelectItem>
|
||||||
|
))}
|
||||||
|
</SelectContent>
|
||||||
|
</Select>
|
||||||
|
<FormMessage />
|
||||||
|
</FormItem>
|
||||||
|
)}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div className="flex gap-2 justify-end mt-6 pt-4 border-t">
|
||||||
|
<Button type="button" variant="outline" onClick={() => handleOpenChange(false)}>
|
||||||
|
Cancel
|
||||||
|
</Button>
|
||||||
|
<Button
|
||||||
|
type="submit"
|
||||||
|
disabled={
|
||||||
|
createProfile.isPending || updateProfile.isPending || addSample.isPending
|
||||||
|
}
|
||||||
|
>
|
||||||
|
{createProfile.isPending || updateProfile.isPending || addSample.isPending
|
||||||
|
? 'Saving...'
|
||||||
|
: editingProfileId
|
||||||
|
? 'Save Changes'
|
||||||
|
: 'Create Profile'}
|
||||||
|
</Button>
|
||||||
|
</div>
|
||||||
|
</form>
|
||||||
|
</Form>
|
||||||
|
</div>
|
||||||
</DialogContent>
|
</DialogContent>
|
||||||
</Dialog>
|
</Dialog>
|
||||||
);
|
);
|
||||||
|
|||||||
@@ -1,11 +1,141 @@
|
|||||||
import { Plus, Trash2, Play } from 'lucide-react';
|
import { Check, Edit, Pause, Play, Plus, Trash2, Volume2, X } from 'lucide-react';
|
||||||
import { useState } from 'react';
|
import { useEffect, useRef, useState } from 'react';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
import { useDeleteSample, useProfileSamples } from '@/lib/hooks/useProfiles';
|
import { CircleButton } from '@/components/ui/circle-button';
|
||||||
import { usePlayerStore } from '@/stores/playerStore';
|
import {
|
||||||
|
Dialog,
|
||||||
|
DialogContent,
|
||||||
|
DialogDescription,
|
||||||
|
DialogFooter,
|
||||||
|
DialogHeader,
|
||||||
|
DialogTitle,
|
||||||
|
} from '@/components/ui/dialog';
|
||||||
|
import { Slider } from '@/components/ui/slider';
|
||||||
|
import { Textarea } from '@/components/ui/textarea';
|
||||||
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { apiClient } from '@/lib/api/client';
|
||||||
|
import { useDeleteSample, useProfileSamples, useUpdateSample } from '@/lib/hooks/useProfiles';
|
||||||
|
import { formatAudioDuration } from '@/lib/utils/audio';
|
||||||
|
import { cn } from '@/lib/utils/cn';
|
||||||
import { SampleUpload } from './SampleUpload';
|
import { SampleUpload } from './SampleUpload';
|
||||||
|
|
||||||
|
interface MiniSamplePlayerProps {
|
||||||
|
audioUrl: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
function MiniSamplePlayer({ audioUrl }: MiniSamplePlayerProps) {
|
||||||
|
const audioRef = useRef<HTMLAudioElement | null>(null);
|
||||||
|
const [isPlaying, setIsPlaying] = useState(false);
|
||||||
|
const [currentTime, setCurrentTime] = useState(0);
|
||||||
|
const [duration, setDuration] = useState(0);
|
||||||
|
const [isLoading, setIsLoading] = useState(true);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
const audio = new Audio(audioUrl);
|
||||||
|
audioRef.current = audio;
|
||||||
|
|
||||||
|
const handleLoadedMetadata = () => {
|
||||||
|
setDuration(audio.duration);
|
||||||
|
setIsLoading(false);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleTimeUpdate = () => {
|
||||||
|
setCurrentTime(audio.currentTime);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleEnded = () => {
|
||||||
|
setIsPlaying(false);
|
||||||
|
setCurrentTime(0);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handlePlay = () => setIsPlaying(true);
|
||||||
|
const handlePause = () => setIsPlaying(false);
|
||||||
|
|
||||||
|
audio.addEventListener('loadedmetadata', handleLoadedMetadata);
|
||||||
|
audio.addEventListener('timeupdate', handleTimeUpdate);
|
||||||
|
audio.addEventListener('ended', handleEnded);
|
||||||
|
audio.addEventListener('play', handlePlay);
|
||||||
|
audio.addEventListener('pause', handlePause);
|
||||||
|
|
||||||
|
return () => {
|
||||||
|
audio.pause();
|
||||||
|
audio.removeEventListener('loadedmetadata', handleLoadedMetadata);
|
||||||
|
audio.removeEventListener('timeupdate', handleTimeUpdate);
|
||||||
|
audio.removeEventListener('ended', handleEnded);
|
||||||
|
audio.removeEventListener('play', handlePlay);
|
||||||
|
audio.removeEventListener('pause', handlePause);
|
||||||
|
audio.src = '';
|
||||||
|
};
|
||||||
|
}, [audioUrl]);
|
||||||
|
|
||||||
|
const handlePlayPause = () => {
|
||||||
|
if (!audioRef.current) return;
|
||||||
|
if (isPlaying) {
|
||||||
|
audioRef.current.pause();
|
||||||
|
} else {
|
||||||
|
audioRef.current.play();
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleSeek = (value: number[]) => {
|
||||||
|
if (!audioRef.current || duration === 0) return;
|
||||||
|
const progress = value[0] / 100;
|
||||||
|
audioRef.current.currentTime = progress * duration;
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleStop = () => {
|
||||||
|
if (audioRef.current) {
|
||||||
|
audioRef.current.pause();
|
||||||
|
audioRef.current.currentTime = 0;
|
||||||
|
}
|
||||||
|
setIsPlaying(false);
|
||||||
|
setCurrentTime(0);
|
||||||
|
};
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="border-t bg-muted/30 px-3 py-2 mt-2">
|
||||||
|
<div className="flex items-center gap-2">
|
||||||
|
<Button
|
||||||
|
type="button"
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-7 w-7 shrink-0"
|
||||||
|
onClick={handlePlayPause}
|
||||||
|
disabled={isLoading}
|
||||||
|
>
|
||||||
|
{isPlaying ? <Pause className="h-3.5 w-3.5" /> : <Play className="h-3.5 w-3.5 ml-0.5" />}
|
||||||
|
</Button>
|
||||||
|
|
||||||
|
<div className="flex-1 min-w-0 flex items-center gap-2">
|
||||||
|
<Slider
|
||||||
|
value={duration > 0 ? [(currentTime / duration) * 100] : [0]}
|
||||||
|
onValueChange={handleSeek}
|
||||||
|
max={100}
|
||||||
|
step={0.1}
|
||||||
|
className="flex-1"
|
||||||
|
/>
|
||||||
|
<div className="flex items-center gap-1 text-xs text-muted-foreground shrink-0 min-w-[70px]">
|
||||||
|
<span className="font-mono">{formatAudioDuration(currentTime)}</span>
|
||||||
|
<span>/</span>
|
||||||
|
<span className="font-mono">{formatAudioDuration(duration)}</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<Button
|
||||||
|
type="button"
|
||||||
|
variant="ghost"
|
||||||
|
size="icon"
|
||||||
|
className="h-7 w-7 shrink-0"
|
||||||
|
onClick={handleStop}
|
||||||
|
title="Stop"
|
||||||
|
>
|
||||||
|
<X className="h-3.5 w-3.5" />
|
||||||
|
</Button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
interface SampleListProps {
|
interface SampleListProps {
|
||||||
profileId: string;
|
profileId: string;
|
||||||
}
|
}
|
||||||
@@ -13,20 +143,62 @@ interface SampleListProps {
|
|||||||
export function SampleList({ profileId }: SampleListProps) {
|
export function SampleList({ profileId }: SampleListProps) {
|
||||||
const { data: samples, isLoading } = useProfileSamples(profileId);
|
const { data: samples, isLoading } = useProfileSamples(profileId);
|
||||||
const deleteSample = useDeleteSample();
|
const deleteSample = useDeleteSample();
|
||||||
|
const updateSample = useUpdateSample();
|
||||||
|
const { toast } = useToast();
|
||||||
const [uploadOpen, setUploadOpen] = useState(false);
|
const [uploadOpen, setUploadOpen] = useState(false);
|
||||||
const setAudio = usePlayerStore((state) => state.setAudio);
|
const [editingSampleId, setEditingSampleId] = useState<string | null>(null);
|
||||||
const currentAudioId = usePlayerStore((state) => state.audioId);
|
const [editedText, setEditedText] = useState<string>('');
|
||||||
const isPlaying = usePlayerStore((state) => state.isPlaying);
|
const [deleteDialogOpen, setDeleteDialogOpen] = useState(false);
|
||||||
|
const [sampleToDelete, setSampleToDelete] = useState<string | null>(null);
|
||||||
|
|
||||||
const handleDelete = (sampleId: string) => {
|
const handleDeleteClick = (sampleId: string) => {
|
||||||
if (confirm('Are you sure you want to delete this sample?')) {
|
setSampleToDelete(sampleId);
|
||||||
deleteSample.mutate(sampleId);
|
setDeleteDialogOpen(true);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleDeleteConfirm = () => {
|
||||||
|
if (sampleToDelete) {
|
||||||
|
deleteSample.mutate(sampleToDelete);
|
||||||
|
setDeleteDialogOpen(false);
|
||||||
|
setSampleToDelete(null);
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
const handlePlay = (referenceText: string, sampleId: string) => {
|
const handleStartEdit = (sampleId: string, currentText: string) => {
|
||||||
const audioUrl = apiClient.getSampleUrl(sampleId);
|
setEditingSampleId(sampleId);
|
||||||
setAudio(audioUrl, sampleId, referenceText.substring(0, 50));
|
setEditedText(currentText);
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleCancelEdit = () => {
|
||||||
|
setEditingSampleId(null);
|
||||||
|
setEditedText('');
|
||||||
|
};
|
||||||
|
|
||||||
|
const handleSaveEdit = async (sampleId: string) => {
|
||||||
|
if (!editedText.trim()) {
|
||||||
|
toast({
|
||||||
|
title: 'Invalid text',
|
||||||
|
description: 'Reference text cannot be empty.',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
await updateSample.mutateAsync({ sampleId, referenceText: editedText.trim() });
|
||||||
|
toast({
|
||||||
|
title: 'Sample updated',
|
||||||
|
description: 'Reference text has been updated successfully.',
|
||||||
|
});
|
||||||
|
setEditingSampleId(null);
|
||||||
|
setEditedText('');
|
||||||
|
} catch (error) {
|
||||||
|
toast({
|
||||||
|
title: 'Update failed',
|
||||||
|
description: error instanceof Error ? error.message : 'Failed to update sample',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
if (isLoading) {
|
if (isLoading) {
|
||||||
@@ -34,57 +206,152 @@ export function SampleList({ profileId }: SampleListProps) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
return (
|
return (
|
||||||
<div className="space-y-4">
|
<div className="space-y-4 pt-4">
|
||||||
<div className="flex items-center justify-between">
|
|
||||||
<h3 className="text-lg font-semibold">Audio Samples</h3>
|
|
||||||
<Button type="button" size="sm" onClick={() => setUploadOpen(true)}>
|
|
||||||
<Plus className="mr-2 h-4 w-4" />
|
|
||||||
Add Sample
|
|
||||||
</Button>
|
|
||||||
</div>
|
|
||||||
|
|
||||||
{samples && samples.length === 0 ? (
|
{samples && samples.length === 0 ? (
|
||||||
<div className="text-sm text-muted-foreground py-4">
|
<div className="flex flex-col items-center justify-center py-8 text-center border border-dashed rounded-lg">
|
||||||
No samples yet. Add your first audio sample.
|
<Volume2 className="h-8 w-8 text-muted-foreground/50 mb-2" />
|
||||||
|
<p className="text-sm text-muted-foreground">No samples yet</p>
|
||||||
|
<p className="text-xs text-muted-foreground/70 mt-1">
|
||||||
|
Add your first audio sample to get started
|
||||||
|
</p>
|
||||||
</div>
|
</div>
|
||||||
) : (
|
) : (
|
||||||
<div className="space-y-2">
|
<div className="space-y-2">
|
||||||
{samples?.map((sample) => (
|
{samples?.map((sample, index) => {
|
||||||
<div
|
const isEditing = editingSampleId === sample.id;
|
||||||
key={sample.id}
|
|
||||||
className="flex items-center justify-between p-3 border rounded-lg"
|
return (
|
||||||
>
|
<div
|
||||||
<div className="flex-1">
|
key={sample.id}
|
||||||
<p className="text-sm font-medium">{sample.reference_text}</p>
|
className={cn(
|
||||||
<p className="text-xs text-muted-foreground mt-1">{sample.audio_path}</p>
|
'group relative rounded-lg border bg-card transition-all duration-200',
|
||||||
|
isEditing ? 'ring-2 ring-primary/20' : 'hover:border-primary/30',
|
||||||
|
)}
|
||||||
|
>
|
||||||
|
{isEditing ? (
|
||||||
|
/* Edit Mode */
|
||||||
|
<div className="p-4 space-y-3">
|
||||||
|
<div className="flex items-center gap-2 text-xs text-muted-foreground mb-2">
|
||||||
|
<Edit className="h-3 w-3" />
|
||||||
|
<span>Editing transcription</span>
|
||||||
|
</div>
|
||||||
|
<Textarea
|
||||||
|
value={editedText}
|
||||||
|
onChange={(e) => setEditedText(e.target.value)}
|
||||||
|
className="min-h-[100px] text-sm resize-none"
|
||||||
|
placeholder="Enter reference text..."
|
||||||
|
autoFocus
|
||||||
|
/>
|
||||||
|
<div className="flex items-center justify-end gap-2 pt-1">
|
||||||
|
<Button
|
||||||
|
type="button"
|
||||||
|
size="sm"
|
||||||
|
variant="ghost"
|
||||||
|
onClick={handleCancelEdit}
|
||||||
|
disabled={updateSample.isPending}
|
||||||
|
>
|
||||||
|
<X className="h-4 w-4 mr-1" />
|
||||||
|
Cancel
|
||||||
|
</Button>
|
||||||
|
<Button
|
||||||
|
type="button"
|
||||||
|
size="sm"
|
||||||
|
onClick={() => handleSaveEdit(sample.id)}
|
||||||
|
disabled={updateSample.isPending}
|
||||||
|
>
|
||||||
|
<Check className="h-4 w-4 mr-1" />
|
||||||
|
{updateSample.isPending ? 'Saving...' : 'Save'}
|
||||||
|
</Button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
) : (
|
||||||
|
<>
|
||||||
|
{/* View Mode */}
|
||||||
|
<div className="flex items-center gap-3 p-3 h-[72px]">
|
||||||
|
{/* Text Content */}
|
||||||
|
<div className="flex-1 min-w-0 py-0.5">
|
||||||
|
<p className="text-sm font-medium line-clamp-2 leading-snug">
|
||||||
|
{sample.reference_text}
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Action Buttons */}
|
||||||
|
<div className="shrink-0 flex items-center gap-0.5 opacity-0 group-hover:opacity-100 transition-opacity">
|
||||||
|
<CircleButton
|
||||||
|
icon={Edit}
|
||||||
|
title="Edit transcription"
|
||||||
|
onClick={() => handleStartEdit(sample.id, sample.reference_text)}
|
||||||
|
/>
|
||||||
|
<CircleButton
|
||||||
|
icon={Trash2}
|
||||||
|
title="Delete sample"
|
||||||
|
onClick={() => handleDeleteClick(sample.id)}
|
||||||
|
disabled={deleteSample.isPending}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Sample Number Badge */}
|
||||||
|
<div className="absolute top-1 right-2 text-[10px] text-muted-foreground/50 font-medium">
|
||||||
|
#{index + 1}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Mini Player - Always visible */}
|
||||||
|
<MiniSamplePlayer audioUrl={apiClient.getSampleUrl(sample.id)} />
|
||||||
|
</>
|
||||||
|
)}
|
||||||
</div>
|
</div>
|
||||||
<div className="flex gap-2">
|
);
|
||||||
<Button
|
})}
|
||||||
type="button"
|
|
||||||
variant="ghost"
|
|
||||||
size="sm"
|
|
||||||
onClick={() => handlePlay(sample.reference_text, sample.id)}
|
|
||||||
className={currentAudioId === sample.id && isPlaying ? 'text-primary' : ''}
|
|
||||||
>
|
|
||||||
<Play className="h-4 w-4 mr-1" />
|
|
||||||
Play
|
|
||||||
</Button>
|
|
||||||
<Button
|
|
||||||
type="button"
|
|
||||||
variant="ghost"
|
|
||||||
size="sm"
|
|
||||||
onClick={() => handleDelete(sample.id)}
|
|
||||||
disabled={deleteSample.isPending}
|
|
||||||
>
|
|
||||||
<Trash2 className="h-4 w-4 text-destructive" />
|
|
||||||
</Button>
|
|
||||||
</div>
|
|
||||||
</div>
|
|
||||||
))}
|
|
||||||
</div>
|
</div>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
|
<Button
|
||||||
|
type="button"
|
||||||
|
variant="outline"
|
||||||
|
className="w-full"
|
||||||
|
onClick={() => setUploadOpen(true)}
|
||||||
|
>
|
||||||
|
<Plus className="mr-2 h-4 w-4" />
|
||||||
|
Add Sample
|
||||||
|
</Button>
|
||||||
|
|
||||||
|
<p className="text-xs text-muted-foreground text-center px-2">
|
||||||
|
Note: A single 30-second sample is the sweet spot. Quality may decrease with multiple
|
||||||
|
samples. In a future update samples might be interchangeable and tagged for varying styles
|
||||||
|
of the same voice.
|
||||||
|
</p>
|
||||||
|
|
||||||
<SampleUpload profileId={profileId} open={uploadOpen} onOpenChange={setUploadOpen} />
|
<SampleUpload profileId={profileId} open={uploadOpen} onOpenChange={setUploadOpen} />
|
||||||
|
|
||||||
|
<Dialog open={deleteDialogOpen} onOpenChange={setDeleteDialogOpen}>
|
||||||
|
<DialogContent>
|
||||||
|
<DialogHeader>
|
||||||
|
<DialogTitle>Delete Sample</DialogTitle>
|
||||||
|
<DialogDescription>
|
||||||
|
Are you sure you want to delete this audio sample? This action cannot be undone.
|
||||||
|
</DialogDescription>
|
||||||
|
</DialogHeader>
|
||||||
|
<DialogFooter>
|
||||||
|
<Button
|
||||||
|
variant="outline"
|
||||||
|
onClick={() => {
|
||||||
|
setDeleteDialogOpen(false);
|
||||||
|
setSampleToDelete(null);
|
||||||
|
}}
|
||||||
|
>
|
||||||
|
Cancel
|
||||||
|
</Button>
|
||||||
|
<Button
|
||||||
|
variant="destructive"
|
||||||
|
onClick={handleDeleteConfirm}
|
||||||
|
disabled={deleteSample.isPending}
|
||||||
|
>
|
||||||
|
{deleteSample.isPending ? 'Deleting...' : 'Delete'}
|
||||||
|
</Button>
|
||||||
|
</DialogFooter>
|
||||||
|
</DialogContent>
|
||||||
|
</Dialog>
|
||||||
</div>
|
</div>
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
import { zodResolver } from '@hookform/resolvers/zod';
|
import { zodResolver } from '@hookform/resolvers/zod';
|
||||||
import { Mic, Monitor, Upload } from 'lucide-react';
|
import { Mic, Monitor, Upload } from 'lucide-react';
|
||||||
import { useState, useEffect } from 'react';
|
import { useEffect, useState } from 'react';
|
||||||
import { useForm } from 'react-hook-form';
|
import { useForm } from 'react-hook-form';
|
||||||
import * as z from 'zod';
|
import * as z from 'zod';
|
||||||
import { Button } from '@/components/ui/button';
|
import { Button } from '@/components/ui/button';
|
||||||
@@ -27,7 +27,7 @@ import { useAudioRecording } from '@/lib/hooks/useAudioRecording';
|
|||||||
import { useAddSample, useProfile } from '@/lib/hooks/useProfiles';
|
import { useAddSample, useProfile } from '@/lib/hooks/useProfiles';
|
||||||
import { useSystemAudioCapture } from '@/lib/hooks/useSystemAudioCapture';
|
import { useSystemAudioCapture } from '@/lib/hooks/useSystemAudioCapture';
|
||||||
import { useTranscription } from '@/lib/hooks/useTranscription';
|
import { useTranscription } from '@/lib/hooks/useTranscription';
|
||||||
import { isTauri } from '@/lib/tauri';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
import { AudioSampleRecording } from './AudioSampleRecording';
|
import { AudioSampleRecording } from './AudioSampleRecording';
|
||||||
import { AudioSampleSystem } from './AudioSampleSystem';
|
import { AudioSampleSystem } from './AudioSampleSystem';
|
||||||
import { AudioSampleUpload } from './AudioSampleUpload';
|
import { AudioSampleUpload } from './AudioSampleUpload';
|
||||||
@@ -49,6 +49,7 @@ interface SampleUploadProps {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export function SampleUpload({ profileId, open, onOpenChange }: SampleUploadProps) {
|
export function SampleUpload({ profileId, open, onOpenChange }: SampleUploadProps) {
|
||||||
|
const platform = usePlatform();
|
||||||
const addSample = useAddSample();
|
const addSample = useAddSample();
|
||||||
const transcribe = useTranscription();
|
const transcribe = useTranscription();
|
||||||
const { data: profile } = useProfile(profileId);
|
const { data: profile } = useProfile(profileId);
|
||||||
@@ -73,7 +74,7 @@ export function SampleUpload({ profileId, open, onOpenChange }: SampleUploadProp
|
|||||||
stopRecording,
|
stopRecording,
|
||||||
cancelRecording,
|
cancelRecording,
|
||||||
} = useAudioRecording({
|
} = useAudioRecording({
|
||||||
maxDurationSeconds: 30,
|
maxDurationSeconds: 29,
|
||||||
onRecordingComplete: (blob, recordedDuration) => {
|
onRecordingComplete: (blob, recordedDuration) => {
|
||||||
// Convert blob to File object
|
// Convert blob to File object
|
||||||
const file = new File([blob], `recording-${Date.now()}.webm`, {
|
const file = new File([blob], `recording-${Date.now()}.webm`, {
|
||||||
@@ -100,7 +101,7 @@ export function SampleUpload({ profileId, open, onOpenChange }: SampleUploadProp
|
|||||||
stopRecording: stopSystemRecording,
|
stopRecording: stopSystemRecording,
|
||||||
cancelRecording: cancelSystemRecording,
|
cancelRecording: cancelSystemRecording,
|
||||||
} = useSystemAudioCapture({
|
} = useSystemAudioCapture({
|
||||||
maxDurationSeconds: 30,
|
maxDurationSeconds: 29,
|
||||||
onRecordingComplete: (blob, recordedDuration) => {
|
onRecordingComplete: (blob, recordedDuration) => {
|
||||||
// Convert blob to File object
|
// Convert blob to File object
|
||||||
const file = new File([blob], `system-audio-${Date.now()}.wav`, {
|
const file = new File([blob], `system-audio-${Date.now()}.wav`, {
|
||||||
@@ -232,7 +233,7 @@ export function SampleUpload({ profileId, open, onOpenChange }: SampleUploadProp
|
|||||||
<form onSubmit={form.handleSubmit(onSubmit)} className="space-y-4">
|
<form onSubmit={form.handleSubmit(onSubmit)} className="space-y-4">
|
||||||
<Tabs value={mode} onValueChange={(v) => setMode(v as 'upload' | 'record' | 'system')}>
|
<Tabs value={mode} onValueChange={(v) => setMode(v as 'upload' | 'record' | 'system')}>
|
||||||
<TabsList
|
<TabsList
|
||||||
className={`grid w-full ${isTauri() && isSystemAudioSupported ? 'grid-cols-3' : 'grid-cols-2'}`}
|
className={`grid w-full ${platform.metadata.isTauri && isSystemAudioSupported ? 'grid-cols-3' : 'grid-cols-2'}`}
|
||||||
>
|
>
|
||||||
<TabsTrigger value="upload" className="flex items-center gap-2">
|
<TabsTrigger value="upload" className="flex items-center gap-2">
|
||||||
<Upload className="h-4 w-4 shrink-0" />
|
<Upload className="h-4 w-4 shrink-0" />
|
||||||
@@ -242,7 +243,7 @@ export function SampleUpload({ profileId, open, onOpenChange }: SampleUploadProp
|
|||||||
<Mic className="h-4 w-4 shrink-0" />
|
<Mic className="h-4 w-4 shrink-0" />
|
||||||
Record
|
Record
|
||||||
</TabsTrigger>
|
</TabsTrigger>
|
||||||
{isTauri() && isSystemAudioSupported && (
|
{platform.metadata.isTauri && isSystemAudioSupported && (
|
||||||
<TabsTrigger value="system" className="flex items-center gap-2">
|
<TabsTrigger value="system" className="flex items-center gap-2">
|
||||||
<Monitor className="h-4 w-4 shrink-0" />
|
<Monitor className="h-4 w-4 shrink-0" />
|
||||||
System Audio
|
System Audio
|
||||||
@@ -289,7 +290,7 @@ export function SampleUpload({ profileId, open, onOpenChange }: SampleUploadProp
|
|||||||
/>
|
/>
|
||||||
</TabsContent>
|
</TabsContent>
|
||||||
|
|
||||||
{isTauri() && isSystemAudioSupported && (
|
{platform.metadata.isTauri && isSystemAudioSupported && (
|
||||||
<TabsContent value="system" className="space-y-4">
|
<TabsContent value="system" className="space-y-4">
|
||||||
<FormField
|
<FormField
|
||||||
control={form.control}
|
control={form.control}
|
||||||
|
|||||||
@@ -6,10 +6,11 @@ export interface CircleButtonProps extends React.ButtonHTMLAttributes<HTMLButton
|
|||||||
}
|
}
|
||||||
|
|
||||||
const CircleButton = React.forwardRef<HTMLButtonElement, CircleButtonProps>(
|
const CircleButton = React.forwardRef<HTMLButtonElement, CircleButtonProps>(
|
||||||
({ className, icon: Icon, ...props }, ref) => {
|
({ className, icon: Icon, type = 'button', ...props }, ref) => {
|
||||||
return (
|
return (
|
||||||
<button
|
<button
|
||||||
ref={ref}
|
ref={ref}
|
||||||
|
type={type}
|
||||||
className={cn(
|
className={cn(
|
||||||
'h-7 w-7 rounded-full flex items-center justify-center flex-shrink-0',
|
'h-7 w-7 rounded-full flex items-center justify-center flex-shrink-0',
|
||||||
'hover:bg-muted transition-colors',
|
'hover:bg-muted transition-colors',
|
||||||
|
|||||||
@@ -0,0 +1,28 @@
|
|||||||
|
import * as PopoverPrimitive from '@radix-ui/react-popover';
|
||||||
|
import * as React from 'react';
|
||||||
|
import { cn } from '@/lib/utils/cn';
|
||||||
|
|
||||||
|
const Popover = PopoverPrimitive.Root;
|
||||||
|
|
||||||
|
const PopoverTrigger = PopoverPrimitive.Trigger;
|
||||||
|
|
||||||
|
const PopoverContent = React.forwardRef<
|
||||||
|
React.ElementRef<typeof PopoverPrimitive.Content>,
|
||||||
|
React.ComponentPropsWithoutRef<typeof PopoverPrimitive.Content>
|
||||||
|
>(({ className, align = 'center', sideOffset = 4, ...props }, ref) => (
|
||||||
|
<PopoverPrimitive.Portal>
|
||||||
|
<PopoverPrimitive.Content
|
||||||
|
ref={ref}
|
||||||
|
align={align}
|
||||||
|
sideOffset={sideOffset}
|
||||||
|
className={cn(
|
||||||
|
'z-50 w-72 rounded-md border bg-popover p-4 text-popover-foreground shadow-md outline-none data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0 data-[state=closed]:zoom-out-95 data-[state=open]:zoom-in-95 data-[side=bottom]:slide-in-from-top-2 data-[side=left]:slide-in-from-right-2 data-[side=right]:slide-in-from-left-2 data-[side=top]:slide-in-from-bottom-2',
|
||||||
|
className,
|
||||||
|
)}
|
||||||
|
{...props}
|
||||||
|
/>
|
||||||
|
</PopoverPrimitive.Portal>
|
||||||
|
));
|
||||||
|
PopoverContent.displayName = PopoverPrimitive.Content.displayName;
|
||||||
|
|
||||||
|
export { Popover, PopoverTrigger, PopoverContent };
|
||||||
Vendored
+3
@@ -0,0 +1,3 @@
|
|||||||
|
interface Window {
|
||||||
|
__voiceboxServerStartedByApp?: boolean;
|
||||||
|
}
|
||||||
+25
-156
@@ -1,172 +1,41 @@
|
|||||||
import { relaunch } from '@tauri-apps/plugin-process';
|
|
||||||
import { check, type Update } from '@tauri-apps/plugin-updater';
|
|
||||||
import { useCallback, useEffect, useState } from 'react';
|
import { useCallback, useEffect, useState } from 'react';
|
||||||
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
import type { UpdateStatus } from '@/platform/types';
|
||||||
|
|
||||||
export interface UpdateStatus {
|
// Re-export UpdateStatus for backwards compatibility
|
||||||
checking: boolean;
|
export type { UpdateStatus };
|
||||||
available: boolean;
|
|
||||||
version?: string;
|
|
||||||
downloading: boolean;
|
|
||||||
installing: boolean;
|
|
||||||
readyToInstall: boolean;
|
|
||||||
error?: string;
|
|
||||||
downloadProgress?: number; // 0-100 percentage
|
|
||||||
downloadedBytes?: number;
|
|
||||||
totalBytes?: number;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Check if we're on Windows (NSIS installer handles restart automatically)
|
|
||||||
const isWindows = () => {
|
|
||||||
return navigator.userAgent.includes('Windows');
|
|
||||||
};
|
|
||||||
|
|
||||||
const isTauri = () => {
|
|
||||||
return '__TAURI_INTERNALS__' in window;
|
|
||||||
};
|
|
||||||
|
|
||||||
export function useAutoUpdater(checkOnMount = false) {
|
export function useAutoUpdater(checkOnMount = false) {
|
||||||
const [status, setStatus] = useState<UpdateStatus>({
|
const platform = usePlatform();
|
||||||
checking: false,
|
const [status, setStatus] = useState<UpdateStatus>(
|
||||||
available: false,
|
platform.updater.getStatus(),
|
||||||
downloading: false,
|
);
|
||||||
installing: false,
|
|
||||||
readyToInstall: false,
|
|
||||||
});
|
|
||||||
|
|
||||||
const [update, setUpdate] = useState<Update | null>(null);
|
// Subscribe to updater status changes
|
||||||
|
useEffect(() => {
|
||||||
|
const unsubscribe = platform.updater.subscribe((newStatus) => {
|
||||||
|
setStatus(newStatus);
|
||||||
|
});
|
||||||
|
return unsubscribe;
|
||||||
|
}, [platform]);
|
||||||
|
|
||||||
const checkForUpdates = useCallback(async () => {
|
const checkForUpdates = useCallback(async () => {
|
||||||
if (!isTauri()) {
|
await platform.updater.checkForUpdates();
|
||||||
return;
|
}, [platform]);
|
||||||
}
|
|
||||||
|
|
||||||
try {
|
const downloadAndInstall = useCallback(async () => {
|
||||||
setStatus((prev) => ({ ...prev, checking: true, error: undefined }));
|
await platform.updater.downloadAndInstall();
|
||||||
|
}, [platform]);
|
||||||
|
|
||||||
const foundUpdate = await check();
|
const restartAndInstall = useCallback(async () => {
|
||||||
|
await platform.updater.restartAndInstall();
|
||||||
if (foundUpdate?.available) {
|
}, [platform]);
|
||||||
setUpdate(foundUpdate);
|
|
||||||
setStatus({
|
|
||||||
checking: false,
|
|
||||||
available: true,
|
|
||||||
version: foundUpdate.version,
|
|
||||||
downloading: false,
|
|
||||||
installing: false,
|
|
||||||
readyToInstall: false,
|
|
||||||
});
|
|
||||||
} else {
|
|
||||||
setStatus({
|
|
||||||
checking: false,
|
|
||||||
available: false,
|
|
||||||
downloading: false,
|
|
||||||
installing: false,
|
|
||||||
readyToInstall: false,
|
|
||||||
});
|
|
||||||
}
|
|
||||||
} catch (error) {
|
|
||||||
setStatus({
|
|
||||||
checking: false,
|
|
||||||
available: false,
|
|
||||||
downloading: false,
|
|
||||||
installing: false,
|
|
||||||
readyToInstall: false,
|
|
||||||
error: error instanceof Error ? error.message : 'Failed to check for updates',
|
|
||||||
});
|
|
||||||
}
|
|
||||||
}, []);
|
|
||||||
|
|
||||||
// Download the update (but don't install yet)
|
|
||||||
const downloadAndInstall = async () => {
|
|
||||||
if (!update || !isTauri()) return;
|
|
||||||
|
|
||||||
try {
|
|
||||||
setStatus((prev) => ({ ...prev, downloading: true, error: undefined }));
|
|
||||||
|
|
||||||
let downloadedBytes = 0;
|
|
||||||
let totalBytes = 0;
|
|
||||||
|
|
||||||
// Just download the update
|
|
||||||
await update.download((event) => {
|
|
||||||
switch (event.event) {
|
|
||||||
case 'Started':
|
|
||||||
totalBytes = event.data.contentLength || 0;
|
|
||||||
downloadedBytes = 0;
|
|
||||||
setStatus((prev) => ({
|
|
||||||
...prev,
|
|
||||||
downloading: true,
|
|
||||||
totalBytes,
|
|
||||||
downloadedBytes: 0,
|
|
||||||
downloadProgress: 0,
|
|
||||||
}));
|
|
||||||
break;
|
|
||||||
case 'Progress': {
|
|
||||||
downloadedBytes += event.data.chunkLength;
|
|
||||||
const progress =
|
|
||||||
totalBytes > 0 ? Math.round((downloadedBytes / totalBytes) * 100) : undefined;
|
|
||||||
setStatus((prev) => ({
|
|
||||||
...prev,
|
|
||||||
downloadedBytes,
|
|
||||||
downloadProgress: progress,
|
|
||||||
}));
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
case 'Finished':
|
|
||||||
setStatus((prev) => ({
|
|
||||||
...prev,
|
|
||||||
downloading: false,
|
|
||||||
readyToInstall: true,
|
|
||||||
downloadProgress: 100,
|
|
||||||
}));
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
});
|
|
||||||
} catch (error) {
|
|
||||||
setStatus((prev) => ({
|
|
||||||
...prev,
|
|
||||||
downloading: false,
|
|
||||||
installing: false,
|
|
||||||
readyToInstall: false,
|
|
||||||
downloadProgress: undefined,
|
|
||||||
downloadedBytes: undefined,
|
|
||||||
totalBytes: undefined,
|
|
||||||
error: error instanceof Error ? error.message : 'Failed to download update',
|
|
||||||
}));
|
|
||||||
}
|
|
||||||
};
|
|
||||||
|
|
||||||
// Install the downloaded update and restart the app
|
|
||||||
const restartAndInstall = async () => {
|
|
||||||
if (!update || !isTauri()) return;
|
|
||||||
|
|
||||||
try {
|
|
||||||
setStatus((prev) => ({ ...prev, installing: true, error: undefined }));
|
|
||||||
|
|
||||||
// Install the update
|
|
||||||
await update.install();
|
|
||||||
|
|
||||||
// On Windows with NSIS, the installer handles the restart automatically.
|
|
||||||
// The process will be killed by the NSIS installer, so we won't reach here.
|
|
||||||
// On macOS/Linux, we need to manually relaunch.
|
|
||||||
if (!isWindows()) {
|
|
||||||
await relaunch();
|
|
||||||
}
|
|
||||||
// If we're on Windows and somehow still running, the NSIS installer
|
|
||||||
// should have already handled everything. Just wait for the process to end.
|
|
||||||
} catch (error) {
|
|
||||||
setStatus((prev) => ({
|
|
||||||
...prev,
|
|
||||||
installing: false,
|
|
||||||
error: error instanceof Error ? error.message : 'Failed to install update',
|
|
||||||
}));
|
|
||||||
}
|
|
||||||
};
|
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (checkOnMount && isTauri()) {
|
if (checkOnMount && platform.metadata.isTauri) {
|
||||||
checkForUpdates();
|
checkForUpdates();
|
||||||
}
|
}
|
||||||
}, [checkOnMount, checkForUpdates]);
|
}, [checkOnMount, checkForUpdates, platform.metadata.isTauri]);
|
||||||
|
|
||||||
return {
|
return {
|
||||||
status,
|
status,
|
||||||
|
|||||||
+145
-1
@@ -1,4 +1,5 @@
|
|||||||
import { useServerStore } from '@/stores/serverStore';
|
import { useServerStore } from '@/stores/serverStore';
|
||||||
|
import type { LanguageCode } from '@/lib/constants/languages';
|
||||||
import type {
|
import type {
|
||||||
VoiceProfileCreate,
|
VoiceProfileCreate,
|
||||||
VoiceProfileResponse,
|
VoiceProfileResponse,
|
||||||
@@ -13,6 +14,16 @@ import type {
|
|||||||
ModelStatusListResponse,
|
ModelStatusListResponse,
|
||||||
ModelDownloadRequest,
|
ModelDownloadRequest,
|
||||||
ActiveTasksResponse,
|
ActiveTasksResponse,
|
||||||
|
StoryCreate,
|
||||||
|
StoryResponse,
|
||||||
|
StoryDetailResponse,
|
||||||
|
StoryItemCreate,
|
||||||
|
StoryItemDetail,
|
||||||
|
StoryItemBatchUpdate,
|
||||||
|
StoryItemReorder,
|
||||||
|
StoryItemMove,
|
||||||
|
StoryItemTrim,
|
||||||
|
StoryItemSplit,
|
||||||
} from './types';
|
} from './types';
|
||||||
|
|
||||||
class ApiClient {
|
class ApiClient {
|
||||||
@@ -110,6 +121,16 @@ class ApiClient {
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
async updateProfileSample(
|
||||||
|
sampleId: string,
|
||||||
|
referenceText: string,
|
||||||
|
): Promise<ProfileSampleResponse> {
|
||||||
|
return this.request<ProfileSampleResponse>(`/profiles/samples/${sampleId}`, {
|
||||||
|
method: 'PUT',
|
||||||
|
body: JSON.stringify({ reference_text: referenceText }),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
async exportProfile(profileId: string): Promise<Blob> {
|
async exportProfile(profileId: string): Promise<Blob> {
|
||||||
const url = `${this.getBaseUrl()}/profiles/${profileId}/export`;
|
const url = `${this.getBaseUrl()}/profiles/${profileId}/export`;
|
||||||
const response = await fetch(url);
|
const response = await fetch(url);
|
||||||
@@ -144,6 +165,32 @@ class ApiClient {
|
|||||||
return response.json();
|
return response.json();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
async uploadAvatar(profileId: string, file: File): Promise<VoiceProfileResponse> {
|
||||||
|
const url = `${this.getBaseUrl()}/profiles/${profileId}/avatar`;
|
||||||
|
const formData = new FormData();
|
||||||
|
formData.append('file', file);
|
||||||
|
|
||||||
|
const response = await fetch(url, {
|
||||||
|
method: 'POST',
|
||||||
|
body: formData,
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!response.ok) {
|
||||||
|
const error = await response.json().catch(() => ({
|
||||||
|
detail: response.statusText,
|
||||||
|
}));
|
||||||
|
throw new Error(error.detail || `HTTP error! status: ${response.status}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
return response.json();
|
||||||
|
}
|
||||||
|
|
||||||
|
async deleteAvatar(profileId: string): Promise<void> {
|
||||||
|
await this.request<void>(`/profiles/${profileId}/avatar`, {
|
||||||
|
method: 'DELETE',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
// Generation
|
// Generation
|
||||||
async generateSpeech(data: GenerationRequest): Promise<GenerationResponse> {
|
async generateSpeech(data: GenerationRequest): Promise<GenerationResponse> {
|
||||||
return this.request<GenerationResponse>('/generate', {
|
return this.request<GenerationResponse>('/generate', {
|
||||||
@@ -234,7 +281,7 @@ class ApiClient {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Transcription
|
// Transcription
|
||||||
async transcribeAudio(file: File, language?: 'en' | 'zh'): Promise<TranscriptionResponse> {
|
async transcribeAudio(file: File, language?: LanguageCode): Promise<TranscriptionResponse> {
|
||||||
const formData = new FormData();
|
const formData = new FormData();
|
||||||
formData.append('file', file);
|
formData.append('file', file);
|
||||||
if (language) {
|
if (language) {
|
||||||
@@ -361,6 +408,103 @@ class ApiClient {
|
|||||||
body: JSON.stringify({ channel_ids: channelIds }),
|
body: JSON.stringify({ channel_ids: channelIds }),
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Stories
|
||||||
|
async listStories(): Promise<StoryResponse[]> {
|
||||||
|
return this.request<StoryResponse[]>('/stories');
|
||||||
|
}
|
||||||
|
|
||||||
|
async createStory(data: StoryCreate): Promise<StoryResponse> {
|
||||||
|
return this.request<StoryResponse>('/stories', {
|
||||||
|
method: 'POST',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async getStory(storyId: string): Promise<StoryDetailResponse> {
|
||||||
|
return this.request<StoryDetailResponse>(`/stories/${storyId}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
async updateStory(storyId: string, data: StoryCreate): Promise<StoryResponse> {
|
||||||
|
return this.request<StoryResponse>(`/stories/${storyId}`, {
|
||||||
|
method: 'PUT',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async deleteStory(storyId: string): Promise<void> {
|
||||||
|
await this.request<void>(`/stories/${storyId}`, {
|
||||||
|
method: 'DELETE',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async addStoryItem(storyId: string, data: StoryItemCreate): Promise<StoryItemDetail> {
|
||||||
|
return this.request<StoryItemDetail>(`/stories/${storyId}/items`, {
|
||||||
|
method: 'POST',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async removeStoryItem(storyId: string, itemId: string): Promise<void> {
|
||||||
|
await this.request<void>(`/stories/${storyId}/items/${itemId}`, {
|
||||||
|
method: 'DELETE',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async updateStoryItemTimes(storyId: string, data: StoryItemBatchUpdate): Promise<void> {
|
||||||
|
await this.request<void>(`/stories/${storyId}/items/times`, {
|
||||||
|
method: 'PUT',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async reorderStoryItems(storyId: string, data: StoryItemReorder): Promise<StoryItemDetail[]> {
|
||||||
|
return this.request<StoryItemDetail[]>(`/stories/${storyId}/items/reorder`, {
|
||||||
|
method: 'PUT',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async moveStoryItem(storyId: string, itemId: string, data: StoryItemMove): Promise<StoryItemDetail> {
|
||||||
|
return this.request<StoryItemDetail>(`/stories/${storyId}/items/${itemId}/move`, {
|
||||||
|
method: 'PUT',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async trimStoryItem(storyId: string, itemId: string, data: StoryItemTrim): Promise<StoryItemDetail> {
|
||||||
|
return this.request<StoryItemDetail>(`/stories/${storyId}/items/${itemId}/trim`, {
|
||||||
|
method: 'PUT',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async splitStoryItem(storyId: string, itemId: string, data: StoryItemSplit): Promise<StoryItemDetail[]> {
|
||||||
|
return this.request<StoryItemDetail[]>(`/stories/${storyId}/items/${itemId}/split`, {
|
||||||
|
method: 'POST',
|
||||||
|
body: JSON.stringify(data),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async duplicateStoryItem(storyId: string, itemId: string): Promise<StoryItemDetail> {
|
||||||
|
return this.request<StoryItemDetail>(`/stories/${storyId}/items/${itemId}/duplicate`, {
|
||||||
|
method: 'POST',
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async exportStoryAudio(storyId: string): Promise<Blob> {
|
||||||
|
const url = `${this.getBaseUrl()}/stories/${storyId}/export-audio`;
|
||||||
|
const response = await fetch(url);
|
||||||
|
|
||||||
|
if (!response.ok) {
|
||||||
|
const error = await response.json().catch(() => ({
|
||||||
|
detail: response.statusText,
|
||||||
|
}));
|
||||||
|
throw new Error(error.detail || `HTTP error! status: ${response.status}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
return response.blob();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
export const apiClient = new ApiClient();
|
export const apiClient = new ApiClient();
|
||||||
|
|||||||
@@ -1,9 +1,10 @@
|
|||||||
// API Types matching backend Pydantic models
|
// API Types matching backend Pydantic models
|
||||||
|
import type { LanguageCode } from '@/lib/constants/languages';
|
||||||
|
|
||||||
export interface VoiceProfileCreate {
|
export interface VoiceProfileCreate {
|
||||||
name: string;
|
name: string;
|
||||||
description?: string;
|
description?: string;
|
||||||
language: 'en' | 'zh';
|
language: LanguageCode;
|
||||||
}
|
}
|
||||||
|
|
||||||
export interface VoiceProfileResponse {
|
export interface VoiceProfileResponse {
|
||||||
@@ -11,6 +12,7 @@ export interface VoiceProfileResponse {
|
|||||||
name: string;
|
name: string;
|
||||||
description?: string;
|
description?: string;
|
||||||
language: string;
|
language: string;
|
||||||
|
avatar_path?: string;
|
||||||
created_at: string;
|
created_at: string;
|
||||||
updated_at: string;
|
updated_at: string;
|
||||||
}
|
}
|
||||||
@@ -29,7 +31,7 @@ export interface ProfileSampleResponse {
|
|||||||
export interface GenerationRequest {
|
export interface GenerationRequest {
|
||||||
profile_id: string;
|
profile_id: string;
|
||||||
text: string;
|
text: string;
|
||||||
language: 'en' | 'zh';
|
language: LanguageCode;
|
||||||
seed?: number;
|
seed?: number;
|
||||||
model_size?: '1.7B' | '0.6B';
|
model_size?: '1.7B' | '0.6B';
|
||||||
}
|
}
|
||||||
@@ -62,7 +64,7 @@ export interface HistoryListResponse {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export interface TranscriptionRequest {
|
export interface TranscriptionRequest {
|
||||||
language?: 'en' | 'zh';
|
language?: LanguageCode;
|
||||||
}
|
}
|
||||||
|
|
||||||
export interface TranscriptionResponse {
|
export interface TranscriptionResponse {
|
||||||
@@ -123,3 +125,79 @@ export interface ActiveTasksResponse {
|
|||||||
downloads: ActiveDownloadTask[];
|
downloads: ActiveDownloadTask[];
|
||||||
generations: ActiveGenerationTask[];
|
generations: ActiveGenerationTask[];
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export interface StoryCreate {
|
||||||
|
name: string;
|
||||||
|
description?: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryResponse {
|
||||||
|
id: string;
|
||||||
|
name: string;
|
||||||
|
description?: string;
|
||||||
|
created_at: string;
|
||||||
|
updated_at: string;
|
||||||
|
item_count: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemDetail {
|
||||||
|
id: string;
|
||||||
|
story_id: string;
|
||||||
|
generation_id: string;
|
||||||
|
start_time_ms: number;
|
||||||
|
track: number;
|
||||||
|
trim_start_ms: number;
|
||||||
|
trim_end_ms: number;
|
||||||
|
created_at: string;
|
||||||
|
profile_id: string;
|
||||||
|
profile_name: string;
|
||||||
|
text: string;
|
||||||
|
language: string;
|
||||||
|
audio_path: string;
|
||||||
|
duration: number;
|
||||||
|
seed?: number;
|
||||||
|
instruct?: string;
|
||||||
|
generation_created_at: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryDetailResponse {
|
||||||
|
id: string;
|
||||||
|
name: string;
|
||||||
|
description?: string;
|
||||||
|
created_at: string;
|
||||||
|
updated_at: string;
|
||||||
|
items: StoryItemDetail[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemCreate {
|
||||||
|
generation_id: string;
|
||||||
|
start_time_ms?: number;
|
||||||
|
track?: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemUpdateTime {
|
||||||
|
generation_id: string;
|
||||||
|
start_time_ms: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemBatchUpdate {
|
||||||
|
updates: StoryItemUpdateTime[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemReorder {
|
||||||
|
generation_ids: string[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemMove {
|
||||||
|
start_time_ms: number;
|
||||||
|
track: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemTrim {
|
||||||
|
trim_start_ms: number;
|
||||||
|
trim_end_ms: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoryItemSplit {
|
||||||
|
split_time_ms: number;
|
||||||
|
}
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
import { useCallback, useEffect, useRef, useState } from 'react';
|
import { useCallback, useEffect, useRef, useState } from 'react';
|
||||||
import { isTauri } from '@/lib/tauri';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
import { convertToWav } from '@/lib/utils/audio';
|
import { convertToWav } from '@/lib/utils/audio';
|
||||||
|
|
||||||
interface UseAudioRecordingOptions {
|
interface UseAudioRecordingOptions {
|
||||||
@@ -8,9 +8,10 @@ interface UseAudioRecordingOptions {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export function useAudioRecording({
|
export function useAudioRecording({
|
||||||
maxDurationSeconds = 30,
|
maxDurationSeconds = 29,
|
||||||
onRecordingComplete,
|
onRecordingComplete,
|
||||||
}: UseAudioRecordingOptions = {}) {
|
}: UseAudioRecordingOptions = {}) {
|
||||||
|
const platform = usePlatform();
|
||||||
const [isRecording, setIsRecording] = useState(false);
|
const [isRecording, setIsRecording] = useState(false);
|
||||||
const [duration, setDuration] = useState(0);
|
const [duration, setDuration] = useState(0);
|
||||||
const [error, setError] = useState<string | null>(null);
|
const [error, setError] = useState<string | null>(null);
|
||||||
@@ -40,15 +41,14 @@ export function useAudioRecording({
|
|||||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||||
|
|
||||||
if (!navigator.mediaDevices || !navigator.mediaDevices.getUserMedia) {
|
if (!navigator.mediaDevices || !navigator.mediaDevices.getUserMedia) {
|
||||||
const isTauriEnv = isTauri();
|
|
||||||
console.error('MediaDevices check:', {
|
console.error('MediaDevices check:', {
|
||||||
hasNavigator: typeof navigator !== 'undefined',
|
hasNavigator: typeof navigator !== 'undefined',
|
||||||
hasMediaDevices: !!navigator?.mediaDevices,
|
hasMediaDevices: !!navigator?.mediaDevices,
|
||||||
hasGetUserMedia: !!navigator?.mediaDevices?.getUserMedia,
|
hasGetUserMedia: !!navigator?.mediaDevices?.getUserMedia,
|
||||||
isTauri: isTauriEnv,
|
isTauri: platform.metadata.isTauri,
|
||||||
});
|
});
|
||||||
|
|
||||||
const errorMsg = isTauriEnv
|
const errorMsg = platform.metadata.isTauri
|
||||||
? 'Microphone access is not available. Please ensure:\n1. The app has microphone permissions in System Settings (macOS: System Settings > Privacy & Security > Microphone)\n2. You restart the app after granting permissions\n3. You are using Tauri v2 with a webview that supports getUserMedia'
|
? 'Microphone access is not available. Please ensure:\n1. The app has microphone permissions in System Settings (macOS: System Settings > Privacy & Security > Microphone)\n2. You restart the app after granting permissions\n3. You are using Tauri v2 with a webview that supports getUserMedia'
|
||||||
: 'Microphone access is not available. Please ensure you are using a secure context (HTTPS or localhost) and that your browser has microphone permissions enabled.';
|
: 'Microphone access is not available. Please ensure you are using a secure context (HTTPS or localhost) and that your browser has microphone permissions enabled.';
|
||||||
setError(errorMsg);
|
setError(errorMsg);
|
||||||
|
|||||||
@@ -0,0 +1,122 @@
|
|||||||
|
import { zodResolver } from '@hookform/resolvers/zod';
|
||||||
|
import { useState } from 'react';
|
||||||
|
import { useForm } from 'react-hook-form';
|
||||||
|
import * as z from 'zod';
|
||||||
|
import { useToast } from '@/components/ui/use-toast';
|
||||||
|
import { apiClient } from '@/lib/api/client';
|
||||||
|
import { LANGUAGE_CODES, type LanguageCode } from '@/lib/constants/languages';
|
||||||
|
import { useGeneration } from '@/lib/hooks/useGeneration';
|
||||||
|
import { useModelDownloadToast } from '@/lib/hooks/useModelDownloadToast';
|
||||||
|
import { useGenerationStore } from '@/stores/generationStore';
|
||||||
|
import { usePlayerStore } from '@/stores/playerStore';
|
||||||
|
|
||||||
|
const generationSchema = z.object({
|
||||||
|
text: z.string().min(1, 'Text is required').max(5000),
|
||||||
|
language: z.enum(LANGUAGE_CODES as [LanguageCode, ...LanguageCode[]]),
|
||||||
|
seed: z.number().int().optional(),
|
||||||
|
modelSize: z.enum(['1.7B', '0.6B']).optional(),
|
||||||
|
instruct: z.string().max(500).optional(),
|
||||||
|
});
|
||||||
|
|
||||||
|
export type GenerationFormValues = z.infer<typeof generationSchema>;
|
||||||
|
|
||||||
|
interface UseGenerationFormOptions {
|
||||||
|
onSuccess?: (generationId: string) => void;
|
||||||
|
defaultValues?: Partial<GenerationFormValues>;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useGenerationForm(options: UseGenerationFormOptions = {}) {
|
||||||
|
const { toast } = useToast();
|
||||||
|
const generation = useGeneration();
|
||||||
|
const setAudioWithAutoPlay = usePlayerStore((state) => state.setAudioWithAutoPlay);
|
||||||
|
const setIsGenerating = useGenerationStore((state) => state.setIsGenerating);
|
||||||
|
const [downloadingModelName, setDownloadingModelName] = useState<string | null>(null);
|
||||||
|
const [downloadingDisplayName, setDownloadingDisplayName] = useState<string | null>(null);
|
||||||
|
|
||||||
|
useModelDownloadToast({
|
||||||
|
modelName: downloadingModelName || '',
|
||||||
|
displayName: downloadingDisplayName || '',
|
||||||
|
enabled: !!downloadingModelName,
|
||||||
|
});
|
||||||
|
|
||||||
|
const form = useForm<GenerationFormValues>({
|
||||||
|
resolver: zodResolver(generationSchema),
|
||||||
|
defaultValues: {
|
||||||
|
text: '',
|
||||||
|
language: 'en',
|
||||||
|
seed: undefined,
|
||||||
|
modelSize: '1.7B',
|
||||||
|
instruct: '',
|
||||||
|
...options.defaultValues,
|
||||||
|
},
|
||||||
|
});
|
||||||
|
|
||||||
|
async function handleSubmit(
|
||||||
|
data: GenerationFormValues,
|
||||||
|
selectedProfileId: string | null,
|
||||||
|
): Promise<void> {
|
||||||
|
if (!selectedProfileId) {
|
||||||
|
toast({
|
||||||
|
title: 'No profile selected',
|
||||||
|
description: 'Please select a voice profile from the cards above.',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
setIsGenerating(true);
|
||||||
|
|
||||||
|
const modelName = `qwen-tts-${data.modelSize}`;
|
||||||
|
const displayName = data.modelSize === '1.7B' ? 'Qwen TTS 1.7B' : 'Qwen TTS 0.6B';
|
||||||
|
|
||||||
|
try {
|
||||||
|
const modelStatus = await apiClient.getModelStatus();
|
||||||
|
const model = modelStatus.models.find((m) => m.model_name === modelName);
|
||||||
|
|
||||||
|
if (model && !model.downloaded) {
|
||||||
|
setDownloadingModelName(modelName);
|
||||||
|
setDownloadingDisplayName(displayName);
|
||||||
|
}
|
||||||
|
} catch (error) {
|
||||||
|
console.error('Failed to check model status:', error);
|
||||||
|
}
|
||||||
|
|
||||||
|
const result = await generation.mutateAsync({
|
||||||
|
profile_id: selectedProfileId,
|
||||||
|
text: data.text,
|
||||||
|
language: data.language,
|
||||||
|
seed: data.seed,
|
||||||
|
model_size: data.modelSize,
|
||||||
|
instruct: data.instruct || undefined,
|
||||||
|
});
|
||||||
|
|
||||||
|
toast({
|
||||||
|
title: 'Generation complete!',
|
||||||
|
description: `Audio generated (${result.duration.toFixed(2)}s)`,
|
||||||
|
});
|
||||||
|
|
||||||
|
const audioUrl = apiClient.getAudioUrl(result.id);
|
||||||
|
setAudioWithAutoPlay(audioUrl, result.id, selectedProfileId, data.text.substring(0, 50));
|
||||||
|
|
||||||
|
form.reset();
|
||||||
|
options.onSuccess?.(result.id);
|
||||||
|
} catch (error) {
|
||||||
|
toast({
|
||||||
|
title: 'Generation failed',
|
||||||
|
description: error instanceof Error ? error.message : 'Failed to generate audio',
|
||||||
|
variant: 'destructive',
|
||||||
|
});
|
||||||
|
} finally {
|
||||||
|
setIsGenerating(false);
|
||||||
|
setDownloadingModelName(null);
|
||||||
|
setDownloadingDisplayName(null);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
form,
|
||||||
|
handleSubmit,
|
||||||
|
isPending: generation.isPending,
|
||||||
|
};
|
||||||
|
}
|
||||||
@@ -1,7 +1,7 @@
|
|||||||
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query';
|
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { apiClient } from '@/lib/api/client';
|
||||||
import type { HistoryQuery } from '@/lib/api/types';
|
import type { HistoryQuery } from '@/lib/api/types';
|
||||||
import { isTauri } from '@/lib/tauri';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
export function useHistory(query?: HistoryQuery) {
|
export function useHistory(query?: HistoryQuery) {
|
||||||
return useQuery({
|
return useQuery({
|
||||||
@@ -30,116 +30,52 @@ export function useDeleteGeneration() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export function useExportGeneration() {
|
export function useExportGeneration() {
|
||||||
|
const platform = usePlatform();
|
||||||
|
|
||||||
return useMutation({
|
return useMutation({
|
||||||
mutationFn: async ({ generationId, text }: { generationId: string; text: string }) => {
|
mutationFn: async ({ generationId, text }: { generationId: string; text: string }) => {
|
||||||
const blob = await apiClient.exportGeneration(generationId);
|
const blob = await apiClient.exportGeneration(generationId);
|
||||||
|
|
||||||
// Create safe filename from text
|
// Create safe filename from text
|
||||||
const safeText = text.substring(0, 30).replace(/[^a-z0-9]/gi, '-').toLowerCase();
|
const safeText = text
|
||||||
|
.substring(0, 30)
|
||||||
|
.replace(/[^a-z0-9]/gi, '-')
|
||||||
|
.toLowerCase();
|
||||||
const filename = `generation-${safeText}.voicebox.zip`;
|
const filename = `generation-${safeText}.voicebox.zip`;
|
||||||
|
|
||||||
if (isTauri()) {
|
await platform.filesystem.saveFile(filename, blob, [
|
||||||
// Use Tauri's native save dialog
|
{
|
||||||
try {
|
name: 'Voicebox Generation',
|
||||||
const { save } = await import('@tauri-apps/plugin-dialog');
|
extensions: ['zip'],
|
||||||
const filePath = await save({
|
},
|
||||||
defaultPath: filename,
|
]);
|
||||||
filters: [
|
|
||||||
{
|
|
||||||
name: 'Voicebox Generation',
|
|
||||||
extensions: ['voicebox.zip', 'zip'],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
if (filePath) {
|
|
||||||
// Write file using Tauri's filesystem API
|
|
||||||
const { writeBinaryFile } = await import('@tauri-apps/plugin-fs');
|
|
||||||
const arrayBuffer = await blob.arrayBuffer();
|
|
||||||
await writeBinaryFile(filePath, new Uint8Array(arrayBuffer));
|
|
||||||
}
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to use Tauri dialog, falling back to browser download:', error);
|
|
||||||
// Fall back to browser download if Tauri dialog fails
|
|
||||||
const url = window.URL.createObjectURL(blob);
|
|
||||||
const a = document.createElement('a');
|
|
||||||
a.href = url;
|
|
||||||
a.download = filename;
|
|
||||||
document.body.appendChild(a);
|
|
||||||
a.click();
|
|
||||||
window.URL.revokeObjectURL(url);
|
|
||||||
document.body.removeChild(a);
|
|
||||||
}
|
|
||||||
} else {
|
|
||||||
// Browser: trigger download
|
|
||||||
const url = window.URL.createObjectURL(blob);
|
|
||||||
const a = document.createElement('a');
|
|
||||||
a.href = url;
|
|
||||||
a.download = filename;
|
|
||||||
document.body.appendChild(a);
|
|
||||||
a.click();
|
|
||||||
window.URL.revokeObjectURL(url);
|
|
||||||
document.body.removeChild(a);
|
|
||||||
}
|
|
||||||
|
|
||||||
return blob;
|
return blob;
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
export function useExportGenerationAudio() {
|
export function useExportGenerationAudio() {
|
||||||
|
const platform = usePlatform();
|
||||||
|
|
||||||
return useMutation({
|
return useMutation({
|
||||||
mutationFn: async ({ generationId, text }: { generationId: string; text: string }) => {
|
mutationFn: async ({ generationId, text }: { generationId: string; text: string }) => {
|
||||||
const blob = await apiClient.exportGenerationAudio(generationId);
|
const blob = await apiClient.exportGenerationAudio(generationId);
|
||||||
|
|
||||||
// Create safe filename from text
|
// Create safe filename from text
|
||||||
const safeText = text.substring(0, 30).replace(/[^a-z0-9]/gi, '-').toLowerCase();
|
const safeText = text
|
||||||
|
.substring(0, 30)
|
||||||
|
.replace(/[^a-z0-9]/gi, '-')
|
||||||
|
.toLowerCase();
|
||||||
const filename = `${safeText}.wav`;
|
const filename = `${safeText}.wav`;
|
||||||
|
|
||||||
if (isTauri()) {
|
await platform.filesystem.saveFile(filename, blob, [
|
||||||
// Use Tauri's native save dialog
|
{
|
||||||
try {
|
name: 'Audio File',
|
||||||
const { save } = await import('@tauri-apps/plugin-dialog');
|
extensions: ['wav'],
|
||||||
const filePath = await save({
|
},
|
||||||
defaultPath: filename,
|
]);
|
||||||
filters: [
|
|
||||||
{
|
|
||||||
name: 'Audio File',
|
|
||||||
extensions: ['wav'],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
if (filePath) {
|
|
||||||
// Write file using Tauri's filesystem API
|
|
||||||
const { writeBinaryFile } = await import('@tauri-apps/plugin-fs');
|
|
||||||
const arrayBuffer = await blob.arrayBuffer();
|
|
||||||
await writeBinaryFile(filePath, new Uint8Array(arrayBuffer));
|
|
||||||
}
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to use Tauri dialog, falling back to browser download:', error);
|
|
||||||
// Fall back to browser download if Tauri dialog fails
|
|
||||||
const url = window.URL.createObjectURL(blob);
|
|
||||||
const a = document.createElement('a');
|
|
||||||
a.href = url;
|
|
||||||
a.download = filename;
|
|
||||||
document.body.appendChild(a);
|
|
||||||
a.click();
|
|
||||||
window.URL.revokeObjectURL(url);
|
|
||||||
document.body.removeChild(a);
|
|
||||||
}
|
|
||||||
} else {
|
|
||||||
// Browser: trigger download
|
|
||||||
const url = window.URL.createObjectURL(blob);
|
|
||||||
const a = document.createElement('a');
|
|
||||||
a.href = url;
|
|
||||||
a.download = filename;
|
|
||||||
document.body.appendChild(a);
|
|
||||||
a.click();
|
|
||||||
window.URL.revokeObjectURL(url);
|
|
||||||
document.body.removeChild(a);
|
|
||||||
}
|
|
||||||
|
|
||||||
return blob;
|
return blob;
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|||||||
@@ -140,8 +140,8 @@ export function useModelDownloadToast({
|
|||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
eventSource.onerror = () => {
|
eventSource.onerror = (error) => {
|
||||||
console.error('SSE error');
|
console.error('SSE error:', error);
|
||||||
eventSource.close();
|
eventSource.close();
|
||||||
eventSourceRef.current = null;
|
eventSourceRef.current = null;
|
||||||
|
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query';
|
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { apiClient } from '@/lib/api/client';
|
||||||
import type { VoiceProfileCreate } from '@/lib/api/types';
|
import type { VoiceProfileCreate } from '@/lib/api/types';
|
||||||
import { isTauri } from '@/lib/tauri';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
export function useProfiles() {
|
export function useProfiles() {
|
||||||
return useQuery({
|
return useQuery({
|
||||||
@@ -98,60 +98,43 @@ export function useDeleteSample() {
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export function useUpdateSample() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ sampleId, referenceText }: { sampleId: string; referenceText: string }) =>
|
||||||
|
apiClient.updateProfileSample(sampleId, referenceText),
|
||||||
|
onSuccess: (data) => {
|
||||||
|
queryClient.invalidateQueries({
|
||||||
|
queryKey: ['profiles', data.profile_id, 'samples'],
|
||||||
|
});
|
||||||
|
queryClient.invalidateQueries({
|
||||||
|
queryKey: ['profiles', data.profile_id],
|
||||||
|
});
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['profiles'] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
export function useExportProfile() {
|
export function useExportProfile() {
|
||||||
|
const platform = usePlatform();
|
||||||
|
|
||||||
return useMutation({
|
return useMutation({
|
||||||
mutationFn: async (profileId: string) => {
|
mutationFn: async (profileId: string) => {
|
||||||
const blob = await apiClient.exportProfile(profileId);
|
const blob = await apiClient.exportProfile(profileId);
|
||||||
|
|
||||||
// Get profile name for filename
|
// Get profile name for filename
|
||||||
const profile = await apiClient.getProfile(profileId);
|
const profile = await apiClient.getProfile(profileId);
|
||||||
const safeName = profile.name.replace(/[^a-z0-9]/gi, '-').toLowerCase();
|
const safeName = profile.name.replace(/[^a-z0-9]/gi, '-').toLowerCase();
|
||||||
const filename = `profile-${safeName}.voicebox.zip`;
|
const filename = `profile-${safeName}.voicebox.zip`;
|
||||||
|
|
||||||
if (isTauri()) {
|
await platform.filesystem.saveFile(filename, blob, [
|
||||||
// Use Tauri's native save dialog
|
{
|
||||||
try {
|
name: 'Voicebox Profile',
|
||||||
const { save } = await import('@tauri-apps/plugin-dialog');
|
extensions: ['zip'],
|
||||||
const filePath = await save({
|
},
|
||||||
defaultPath: filename,
|
]);
|
||||||
filters: [
|
|
||||||
{
|
|
||||||
name: 'Voicebox Profile',
|
|
||||||
extensions: ['voicebox.zip', 'zip'],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
if (filePath) {
|
|
||||||
// Write file using Tauri's filesystem API
|
|
||||||
const { writeBinaryFile } = await import('@tauri-apps/plugin-fs');
|
|
||||||
const arrayBuffer = await blob.arrayBuffer();
|
|
||||||
await writeBinaryFile(filePath, new Uint8Array(arrayBuffer));
|
|
||||||
}
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to use Tauri dialog, falling back to browser download:', error);
|
|
||||||
// Fall back to browser download if Tauri dialog fails
|
|
||||||
const url = window.URL.createObjectURL(blob);
|
|
||||||
const a = document.createElement('a');
|
|
||||||
a.href = url;
|
|
||||||
a.download = filename;
|
|
||||||
document.body.appendChild(a);
|
|
||||||
a.click();
|
|
||||||
window.URL.revokeObjectURL(url);
|
|
||||||
document.body.removeChild(a);
|
|
||||||
}
|
|
||||||
} else {
|
|
||||||
// Browser: trigger download
|
|
||||||
const url = window.URL.createObjectURL(blob);
|
|
||||||
const a = document.createElement('a');
|
|
||||||
a.href = url;
|
|
||||||
a.download = filename;
|
|
||||||
document.body.appendChild(a);
|
|
||||||
a.click();
|
|
||||||
window.URL.revokeObjectURL(url);
|
|
||||||
document.body.removeChild(a);
|
|
||||||
}
|
|
||||||
|
|
||||||
return blob;
|
return blob;
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
@@ -167,3 +150,32 @@ export function useImportProfile() {
|
|||||||
},
|
},
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export function useUploadAvatar() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ profileId, file }: { profileId: string; file: File }) =>
|
||||||
|
apiClient.uploadAvatar(profileId, file),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['profiles'] });
|
||||||
|
queryClient.invalidateQueries({
|
||||||
|
queryKey: ['profiles', variables.profileId],
|
||||||
|
});
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useDeleteAvatar() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: (profileId: string) => apiClient.deleteAvatar(profileId),
|
||||||
|
onSuccess: (_, profileId) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['profiles'] });
|
||||||
|
queryClient.invalidateQueries({
|
||||||
|
queryKey: ['profiles', profileId],
|
||||||
|
});
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,181 @@
|
|||||||
|
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query';
|
||||||
|
import { apiClient } from '@/lib/api/client';
|
||||||
|
import type { StoryCreate, StoryItemCreate, StoryItemBatchUpdate, StoryItemReorder, StoryItemMove, StoryItemTrim, StoryItemSplit } from '@/lib/api/types';
|
||||||
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
|
|
||||||
|
export function useStories() {
|
||||||
|
return useQuery({
|
||||||
|
queryKey: ['stories'],
|
||||||
|
queryFn: () => apiClient.listStories(),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useStory(storyId: string | null) {
|
||||||
|
return useQuery({
|
||||||
|
queryKey: ['stories', storyId],
|
||||||
|
queryFn: () => apiClient.getStory(storyId!),
|
||||||
|
enabled: !!storyId,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useCreateStory() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: (data: StoryCreate) => apiClient.createStory(data),
|
||||||
|
onSuccess: () => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useUpdateStory() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, data }: { storyId: string; data: StoryCreate }) =>
|
||||||
|
apiClient.updateStory(storyId, data),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useDeleteStory() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: (storyId: string) => apiClient.deleteStory(storyId),
|
||||||
|
onSuccess: () => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useAddStoryItem() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, data }: { storyId: string; data: StoryItemCreate }) =>
|
||||||
|
apiClient.addStoryItem(storyId, data),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useRemoveStoryItem() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, itemId }: { storyId: string; itemId: string }) =>
|
||||||
|
apiClient.removeStoryItem(storyId, itemId),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useUpdateStoryItemTimes() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, data }: { storyId: string; data: StoryItemBatchUpdate }) =>
|
||||||
|
apiClient.updateStoryItemTimes(storyId, data),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useReorderStoryItems() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, data }: { storyId: string; data: StoryItemReorder }) =>
|
||||||
|
apiClient.reorderStoryItems(storyId, data),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useMoveStoryItem() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, itemId, data }: { storyId: string; itemId: string; data: StoryItemMove }) =>
|
||||||
|
apiClient.moveStoryItem(storyId, itemId, data),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useTrimStoryItem() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, itemId, data }: { storyId: string; itemId: string; data: StoryItemTrim }) =>
|
||||||
|
apiClient.trimStoryItem(storyId, itemId, data),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useSplitStoryItem() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, itemId, data }: { storyId: string; itemId: string; data: StoryItemSplit }) =>
|
||||||
|
apiClient.splitStoryItem(storyId, itemId, data),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useDuplicateStoryItem() {
|
||||||
|
const queryClient = useQueryClient();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: ({ storyId, itemId }: { storyId: string; itemId: string }) =>
|
||||||
|
apiClient.duplicateStoryItem(storyId, itemId),
|
||||||
|
onSuccess: (_, variables) => {
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories'] });
|
||||||
|
queryClient.invalidateQueries({ queryKey: ['stories', variables.storyId] });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export function useExportStoryAudio() {
|
||||||
|
const platform = usePlatform();
|
||||||
|
|
||||||
|
return useMutation({
|
||||||
|
mutationFn: async ({ storyId, storyName }: { storyId: string; storyName: string }) => {
|
||||||
|
const blob = await apiClient.exportStoryAudio(storyId);
|
||||||
|
|
||||||
|
// Create safe filename
|
||||||
|
const safeName = storyName.substring(0, 50).replace(/[^a-z0-9]/gi, '-').toLowerCase();
|
||||||
|
const filename = `${safeName || 'story'}.wav`;
|
||||||
|
|
||||||
|
await platform.filesystem.saveFile(filename, blob, [
|
||||||
|
{
|
||||||
|
name: 'Audio File',
|
||||||
|
extensions: ['wav'],
|
||||||
|
},
|
||||||
|
]);
|
||||||
|
|
||||||
|
return blob;
|
||||||
|
},
|
||||||
|
});
|
||||||
|
}
|
||||||
@@ -0,0 +1,389 @@
|
|||||||
|
import { useCallback, useEffect, useRef } from 'react';
|
||||||
|
import { apiClient } from '@/lib/api/client';
|
||||||
|
import type { StoryItemDetail } from '@/lib/api/types';
|
||||||
|
import { useStoryStore } from '@/stores/storyStore';
|
||||||
|
|
||||||
|
interface ActiveSource {
|
||||||
|
source: AudioBufferSourceNode;
|
||||||
|
itemId: string;
|
||||||
|
generationId: string;
|
||||||
|
startTimeMs: number;
|
||||||
|
endTimeMs: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Hook for managing timecode-based story playback using Web Audio API.
|
||||||
|
* Supports multiple simultaneous audio sources for overlapping clips on different tracks.
|
||||||
|
* Uses AudioContext for sample-accurate timing synchronization.
|
||||||
|
*/
|
||||||
|
export function useStoryPlayback(items: StoryItemDetail[] | undefined) {
|
||||||
|
const isPlaying = useStoryStore((state) => state.isPlaying);
|
||||||
|
const playbackItems = useStoryStore((state) => state.playbackItems);
|
||||||
|
const playbackStartContextTime = useStoryStore((state) => state.playbackStartContextTime);
|
||||||
|
const playbackStartStoryTime = useStoryStore((state) => state.playbackStartStoryTime);
|
||||||
|
const setPlaybackTiming = useStoryStore((state) => state.setPlaybackTiming);
|
||||||
|
|
||||||
|
// AudioContext instance (created once)
|
||||||
|
const audioContextRef = useRef<AudioContext | null>(null);
|
||||||
|
// Master gain for volume control
|
||||||
|
const masterGainRef = useRef<GainNode | null>(null);
|
||||||
|
// Preloaded AudioBuffers by generation_id (audio file is shared between split clips)
|
||||||
|
const audioBuffersRef = useRef<Map<string, AudioBuffer>>(new Map());
|
||||||
|
// Currently playing AudioBufferSourceNodes by item.id (unique per clip)
|
||||||
|
const activeSourcesRef = useRef<Map<string, ActiveSource>>(new Map());
|
||||||
|
// Animation frame for syncing visual playhead
|
||||||
|
const animationFrameRef = useRef<number | null>(null);
|
||||||
|
|
||||||
|
// Get or create AudioContext and audio graph
|
||||||
|
const getAudioContext = useCallback(() => {
|
||||||
|
if (!audioContextRef.current) {
|
||||||
|
audioContextRef.current = new AudioContext();
|
||||||
|
console.log(
|
||||||
|
'[StoryPlayback] Created AudioContext, sample rate:',
|
||||||
|
audioContextRef.current.sampleRate,
|
||||||
|
);
|
||||||
|
|
||||||
|
// Create master gain node for volume control
|
||||||
|
masterGainRef.current = audioContextRef.current.createGain();
|
||||||
|
masterGainRef.current.gain.value = 1;
|
||||||
|
masterGainRef.current.connect(audioContextRef.current.destination);
|
||||||
|
}
|
||||||
|
// Resume context if suspended (browser autoplay policy)
|
||||||
|
if (audioContextRef.current.state === 'suspended') {
|
||||||
|
audioContextRef.current.resume().catch(() => {
|
||||||
|
// Ignore resume errors
|
||||||
|
});
|
||||||
|
}
|
||||||
|
return audioContextRef.current;
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
// Stop a source by item id
|
||||||
|
const stopSource = useCallback((itemId: string) => {
|
||||||
|
const activeSource = activeSourcesRef.current.get(itemId);
|
||||||
|
if (activeSource) {
|
||||||
|
try {
|
||||||
|
activeSource.source.stop();
|
||||||
|
} catch {
|
||||||
|
// Source may have already stopped
|
||||||
|
}
|
||||||
|
activeSourcesRef.current.delete(itemId);
|
||||||
|
}
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
// Preload audio files as AudioBuffers
|
||||||
|
useEffect(() => {
|
||||||
|
if (!items || items.length === 0) {
|
||||||
|
// Clear preloaded buffers when no items
|
||||||
|
audioBuffersRef.current.clear();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const currentIds = new Set(items.map((item) => item.generation_id));
|
||||||
|
const audioContext = getAudioContext();
|
||||||
|
|
||||||
|
// Remove buffers for items that no longer exist
|
||||||
|
for (const [id] of audioBuffersRef.current) {
|
||||||
|
if (!currentIds.has(id)) {
|
||||||
|
audioBuffersRef.current.delete(id);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Preload audio for new items
|
||||||
|
const preloadPromises: Promise<void>[] = [];
|
||||||
|
for (const item of items) {
|
||||||
|
if (!audioBuffersRef.current.has(item.generation_id)) {
|
||||||
|
const audioUrl = apiClient.getAudioUrl(item.generation_id);
|
||||||
|
console.log('[StoryPlayback] Preloading audio buffer:', item.generation_id);
|
||||||
|
|
||||||
|
const preloadPromise = fetch(audioUrl)
|
||||||
|
.then((response) => response.arrayBuffer())
|
||||||
|
.then((arrayBuffer) => audioContext.decodeAudioData(arrayBuffer))
|
||||||
|
.then((audioBuffer) => {
|
||||||
|
audioBuffersRef.current.set(item.generation_id, audioBuffer);
|
||||||
|
console.log(
|
||||||
|
'[StoryPlayback] Preloaded buffer:',
|
||||||
|
item.generation_id,
|
||||||
|
'duration:',
|
||||||
|
audioBuffer.duration,
|
||||||
|
);
|
||||||
|
})
|
||||||
|
.catch((err) => {
|
||||||
|
console.error('[StoryPlayback] Failed to preload audio:', item.generation_id, err);
|
||||||
|
});
|
||||||
|
|
||||||
|
preloadPromises.push(preloadPromise);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
Promise.all(preloadPromises).then(() => {
|
||||||
|
console.log('[StoryPlayback] Preloaded', audioBuffersRef.current.size, 'audio buffers');
|
||||||
|
});
|
||||||
|
}, [items, getAudioContext]);
|
||||||
|
|
||||||
|
// Cleanup AudioContext on unmount
|
||||||
|
useEffect(() => {
|
||||||
|
return () => {
|
||||||
|
// Stop all sources
|
||||||
|
for (const [itemId] of activeSourcesRef.current) {
|
||||||
|
stopSource(itemId);
|
||||||
|
}
|
||||||
|
activeSourcesRef.current.clear();
|
||||||
|
|
||||||
|
// Clean up audio graph
|
||||||
|
if (masterGainRef.current) {
|
||||||
|
masterGainRef.current.disconnect();
|
||||||
|
masterGainRef.current = null;
|
||||||
|
}
|
||||||
|
if (audioContextRef.current && audioContextRef.current.state !== 'closed') {
|
||||||
|
audioContextRef.current.close().catch(() => {
|
||||||
|
// Ignore errors when closing
|
||||||
|
});
|
||||||
|
audioContextRef.current = null;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (animationFrameRef.current !== null) {
|
||||||
|
cancelAnimationFrame(animationFrameRef.current);
|
||||||
|
}
|
||||||
|
};
|
||||||
|
}, [stopSource]);
|
||||||
|
|
||||||
|
// Find ALL items that should be playing at a given story time
|
||||||
|
const findActiveItems = useCallback(
|
||||||
|
(storyTimeMs: number, itemList: StoryItemDetail[]): StoryItemDetail[] => {
|
||||||
|
return itemList.filter((item) => {
|
||||||
|
const itemStart = item.start_time_ms;
|
||||||
|
// Use effective duration (accounting for trims)
|
||||||
|
const trimStartMs = item.trim_start_ms || 0;
|
||||||
|
const trimEndMs = item.trim_end_ms || 0;
|
||||||
|
const effectiveDurationMs = item.duration * 1000 - trimStartMs - trimEndMs;
|
||||||
|
const itemEnd = item.start_time_ms + effectiveDurationMs;
|
||||||
|
return storyTimeMs >= itemStart && storyTimeMs < itemEnd;
|
||||||
|
});
|
||||||
|
},
|
||||||
|
[],
|
||||||
|
);
|
||||||
|
|
||||||
|
// Convert AudioContext time to story time (ms)
|
||||||
|
const contextTimeToStoryTime = useCallback(
|
||||||
|
(contextTime: number): number => {
|
||||||
|
if (playbackStartContextTime === null || playbackStartStoryTime === null) {
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
const elapsedContextTime = contextTime - playbackStartContextTime;
|
||||||
|
return playbackStartStoryTime + elapsedContextTime * 1000;
|
||||||
|
},
|
||||||
|
[playbackStartContextTime, playbackStartStoryTime],
|
||||||
|
);
|
||||||
|
|
||||||
|
// Convert story time (ms) to AudioContext time
|
||||||
|
const storyTimeToContextTime = useCallback(
|
||||||
|
(storyTimeMs: number): number => {
|
||||||
|
if (playbackStartContextTime === null || playbackStartStoryTime === null) {
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
const elapsedStoryTime = (storyTimeMs - playbackStartStoryTime) / 1000;
|
||||||
|
return playbackStartContextTime + elapsedStoryTime;
|
||||||
|
},
|
||||||
|
[playbackStartContextTime, playbackStartStoryTime],
|
||||||
|
);
|
||||||
|
|
||||||
|
// Stop all sources
|
||||||
|
const stopAllSources = useCallback(() => {
|
||||||
|
console.log('[StoryPlayback] Stopping all sources');
|
||||||
|
for (const [itemId] of activeSourcesRef.current) {
|
||||||
|
stopSource(itemId);
|
||||||
|
}
|
||||||
|
activeSourcesRef.current.clear();
|
||||||
|
}, [stopSource]);
|
||||||
|
|
||||||
|
// Schedule playback for all items that should be playing
|
||||||
|
const schedulePlayback = useCallback(
|
||||||
|
(storyTimeMs: number, itemList: StoryItemDetail[]) => {
|
||||||
|
const audioContext = getAudioContext();
|
||||||
|
const currentContextTime = audioContext.currentTime;
|
||||||
|
|
||||||
|
// Find all items that should be playing
|
||||||
|
const shouldBePlaying = findActiveItems(storyTimeMs, itemList);
|
||||||
|
const shouldBePlayingIds = new Set(shouldBePlaying.map((item) => item.id));
|
||||||
|
|
||||||
|
// Stop sources that shouldn't be playing anymore
|
||||||
|
for (const [itemId] of activeSourcesRef.current) {
|
||||||
|
if (!shouldBePlayingIds.has(itemId)) {
|
||||||
|
stopSource(itemId);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Schedule new sources for items that should be playing
|
||||||
|
for (const item of shouldBePlaying) {
|
||||||
|
if (!activeSourcesRef.current.has(item.id)) {
|
||||||
|
const buffer = audioBuffersRef.current.get(item.generation_id);
|
||||||
|
if (!buffer) {
|
||||||
|
console.warn('[StoryPlayback] Buffer not loaded for:', item.generation_id);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Calculate when this item should start in AudioContext time
|
||||||
|
const itemStartContextTime = storyTimeToContextTime(item.start_time_ms);
|
||||||
|
|
||||||
|
// Calculate effective duration and trim offsets
|
||||||
|
const trimStartSec = (item.trim_start_ms || 0) / 1000;
|
||||||
|
const trimEndSec = (item.trim_end_ms || 0) / 1000;
|
||||||
|
const effectiveDuration = item.duration - trimStartSec - trimEndSec;
|
||||||
|
const itemEndStoryTime = item.start_time_ms + effectiveDuration * 1000;
|
||||||
|
|
||||||
|
// Calculate offset into the buffer (if seeking mid-way)
|
||||||
|
// Offset is relative to the trimmed start of the clip
|
||||||
|
const offsetIntoEffectiveClip = Math.max(0, (storyTimeMs - item.start_time_ms) / 1000);
|
||||||
|
const offsetIntoBuffer = trimStartSec + offsetIntoEffectiveClip;
|
||||||
|
const duration = effectiveDuration - offsetIntoEffectiveClip;
|
||||||
|
|
||||||
|
// If the item should have already started, schedule it to start immediately
|
||||||
|
const startAtContextTime = Math.max(currentContextTime, itemStartContextTime);
|
||||||
|
|
||||||
|
console.log('[StoryPlayback] Scheduling source:', {
|
||||||
|
itemId: item.id,
|
||||||
|
generationId: item.generation_id,
|
||||||
|
storyTimeMs,
|
||||||
|
itemStart: item.start_time_ms,
|
||||||
|
offsetIntoBuffer,
|
||||||
|
startAtContextTime,
|
||||||
|
duration,
|
||||||
|
});
|
||||||
|
|
||||||
|
const source = audioContext.createBufferSource();
|
||||||
|
source.buffer = buffer;
|
||||||
|
source.connect(masterGainRef.current || audioContext.destination);
|
||||||
|
|
||||||
|
const activeSource: ActiveSource = {
|
||||||
|
source,
|
||||||
|
itemId: item.id,
|
||||||
|
generationId: item.generation_id,
|
||||||
|
startTimeMs: item.start_time_ms,
|
||||||
|
endTimeMs: itemEndStoryTime,
|
||||||
|
};
|
||||||
|
|
||||||
|
activeSourcesRef.current.set(item.id, activeSource);
|
||||||
|
|
||||||
|
// Schedule playback
|
||||||
|
source.start(startAtContextTime, offsetIntoBuffer, duration);
|
||||||
|
|
||||||
|
// Clean up when source ends
|
||||||
|
source.onended = () => {
|
||||||
|
console.log('[StoryPlayback] Source ended:', item.id);
|
||||||
|
activeSourcesRef.current.delete(item.id);
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
[getAudioContext, findActiveItems, storyTimeToContextTime, stopSource],
|
||||||
|
);
|
||||||
|
|
||||||
|
// Sync visual playhead from AudioContext time
|
||||||
|
useEffect(() => {
|
||||||
|
if (!isPlaying || playbackStartContextTime === null || playbackStartStoryTime === null) {
|
||||||
|
if (animationFrameRef.current !== null) {
|
||||||
|
cancelAnimationFrame(animationFrameRef.current);
|
||||||
|
animationFrameRef.current = null;
|
||||||
|
}
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const audioContext = getAudioContext();
|
||||||
|
const itemList = playbackItems || [];
|
||||||
|
|
||||||
|
const syncPlayhead = () => {
|
||||||
|
if (!useStoryStore.getState().isPlaying) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const currentContextTime = audioContext.currentTime;
|
||||||
|
const currentStoryTime = contextTimeToStoryTime(currentContextTime);
|
||||||
|
const totalDuration = useStoryStore.getState().totalDurationMs;
|
||||||
|
|
||||||
|
// Update store with current story time
|
||||||
|
useStoryStore.setState({ currentTimeMs: Math.min(currentStoryTime, totalDuration) });
|
||||||
|
|
||||||
|
// Schedule any items that should be playing
|
||||||
|
schedulePlayback(currentStoryTime, itemList);
|
||||||
|
|
||||||
|
// Check if we've reached the end
|
||||||
|
if (currentStoryTime >= totalDuration) {
|
||||||
|
// Check if all sources have ended
|
||||||
|
if (activeSourcesRef.current.size === 0) {
|
||||||
|
console.log('[StoryPlayback] Reached end');
|
||||||
|
useStoryStore.getState().stop();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Continue sync loop
|
||||||
|
animationFrameRef.current = requestAnimationFrame(syncPlayhead);
|
||||||
|
};
|
||||||
|
|
||||||
|
// Initial sync
|
||||||
|
const currentContextTime = audioContext.currentTime;
|
||||||
|
const currentStoryTime = contextTimeToStoryTime(currentContextTime);
|
||||||
|
schedulePlayback(currentStoryTime, itemList);
|
||||||
|
|
||||||
|
// Start sync loop
|
||||||
|
animationFrameRef.current = requestAnimationFrame(syncPlayhead);
|
||||||
|
|
||||||
|
return () => {
|
||||||
|
if (animationFrameRef.current !== null) {
|
||||||
|
cancelAnimationFrame(animationFrameRef.current);
|
||||||
|
animationFrameRef.current = null;
|
||||||
|
}
|
||||||
|
};
|
||||||
|
}, [
|
||||||
|
isPlaying,
|
||||||
|
playbackItems,
|
||||||
|
playbackStartContextTime,
|
||||||
|
playbackStartStoryTime,
|
||||||
|
getAudioContext,
|
||||||
|
contextTimeToStoryTime,
|
||||||
|
schedulePlayback,
|
||||||
|
]);
|
||||||
|
|
||||||
|
// Handle play/pause changes - stop sources when paused
|
||||||
|
useEffect(() => {
|
||||||
|
if (!isPlaying) {
|
||||||
|
console.log('[StoryPlayback] Stopping playback');
|
||||||
|
stopAllSources();
|
||||||
|
}
|
||||||
|
}, [isPlaying, stopAllSources]);
|
||||||
|
|
||||||
|
// Handle seek - reset timing anchors when they become null (triggered by seek)
|
||||||
|
useEffect(() => {
|
||||||
|
if (!isPlaying || !playbackItems || playbackItems.length === 0) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Only run when timing anchors are null (after a seek)
|
||||||
|
if (playbackStartContextTime !== null && playbackStartStoryTime !== null) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const audioContext = getAudioContext();
|
||||||
|
const currentContextTime = audioContext.currentTime;
|
||||||
|
const currentStoryTime = useStoryStore.getState().currentTimeMs;
|
||||||
|
|
||||||
|
console.log('[StoryPlayback] Setting timing anchors after seek:', {
|
||||||
|
contextTime: currentContextTime,
|
||||||
|
storyTime: currentStoryTime,
|
||||||
|
});
|
||||||
|
setPlaybackTiming(currentContextTime, currentStoryTime);
|
||||||
|
|
||||||
|
// Stop all existing sources and reschedule from new position
|
||||||
|
stopAllSources();
|
||||||
|
schedulePlayback(currentStoryTime, playbackItems);
|
||||||
|
}, [
|
||||||
|
isPlaying,
|
||||||
|
playbackItems,
|
||||||
|
playbackStartContextTime,
|
||||||
|
playbackStartStoryTime,
|
||||||
|
getAudioContext,
|
||||||
|
stopAllSources,
|
||||||
|
schedulePlayback,
|
||||||
|
setPlaybackTiming,
|
||||||
|
]);
|
||||||
|
}
|
||||||
@@ -1,6 +1,5 @@
|
|||||||
import { useState, useRef, useCallback, useEffect } from 'react';
|
import { useState, useRef, useCallback, useEffect } from 'react';
|
||||||
import { invoke } from '@tauri-apps/api/core';
|
import { usePlatform } from '@/platform/PlatformContext';
|
||||||
import { isTauri } from '@/lib/tauri';
|
|
||||||
|
|
||||||
interface UseSystemAudioCaptureOptions {
|
interface UseSystemAudioCaptureOptions {
|
||||||
maxDurationSeconds?: number;
|
maxDurationSeconds?: number;
|
||||||
@@ -12,9 +11,10 @@ interface UseSystemAudioCaptureOptions {
|
|||||||
* Uses ScreenCaptureKit on macOS and WASAPI loopback on Windows.
|
* Uses ScreenCaptureKit on macOS and WASAPI loopback on Windows.
|
||||||
*/
|
*/
|
||||||
export function useSystemAudioCapture({
|
export function useSystemAudioCapture({
|
||||||
maxDurationSeconds = 30,
|
maxDurationSeconds = 29,
|
||||||
onRecordingComplete,
|
onRecordingComplete,
|
||||||
}: UseSystemAudioCaptureOptions = {}) {
|
}: UseSystemAudioCaptureOptions = {}) {
|
||||||
|
const platform = usePlatform();
|
||||||
const [isRecording, setIsRecording] = useState(false);
|
const [isRecording, setIsRecording] = useState(false);
|
||||||
const [duration, setDuration] = useState(0);
|
const [duration, setDuration] = useState(0);
|
||||||
const [error, setError] = useState<string | null>(null);
|
const [error, setError] = useState<string | null>(null);
|
||||||
@@ -26,22 +26,12 @@ export function useSystemAudioCapture({
|
|||||||
|
|
||||||
// Check if system audio capture is supported
|
// Check if system audio capture is supported
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!isTauri()) {
|
const supported = platform.audio.isSystemAudioSupported();
|
||||||
setIsSupported(false);
|
setIsSupported(supported);
|
||||||
return;
|
}, [platform]);
|
||||||
}
|
|
||||||
|
|
||||||
invoke<boolean>('is_system_audio_supported')
|
|
||||||
.then((supported) => {
|
|
||||||
setIsSupported(supported);
|
|
||||||
})
|
|
||||||
.catch(() => {
|
|
||||||
setIsSupported(false);
|
|
||||||
});
|
|
||||||
}, []);
|
|
||||||
|
|
||||||
const startRecording = useCallback(async () => {
|
const startRecording = useCallback(async () => {
|
||||||
if (!isTauri()) {
|
if (!platform.metadata.isTauri) {
|
||||||
const errorMsg = 'System audio capture is only available in the desktop app.';
|
const errorMsg = 'System audio capture is only available in the desktop app.';
|
||||||
setError(errorMsg);
|
setError(errorMsg);
|
||||||
return;
|
return;
|
||||||
@@ -58,9 +48,7 @@ export function useSystemAudioCapture({
|
|||||||
setDuration(0);
|
setDuration(0);
|
||||||
|
|
||||||
// Start native capture
|
// Start native capture
|
||||||
await invoke('start_system_audio_capture', {
|
await platform.audio.startSystemAudioCapture(maxDurationSeconds);
|
||||||
maxDurationSecs: maxDurationSeconds,
|
|
||||||
});
|
|
||||||
|
|
||||||
setIsRecording(true);
|
setIsRecording(true);
|
||||||
isRecordingRef.current = true;
|
isRecordingRef.current = true;
|
||||||
@@ -86,10 +74,10 @@ export function useSystemAudioCapture({
|
|||||||
setError(errorMessage);
|
setError(errorMessage);
|
||||||
setIsRecording(false);
|
setIsRecording(false);
|
||||||
}
|
}
|
||||||
}, [maxDurationSeconds, isSupported]);
|
}, [maxDurationSeconds, isSupported, platform]);
|
||||||
|
|
||||||
const stopRecording = useCallback(async () => {
|
const stopRecording = useCallback(async () => {
|
||||||
if (!isRecording || !isTauri()) {
|
if (!isRecording || !platform.metadata.isTauri) {
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -102,17 +90,9 @@ export function useSystemAudioCapture({
|
|||||||
timerRef.current = null;
|
timerRef.current = null;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Stop capture and get base64 WAV data
|
// Stop capture and get Blob
|
||||||
const base64Data = await invoke<string>('stop_system_audio_capture');
|
const blob = await platform.audio.stopSystemAudioCapture();
|
||||||
|
|
||||||
// Convert base64 to Blob
|
|
||||||
const binaryString = atob(base64Data);
|
|
||||||
const bytes = new Uint8Array(binaryString.length);
|
|
||||||
for (let i = 0; i < binaryString.length; i++) {
|
|
||||||
bytes[i] = binaryString.charCodeAt(i);
|
|
||||||
}
|
|
||||||
|
|
||||||
const blob = new Blob([bytes], { type: 'audio/wav' });
|
|
||||||
// Pass the actual recorded duration
|
// Pass the actual recorded duration
|
||||||
const recordedDuration = startTimeRef.current
|
const recordedDuration = startTimeRef.current
|
||||||
? (Date.now() - startTimeRef.current) / 1000
|
? (Date.now() - startTimeRef.current) / 1000
|
||||||
@@ -125,7 +105,7 @@ export function useSystemAudioCapture({
|
|||||||
: 'Failed to stop system audio capture.';
|
: 'Failed to stop system audio capture.';
|
||||||
setError(errorMessage);
|
setError(errorMessage);
|
||||||
}
|
}
|
||||||
}, [isRecording, onRecordingComplete]);
|
}, [isRecording, onRecordingComplete, platform]);
|
||||||
|
|
||||||
// Store stopRecording in ref for use in timer
|
// Store stopRecording in ref for use in timer
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
@@ -155,15 +135,15 @@ export function useSystemAudioCapture({
|
|||||||
timerRef.current = null;
|
timerRef.current = null;
|
||||||
}
|
}
|
||||||
// Cancel recording on unmount if still recording
|
// Cancel recording on unmount if still recording
|
||||||
if (isRecordingRef.current && isTauri()) {
|
if (isRecordingRef.current && platform.metadata.isTauri) {
|
||||||
// Call stop directly without the callback to avoid stale closure
|
// Call stop directly without the callback to avoid stale closure
|
||||||
invoke('stop_system_audio_capture').catch((err) => {
|
platform.audio.stopSystemAudioCapture().catch((err) => {
|
||||||
console.error('Error stopping audio capture on unmount:', err);
|
console.error('Error stopping audio capture on unmount:', err);
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
// biome-ignore lint/correctness/useExhaustiveDependencies: Only run on unmount
|
// biome-ignore lint/correctness/useExhaustiveDependencies: Only run on unmount
|
||||||
}, []);
|
}, [platform]);
|
||||||
|
|
||||||
return {
|
return {
|
||||||
isRecording,
|
isRecording,
|
||||||
|
|||||||
@@ -1,9 +1,10 @@
|
|||||||
import { useMutation } from '@tanstack/react-query';
|
import { useMutation } from '@tanstack/react-query';
|
||||||
import { apiClient } from '@/lib/api/client';
|
import { apiClient } from '@/lib/api/client';
|
||||||
|
import type { LanguageCode } from '@/lib/constants/languages';
|
||||||
|
|
||||||
export function useTranscription() {
|
export function useTranscription() {
|
||||||
return useMutation({
|
return useMutation({
|
||||||
mutationFn: ({ file, language }: { file: File; language?: 'en' | 'zh' }) =>
|
mutationFn: ({ file, language }: { file: File; language?: LanguageCode }) =>
|
||||||
apiClient.transcribeAudio(file, language),
|
apiClient.transcribeAudio(file, language),
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,108 +0,0 @@
|
|||||||
/**
|
|
||||||
* Tauri integration utilities
|
|
||||||
*/
|
|
||||||
|
|
||||||
import { invoke } from '@tauri-apps/api/core';
|
|
||||||
import { listen, emit } from '@tauri-apps/api/event';
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Check if running in Tauri environment
|
|
||||||
*/
|
|
||||||
export function isTauri(): boolean {
|
|
||||||
return '__TAURI_INTERNALS__' in window;
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Check if running on macOS
|
|
||||||
*/
|
|
||||||
export function isMacOS(): boolean {
|
|
||||||
return navigator.platform.toLowerCase().includes('mac');
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Start the bundled Python server (Tauri only)
|
|
||||||
*/
|
|
||||||
export async function startServer(remote = false): Promise<string> {
|
|
||||||
if (!isTauri()) {
|
|
||||||
throw new Error('Not running in Tauri environment');
|
|
||||||
}
|
|
||||||
|
|
||||||
try {
|
|
||||||
const result = await invoke<string>('start_server', { remote });
|
|
||||||
console.log('Server started:', result);
|
|
||||||
return result;
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to start server:', error);
|
|
||||||
throw error;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Stop the bundled Python server (Tauri only)
|
|
||||||
*/
|
|
||||||
export async function stopServer(): Promise<void> {
|
|
||||||
if (!isTauri()) {
|
|
||||||
throw new Error('Not running in Tauri environment');
|
|
||||||
}
|
|
||||||
|
|
||||||
try {
|
|
||||||
await invoke('stop_server');
|
|
||||||
console.log('Server stopped');
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to stop server:', error);
|
|
||||||
throw error;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Set whether the server should keep running when the app closes (Tauri only)
|
|
||||||
*/
|
|
||||||
export async function setKeepServerRunning(keepRunning: boolean): Promise<void> {
|
|
||||||
if (!isTauri()) {
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
|
|
||||||
try {
|
|
||||||
await invoke('set_keep_server_running', { keepRunning });
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to set keep server running setting:', error);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Setup window close handler to check setting and stop server if needed
|
|
||||||
*/
|
|
||||||
export async function setupWindowCloseHandler(): Promise<void> {
|
|
||||||
if (!isTauri()) {
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
|
|
||||||
try {
|
|
||||||
// Listen for window close request from Rust
|
|
||||||
await listen<null>('window-close-requested', async () => {
|
|
||||||
// Import store here to avoid circular dependency
|
|
||||||
const { useServerStore } = await import('@/stores/serverStore');
|
|
||||||
const keepRunning = useServerStore.getState().keepServerRunningOnClose;
|
|
||||||
|
|
||||||
// Check if server was started by this app instance
|
|
||||||
// In dev mode, serverStartedByApp will be false, so we won't try to stop a separately-run server
|
|
||||||
// We need to access the module-level variable - this is a bit hacky but works
|
|
||||||
// @ts-expect-error - accessing module-level variable from another module
|
|
||||||
const serverStartedByApp = window.__voiceboxServerStartedByApp ?? false;
|
|
||||||
|
|
||||||
if (!keepRunning && serverStartedByApp) {
|
|
||||||
// Stop server before closing (only if we started it)
|
|
||||||
try {
|
|
||||||
await stopServer();
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to stop server on close:', error);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Emit event back to Rust to allow close
|
|
||||||
await emit('window-close-allowed');
|
|
||||||
});
|
|
||||||
} catch (error) {
|
|
||||||
console.error('Failed to setup window close handler:', error);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -17,6 +17,41 @@ export function formatAudioDuration(seconds: number): string {
|
|||||||
return `${mins}:${secs.toString().padStart(2, '0')}`;
|
return `${mins}:${secs.toString().padStart(2, '0')}`;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Get audio duration from a File.
|
||||||
|
* If the file has a recordedDuration property (from recording hooks),
|
||||||
|
* use that instead of trying to read metadata. This fixes issues on Windows
|
||||||
|
* where WebM files from MediaRecorder don't have proper duration metadata.
|
||||||
|
*/
|
||||||
|
export async function getAudioDuration(
|
||||||
|
file: File & { recordedDuration?: number },
|
||||||
|
): Promise<number> {
|
||||||
|
if (file.recordedDuration !== undefined && Number.isFinite(file.recordedDuration)) {
|
||||||
|
return file.recordedDuration;
|
||||||
|
}
|
||||||
|
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const audio = new Audio();
|
||||||
|
const url = URL.createObjectURL(file);
|
||||||
|
|
||||||
|
audio.addEventListener('loadedmetadata', () => {
|
||||||
|
URL.revokeObjectURL(url);
|
||||||
|
if (Number.isFinite(audio.duration) && audio.duration > 0) {
|
||||||
|
resolve(audio.duration);
|
||||||
|
} else {
|
||||||
|
reject(new Error('Audio file has invalid duration metadata'));
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
audio.addEventListener('error', () => {
|
||||||
|
URL.revokeObjectURL(url);
|
||||||
|
reject(new Error('Failed to load audio file'));
|
||||||
|
});
|
||||||
|
|
||||||
|
audio.src = url;
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Convert any audio blob to WAV format using Web Audio API.
|
* Convert any audio blob to WAV format using Web Audio API.
|
||||||
* This ensures compatibility without requiring ffmpeg on the backend.
|
* This ensures compatibility without requiring ffmpeg on the backend.
|
||||||
|
|||||||
@@ -0,0 +1,19 @@
|
|||||||
|
const DEBUG = import.meta.env.DEV;
|
||||||
|
|
||||||
|
export const debug = {
|
||||||
|
log: (...args: unknown[]) => {
|
||||||
|
if (DEBUG) {
|
||||||
|
console.log(...args);
|
||||||
|
}
|
||||||
|
},
|
||||||
|
error: (...args: unknown[]) => {
|
||||||
|
if (DEBUG) {
|
||||||
|
console.error(...args);
|
||||||
|
}
|
||||||
|
},
|
||||||
|
warn: (...args: unknown[]) => {
|
||||||
|
if (DEBUG) {
|
||||||
|
console.warn(...args);
|
||||||
|
}
|
||||||
|
},
|
||||||
|
};
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
import { createContext, useContext, type ReactNode } from 'react';
|
||||||
|
import type { Platform } from './types';
|
||||||
|
|
||||||
|
const PlatformContext = createContext<Platform | null>(null);
|
||||||
|
|
||||||
|
export interface PlatformProviderProps {
|
||||||
|
platform: Platform;
|
||||||
|
children: ReactNode;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function PlatformProvider({ platform, children }: PlatformProviderProps) {
|
||||||
|
return (
|
||||||
|
<PlatformContext.Provider value={platform}>
|
||||||
|
{children}
|
||||||
|
</PlatformContext.Provider>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function usePlatform(): Platform {
|
||||||
|
const platform = useContext(PlatformContext);
|
||||||
|
if (!platform) {
|
||||||
|
throw new Error('usePlatform must be used within PlatformProvider');
|
||||||
|
}
|
||||||
|
return platform;
|
||||||
|
}
|
||||||
@@ -0,0 +1,70 @@
|
|||||||
|
/**
|
||||||
|
* Platform abstraction types
|
||||||
|
* These interfaces define the contract that platform implementations must fulfill
|
||||||
|
*/
|
||||||
|
|
||||||
|
export interface FileFilter {
|
||||||
|
name: string;
|
||||||
|
extensions: string[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface PlatformFilesystem {
|
||||||
|
saveFile(filename: string, blob: Blob, filters?: FileFilter[]): Promise<void>;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface UpdateStatus {
|
||||||
|
checking: boolean;
|
||||||
|
available: boolean;
|
||||||
|
version?: string;
|
||||||
|
downloading: boolean;
|
||||||
|
installing: boolean;
|
||||||
|
readyToInstall: boolean;
|
||||||
|
error?: string;
|
||||||
|
downloadProgress?: number; // 0-100 percentage
|
||||||
|
downloadedBytes?: number;
|
||||||
|
totalBytes?: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface PlatformUpdater {
|
||||||
|
checkForUpdates(): Promise<void>;
|
||||||
|
downloadAndInstall(): Promise<void>;
|
||||||
|
restartAndInstall(): Promise<void>;
|
||||||
|
getStatus(): UpdateStatus;
|
||||||
|
subscribe(callback: (status: UpdateStatus) => void): () => void;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface AudioDevice {
|
||||||
|
id: string;
|
||||||
|
name: string;
|
||||||
|
is_default: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface PlatformAudio {
|
||||||
|
isSystemAudioSupported(): boolean;
|
||||||
|
startSystemAudioCapture(maxDurationSecs: number): Promise<void>;
|
||||||
|
stopSystemAudioCapture(): Promise<Blob>;
|
||||||
|
listOutputDevices(): Promise<AudioDevice[]>;
|
||||||
|
playToDevices(audioData: Uint8Array, deviceIds: string[]): Promise<void>;
|
||||||
|
stopPlayback(): void;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface PlatformLifecycle {
|
||||||
|
startServer(remote?: boolean): Promise<string>;
|
||||||
|
stopServer(): Promise<void>;
|
||||||
|
setKeepServerRunning(keep: boolean): Promise<void>;
|
||||||
|
setupWindowCloseHandler(): Promise<void>;
|
||||||
|
onServerReady?: () => void;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface PlatformMetadata {
|
||||||
|
getVersion(): Promise<string>;
|
||||||
|
isTauri: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface Platform {
|
||||||
|
filesystem: PlatformFilesystem;
|
||||||
|
updater: PlatformUpdater;
|
||||||
|
audio: PlatformAudio;
|
||||||
|
lifecycle: PlatformLifecycle;
|
||||||
|
metadata: PlatformMetadata;
|
||||||
|
}
|
||||||
@@ -0,0 +1,135 @@
|
|||||||
|
import { createRootRoute, createRoute, createRouter, Outlet } from '@tanstack/react-router';
|
||||||
|
import { AppFrame } from '@/components/AppFrame/AppFrame';
|
||||||
|
import { AudioTab } from '@/components/AudioTab/AudioTab';
|
||||||
|
import { MainEditor } from '@/components/MainEditor/MainEditor';
|
||||||
|
import { ModelsTab } from '@/components/ModelsTab/ModelsTab';
|
||||||
|
import { ServerTab } from '@/components/ServerTab/ServerTab';
|
||||||
|
import { Sidebar } from '@/components/Sidebar';
|
||||||
|
import { StoriesTab } from '@/components/StoriesTab/StoriesTab';
|
||||||
|
import { Toaster } from '@/components/ui/toaster';
|
||||||
|
import { VoicesTab } from '@/components/VoicesTab/VoicesTab';
|
||||||
|
import { useModelDownloadToast } from '@/lib/hooks/useModelDownloadToast';
|
||||||
|
import { MODEL_DISPLAY_NAMES, useRestoreActiveTasks } from '@/lib/hooks/useRestoreActiveTasks';
|
||||||
|
// Simple platform check that works in both web and Tauri
|
||||||
|
const isMacOS = () => navigator.platform.toLowerCase().includes('mac');
|
||||||
|
|
||||||
|
// Root layout component
|
||||||
|
function RootLayout() {
|
||||||
|
// Monitor active downloads/generations and show toasts for them
|
||||||
|
const activeDownloads = useRestoreActiveTasks();
|
||||||
|
|
||||||
|
return (
|
||||||
|
<AppFrame>
|
||||||
|
<div className="flex flex-1 min-h-0 overflow-hidden">
|
||||||
|
<Sidebar isMacOS={isMacOS()} />
|
||||||
|
|
||||||
|
<main className="flex-1 ml-20 overflow-hidden flex flex-col">
|
||||||
|
<div className="container mx-auto px-8 max-w-[1800px] h-full overflow-hidden flex flex-col">
|
||||||
|
<Outlet />
|
||||||
|
</div>
|
||||||
|
</main>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Show download toasts for any active downloads (from anywhere) */}
|
||||||
|
{activeDownloads.map((download) => {
|
||||||
|
const displayName = MODEL_DISPLAY_NAMES[download.model_name] || download.model_name;
|
||||||
|
return (
|
||||||
|
<DownloadToastRestorer
|
||||||
|
key={download.model_name}
|
||||||
|
modelName={download.model_name}
|
||||||
|
displayName={displayName}
|
||||||
|
/>
|
||||||
|
);
|
||||||
|
})}
|
||||||
|
|
||||||
|
<Toaster />
|
||||||
|
</AppFrame>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Component that restores a download toast for a specific model.
|
||||||
|
*/
|
||||||
|
function DownloadToastRestorer({
|
||||||
|
modelName,
|
||||||
|
displayName,
|
||||||
|
}: {
|
||||||
|
modelName: string;
|
||||||
|
displayName: string;
|
||||||
|
}) {
|
||||||
|
// Use the download toast hook to restore the toast
|
||||||
|
useModelDownloadToast({
|
||||||
|
modelName,
|
||||||
|
displayName,
|
||||||
|
enabled: true,
|
||||||
|
});
|
||||||
|
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Root route with layout
|
||||||
|
const rootRoute = createRootRoute({
|
||||||
|
component: RootLayout,
|
||||||
|
});
|
||||||
|
|
||||||
|
// Index route (main/generate)
|
||||||
|
const indexRoute = createRoute({
|
||||||
|
getParentRoute: () => rootRoute,
|
||||||
|
path: '/',
|
||||||
|
component: MainEditor,
|
||||||
|
});
|
||||||
|
|
||||||
|
// Stories route
|
||||||
|
const storiesRoute = createRoute({
|
||||||
|
getParentRoute: () => rootRoute,
|
||||||
|
path: '/stories',
|
||||||
|
component: StoriesTab,
|
||||||
|
});
|
||||||
|
|
||||||
|
// Voices route
|
||||||
|
const voicesRoute = createRoute({
|
||||||
|
getParentRoute: () => rootRoute,
|
||||||
|
path: '/voices',
|
||||||
|
component: VoicesTab,
|
||||||
|
});
|
||||||
|
|
||||||
|
// Audio route
|
||||||
|
const audioRoute = createRoute({
|
||||||
|
getParentRoute: () => rootRoute,
|
||||||
|
path: '/audio',
|
||||||
|
component: AudioTab,
|
||||||
|
});
|
||||||
|
|
||||||
|
// Models route
|
||||||
|
const modelsRoute = createRoute({
|
||||||
|
getParentRoute: () => rootRoute,
|
||||||
|
path: '/models',
|
||||||
|
component: ModelsTab,
|
||||||
|
});
|
||||||
|
|
||||||
|
// Server route
|
||||||
|
const serverRoute = createRoute({
|
||||||
|
getParentRoute: () => rootRoute,
|
||||||
|
path: '/server',
|
||||||
|
component: ServerTab,
|
||||||
|
});
|
||||||
|
|
||||||
|
// Route tree
|
||||||
|
const routeTree = rootRoute.addChildren([
|
||||||
|
indexRoute,
|
||||||
|
storiesRoute,
|
||||||
|
voicesRoute,
|
||||||
|
audioRoute,
|
||||||
|
modelsRoute,
|
||||||
|
serverRoute,
|
||||||
|
]);
|
||||||
|
|
||||||
|
// Create router
|
||||||
|
export const router = createRouter({ routeTree });
|
||||||
|
|
||||||
|
// Register router for type safety
|
||||||
|
declare module '@tanstack/react-router' {
|
||||||
|
interface Register {
|
||||||
|
router: typeof router;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -11,8 +11,11 @@ interface PlayerState {
|
|||||||
volume: number;
|
volume: number;
|
||||||
isLooping: boolean;
|
isLooping: boolean;
|
||||||
shouldRestart: boolean;
|
shouldRestart: boolean;
|
||||||
|
shouldAutoPlay: boolean;
|
||||||
|
onFinish: (() => void) | null;
|
||||||
|
|
||||||
setAudio: (url: string, id: string, profileId: string | null, title?: string) => void;
|
setAudio: (url: string, id: string, profileId: string | null, title?: string) => void;
|
||||||
|
setAudioWithAutoPlay: (url: string, id: string, profileId: string | null, title?: string) => void;
|
||||||
setIsPlaying: (playing: boolean) => void;
|
setIsPlaying: (playing: boolean) => void;
|
||||||
setCurrentTime: (time: number) => void;
|
setCurrentTime: (time: number) => void;
|
||||||
setDuration: (duration: number) => void;
|
setDuration: (duration: number) => void;
|
||||||
@@ -20,6 +23,8 @@ interface PlayerState {
|
|||||||
toggleLoop: () => void;
|
toggleLoop: () => void;
|
||||||
restartCurrentAudio: () => void;
|
restartCurrentAudio: () => void;
|
||||||
clearRestartFlag: () => void;
|
clearRestartFlag: () => void;
|
||||||
|
clearAutoPlayFlag: () => void;
|
||||||
|
setOnFinish: (callback: (() => void) | null) => void;
|
||||||
reset: () => void;
|
reset: () => void;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -34,6 +39,8 @@ export const usePlayerStore = create<PlayerState>((set) => ({
|
|||||||
volume: 1,
|
volume: 1,
|
||||||
isLooping: false,
|
isLooping: false,
|
||||||
shouldRestart: false,
|
shouldRestart: false,
|
||||||
|
shouldAutoPlay: false,
|
||||||
|
onFinish: null,
|
||||||
|
|
||||||
setAudio: (url, id, profileId, title) =>
|
setAudio: (url, id, profileId, title) =>
|
||||||
set({
|
set({
|
||||||
@@ -44,6 +51,18 @@ export const usePlayerStore = create<PlayerState>((set) => ({
|
|||||||
currentTime: 0,
|
currentTime: 0,
|
||||||
isPlaying: false,
|
isPlaying: false,
|
||||||
shouldRestart: false,
|
shouldRestart: false,
|
||||||
|
shouldAutoPlay: false,
|
||||||
|
}),
|
||||||
|
setAudioWithAutoPlay: (url, id, profileId, title) =>
|
||||||
|
set({
|
||||||
|
audioUrl: url,
|
||||||
|
audioId: id,
|
||||||
|
profileId: profileId || null,
|
||||||
|
title: title || null,
|
||||||
|
currentTime: 0,
|
||||||
|
isPlaying: false,
|
||||||
|
shouldRestart: false,
|
||||||
|
shouldAutoPlay: true,
|
||||||
}),
|
}),
|
||||||
setIsPlaying: (playing) => set({ isPlaying: playing }),
|
setIsPlaying: (playing) => set({ isPlaying: playing }),
|
||||||
setCurrentTime: (time) => set({ currentTime: time }),
|
setCurrentTime: (time) => set({ currentTime: time }),
|
||||||
@@ -52,6 +71,8 @@ export const usePlayerStore = create<PlayerState>((set) => ({
|
|||||||
toggleLoop: () => set((state) => ({ isLooping: !state.isLooping })),
|
toggleLoop: () => set((state) => ({ isLooping: !state.isLooping })),
|
||||||
restartCurrentAudio: () => set({ shouldRestart: true }),
|
restartCurrentAudio: () => set({ shouldRestart: true }),
|
||||||
clearRestartFlag: () => set({ shouldRestart: false }),
|
clearRestartFlag: () => set({ shouldRestart: false }),
|
||||||
|
clearAutoPlayFlag: () => set({ shouldAutoPlay: false }),
|
||||||
|
setOnFinish: (callback) => set({ onFinish: callback }),
|
||||||
reset: () =>
|
reset: () =>
|
||||||
set({
|
set({
|
||||||
audioUrl: null,
|
audioUrl: null,
|
||||||
@@ -63,5 +84,7 @@ export const usePlayerStore = create<PlayerState>((set) => ({
|
|||||||
duration: 0,
|
duration: 0,
|
||||||
isLooping: false,
|
isLooping: false,
|
||||||
shouldRestart: false,
|
shouldRestart: false,
|
||||||
|
shouldAutoPlay: false,
|
||||||
|
onFinish: null,
|
||||||
}),
|
}),
|
||||||
}));
|
}));
|
||||||
|
|||||||
@@ -0,0 +1,148 @@
|
|||||||
|
import { create } from 'zustand';
|
||||||
|
import type { StoryItemDetail } from '@/lib/api/types';
|
||||||
|
|
||||||
|
interface StoryPlaybackState {
|
||||||
|
// Selection
|
||||||
|
selectedStoryId: string | null;
|
||||||
|
setSelectedStoryId: (id: string | null) => void;
|
||||||
|
selectedClipId: string | null;
|
||||||
|
setSelectedClipId: (id: string | null) => void;
|
||||||
|
|
||||||
|
// Track editor UI state
|
||||||
|
trackEditorHeight: number;
|
||||||
|
setTrackEditorHeight: (height: number) => void;
|
||||||
|
|
||||||
|
// Playback state
|
||||||
|
isPlaying: boolean;
|
||||||
|
currentTimeMs: number;
|
||||||
|
totalDurationMs: number;
|
||||||
|
playbackStoryId: string | null;
|
||||||
|
playbackItems: StoryItemDetail[] | null;
|
||||||
|
// Web Audio API timing (null when not playing)
|
||||||
|
playbackStartContextTime: number | null; // AudioContext.currentTime when playback started
|
||||||
|
playbackStartStoryTime: number | null; // Story time (ms) when playback started
|
||||||
|
|
||||||
|
// Actions
|
||||||
|
play: (storyId: string, items: StoryItemDetail[]) => void;
|
||||||
|
pause: () => void;
|
||||||
|
stop: () => void;
|
||||||
|
seek: (timeMs: number) => void;
|
||||||
|
setPlaybackTiming: (contextTime: number, storyTime: number) => void; // Set timing anchors for Web Audio API
|
||||||
|
setActiveStory: (storyId: string, items: StoryItemDetail[], totalDurationMs: number) => void; // Activate story for seeking without playing
|
||||||
|
}
|
||||||
|
|
||||||
|
const DEFAULT_TRACK_EDITOR_HEIGHT = 250;
|
||||||
|
|
||||||
|
export const useStoryStore = create<StoryPlaybackState>((set, get) => ({
|
||||||
|
// Selection
|
||||||
|
selectedStoryId: null,
|
||||||
|
setSelectedStoryId: (id) => set({ selectedStoryId: id }),
|
||||||
|
selectedClipId: null,
|
||||||
|
setSelectedClipId: (id) => set({ selectedClipId: id }),
|
||||||
|
|
||||||
|
// Track editor UI state
|
||||||
|
trackEditorHeight: DEFAULT_TRACK_EDITOR_HEIGHT,
|
||||||
|
setTrackEditorHeight: (height) => set({ trackEditorHeight: height }),
|
||||||
|
|
||||||
|
// Playback state
|
||||||
|
isPlaying: false,
|
||||||
|
currentTimeMs: 0,
|
||||||
|
totalDurationMs: 0,
|
||||||
|
playbackStoryId: null,
|
||||||
|
playbackItems: null,
|
||||||
|
playbackStartContextTime: null,
|
||||||
|
playbackStartStoryTime: null,
|
||||||
|
|
||||||
|
// Actions
|
||||||
|
play: (storyId, items) => {
|
||||||
|
// Calculate total duration from items
|
||||||
|
const maxEndTimeMs = Math.max(
|
||||||
|
...items.map((item) => item.start_time_ms + item.duration * 1000),
|
||||||
|
0,
|
||||||
|
);
|
||||||
|
|
||||||
|
// Find the minimum start time (first item)
|
||||||
|
const minStartTimeMs = Math.min(...items.map((item) => item.start_time_ms), 0);
|
||||||
|
|
||||||
|
// If resuming the same story, keep position; otherwise start at first item
|
||||||
|
const currentState = get();
|
||||||
|
const shouldResume = currentState.playbackStoryId === storyId && currentState.currentTimeMs > 0;
|
||||||
|
const startTimeMs = shouldResume ? currentState.currentTimeMs : minStartTimeMs;
|
||||||
|
|
||||||
|
console.log('[StoryStore] Play called:', {
|
||||||
|
storyId,
|
||||||
|
itemCount: items.length,
|
||||||
|
items: items.map((i) => ({
|
||||||
|
id: i.generation_id,
|
||||||
|
start: i.start_time_ms,
|
||||||
|
duration: i.duration,
|
||||||
|
})),
|
||||||
|
maxEndTimeMs,
|
||||||
|
minStartTimeMs,
|
||||||
|
startTimeMs,
|
||||||
|
shouldResume,
|
||||||
|
});
|
||||||
|
|
||||||
|
set({
|
||||||
|
isPlaying: true,
|
||||||
|
playbackStoryId: storyId,
|
||||||
|
playbackItems: items,
|
||||||
|
totalDurationMs: maxEndTimeMs,
|
||||||
|
currentTimeMs: startTimeMs,
|
||||||
|
// Reset timing anchors - will be set fresh by the playback hook
|
||||||
|
playbackStartContextTime: null,
|
||||||
|
playbackStartStoryTime: null,
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
pause: () => {
|
||||||
|
set({
|
||||||
|
isPlaying: false,
|
||||||
|
// Keep timing anchors so we can resume from same position
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
stop: () => {
|
||||||
|
set({
|
||||||
|
isPlaying: false,
|
||||||
|
currentTimeMs: 0,
|
||||||
|
playbackStoryId: null,
|
||||||
|
playbackItems: null,
|
||||||
|
totalDurationMs: 0,
|
||||||
|
playbackStartContextTime: null,
|
||||||
|
playbackStartStoryTime: null,
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
seek: (timeMs) => {
|
||||||
|
const state = get();
|
||||||
|
const clampedTime = Math.max(0, Math.min(timeMs, state.totalDurationMs));
|
||||||
|
set({
|
||||||
|
currentTimeMs: clampedTime,
|
||||||
|
// Reset timing anchors - will be set by hook when playback resumes
|
||||||
|
playbackStartContextTime: null,
|
||||||
|
playbackStartStoryTime: null,
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
setPlaybackTiming: (contextTime, storyTime) => {
|
||||||
|
set({
|
||||||
|
playbackStartContextTime: contextTime,
|
||||||
|
playbackStartStoryTime: storyTime,
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
setActiveStory: (storyId, items, totalDurationMs) => {
|
||||||
|
const currentState = get();
|
||||||
|
// Only update if switching to a different story
|
||||||
|
if (currentState.playbackStoryId !== storyId) {
|
||||||
|
set({
|
||||||
|
playbackStoryId: storyId,
|
||||||
|
playbackItems: items,
|
||||||
|
totalDurationMs,
|
||||||
|
currentTimeMs: 0,
|
||||||
|
isPlaying: false,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
},
|
||||||
|
}));
|
||||||
@@ -1,5 +1,18 @@
|
|||||||
import { create } from 'zustand';
|
import { create } from 'zustand';
|
||||||
|
|
||||||
|
// Draft state for the create voice profile form
|
||||||
|
export interface ProfileFormDraft {
|
||||||
|
name: string;
|
||||||
|
description: string;
|
||||||
|
language: string;
|
||||||
|
referenceText: string;
|
||||||
|
sampleMode: 'upload' | 'record' | 'system';
|
||||||
|
// Note: File objects can't be persisted, so we store metadata
|
||||||
|
sampleFileName?: string;
|
||||||
|
sampleFileType?: string;
|
||||||
|
sampleFileData?: string; // Base64 encoded
|
||||||
|
}
|
||||||
|
|
||||||
interface UIStore {
|
interface UIStore {
|
||||||
// Sidebar
|
// Sidebar
|
||||||
sidebarOpen: boolean;
|
sidebarOpen: boolean;
|
||||||
@@ -18,6 +31,10 @@ interface UIStore {
|
|||||||
selectedProfileId: string | null;
|
selectedProfileId: string | null;
|
||||||
setSelectedProfileId: (id: string | null) => void;
|
setSelectedProfileId: (id: string | null) => void;
|
||||||
|
|
||||||
|
// Profile form draft (for persisting create voice modal state)
|
||||||
|
profileFormDraft: ProfileFormDraft | null;
|
||||||
|
setProfileFormDraft: (draft: ProfileFormDraft | null) => void;
|
||||||
|
|
||||||
// Theme
|
// Theme
|
||||||
theme: 'light' | 'dark';
|
theme: 'light' | 'dark';
|
||||||
setTheme: (theme: 'light' | 'dark') => void;
|
setTheme: (theme: 'light' | 'dark') => void;
|
||||||
@@ -38,6 +55,9 @@ export const useUIStore = create<UIStore>((set) => ({
|
|||||||
selectedProfileId: null,
|
selectedProfileId: null,
|
||||||
setSelectedProfileId: (id) => set({ selectedProfileId: id }),
|
setSelectedProfileId: (id) => set({ selectedProfileId: id }),
|
||||||
|
|
||||||
|
profileFormDraft: null,
|
||||||
|
setProfileFormDraft: (draft) => set({ profileFormDraft: draft }),
|
||||||
|
|
||||||
theme: 'light',
|
theme: 'light',
|
||||||
setTheme: (theme) => {
|
setTheme: (theme) => {
|
||||||
set({ theme });
|
set({ theme });
|
||||||
|
|||||||
+28
-7
@@ -19,8 +19,13 @@ Production-quality FastAPI backend for Qwen3-TTS voice cloning.
|
|||||||
backend/
|
backend/
|
||||||
├── main.py # FastAPI app with all routes
|
├── main.py # FastAPI app with all routes
|
||||||
├── models.py # Pydantic request/response models
|
├── models.py # Pydantic request/response models
|
||||||
├── tts.py # Qwen3-TTS inference
|
├── platform_detect.py # Platform detection for backend selection
|
||||||
├── transcribe.py # Whisper ASR
|
├── tts.py # TTS backend abstraction (delegates to MLX or PyTorch)
|
||||||
|
├── transcribe.py # STT backend abstraction (delegates to MLX or PyTorch)
|
||||||
|
├── backends/ # Backend implementations
|
||||||
|
│ ├── __init__.py # Backend factory and protocols
|
||||||
|
│ ├── mlx_backend.py # MLX backend (Apple Silicon)
|
||||||
|
│ └── pytorch_backend.py # PyTorch backend (Windows/Linux/Intel)
|
||||||
├── profiles.py # Voice profile CRUD
|
├── profiles.py # Voice profile CRUD
|
||||||
├── history.py # Generation history
|
├── history.py # Generation history
|
||||||
├── studio.py # Audio editing (TODO)
|
├── studio.py # Audio editing (TODO)
|
||||||
@@ -31,6 +36,15 @@ backend/
|
|||||||
└── validation.py # Input validation
|
└── validation.py # Input validation
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Backend Selection
|
||||||
|
|
||||||
|
Voicebox automatically selects the best backend based on platform:
|
||||||
|
|
||||||
|
- **Apple Silicon (M1/M2/M3)**: Uses MLX backend with native Metal acceleration (4-5x faster)
|
||||||
|
- **Windows/Linux/Intel Mac**: Uses PyTorch backend (CUDA GPU if available, CPU fallback)
|
||||||
|
|
||||||
|
The backend is detected at runtime via `platform_detect.py`. Both backends implement the same interface, so the API remains consistent across platforms.
|
||||||
|
|
||||||
## API Endpoints
|
## API Endpoints
|
||||||
|
|
||||||
### Health & Info
|
### Health & Info
|
||||||
@@ -47,12 +61,20 @@ Health check with model status.
|
|||||||
"status": "healthy",
|
"status": "healthy",
|
||||||
"model_loaded": true,
|
"model_loaded": true,
|
||||||
"gpu_available": true,
|
"gpu_available": true,
|
||||||
"vram_used_mb": 1024.5
|
"gpu_type": "Metal (Apple Silicon via MLX)",
|
||||||
|
"backend_type": "mlx",
|
||||||
|
"vram_used_mb": null
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**Backend Types:**
|
||||||
|
- `"mlx"` - MLX backend (Apple Silicon with Metal acceleration)
|
||||||
|
- `"pytorch"` - PyTorch backend (Windows/Linux/Intel Mac)
|
||||||
|
|
||||||
### Voice Profiles
|
### Voice Profiles
|
||||||
|
|
||||||
|
**Note:** The database is automatically initialized when the server starts. No manual setup required.
|
||||||
|
|
||||||
#### `POST /profiles`
|
#### `POST /profiles`
|
||||||
Create a new voice profile.
|
Create a new voice profile.
|
||||||
|
|
||||||
@@ -266,13 +288,12 @@ data/
|
|||||||
pip install -r requirements.txt
|
pip install -r requirements.txt
|
||||||
```
|
```
|
||||||
|
|
||||||
### 2. Initialize Database
|
**Note:** On Apple Silicon, also install MLX dependencies for faster inference:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python -c "from database import init_db; init_db()"
|
pip install -r requirements-mlx.txt
|
||||||
```
|
```
|
||||||
|
|
||||||
### 3. Download Models (Automatic)
|
### 2. Download Models (Automatic)
|
||||||
|
|
||||||
The Qwen3-TTS models are automatically downloaded from HuggingFace Hub on first use, similar to how Whisper models work.
|
The Qwen3-TTS models are automatically downloaded from HuggingFace Hub on first use, similar to how Whisper models work.
|
||||||
|
|
||||||
|
|||||||
@@ -1 +1,3 @@
|
|||||||
# Backend package
|
# Backend package
|
||||||
|
|
||||||
|
__version__ = "0.1.11"
|
||||||
|
|||||||
@@ -0,0 +1,166 @@
|
|||||||
|
"""
|
||||||
|
Backend abstraction layer for TTS and STT.
|
||||||
|
|
||||||
|
Provides a unified interface for MLX and PyTorch backends.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from typing import Protocol, Optional, Tuple, List
|
||||||
|
from typing_extensions import runtime_checkable
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from ..platform_detect import get_backend_type
|
||||||
|
|
||||||
|
|
||||||
|
@runtime_checkable
|
||||||
|
class TTSBackend(Protocol):
|
||||||
|
"""Protocol for TTS backend implementations."""
|
||||||
|
|
||||||
|
async def load_model(self, model_size: str) -> None:
|
||||||
|
"""Load TTS model."""
|
||||||
|
...
|
||||||
|
|
||||||
|
async def create_voice_prompt(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
reference_text: str,
|
||||||
|
use_cache: bool = True,
|
||||||
|
) -> Tuple[dict, bool]:
|
||||||
|
"""
|
||||||
|
Create voice prompt from reference audio.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (voice_prompt_dict, was_cached)
|
||||||
|
"""
|
||||||
|
...
|
||||||
|
|
||||||
|
async def combine_voice_prompts(
|
||||||
|
self,
|
||||||
|
audio_paths: List[str],
|
||||||
|
reference_texts: List[str],
|
||||||
|
) -> Tuple[np.ndarray, str]:
|
||||||
|
"""
|
||||||
|
Combine multiple voice prompts.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (combined_audio_array, combined_text)
|
||||||
|
"""
|
||||||
|
...
|
||||||
|
|
||||||
|
async def generate(
|
||||||
|
self,
|
||||||
|
text: str,
|
||||||
|
voice_prompt: dict,
|
||||||
|
language: str = "en",
|
||||||
|
seed: Optional[int] = None,
|
||||||
|
instruct: Optional[str] = None,
|
||||||
|
) -> Tuple[np.ndarray, int]:
|
||||||
|
"""
|
||||||
|
Generate audio from text.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (audio_array, sample_rate)
|
||||||
|
"""
|
||||||
|
...
|
||||||
|
|
||||||
|
def unload_model(self) -> None:
|
||||||
|
"""Unload model to free memory."""
|
||||||
|
...
|
||||||
|
|
||||||
|
def is_loaded(self) -> bool:
|
||||||
|
"""Check if model is loaded."""
|
||||||
|
...
|
||||||
|
|
||||||
|
def _get_model_path(self, model_size: str) -> str:
|
||||||
|
"""
|
||||||
|
Get model path for a given size.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Model path or HuggingFace Hub ID
|
||||||
|
"""
|
||||||
|
...
|
||||||
|
|
||||||
|
|
||||||
|
@runtime_checkable
|
||||||
|
class STTBackend(Protocol):
|
||||||
|
"""Protocol for STT (Speech-to-Text) backend implementations."""
|
||||||
|
|
||||||
|
async def load_model(self, model_size: str) -> None:
|
||||||
|
"""Load STT model."""
|
||||||
|
...
|
||||||
|
|
||||||
|
async def transcribe(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
language: Optional[str] = None,
|
||||||
|
) -> str:
|
||||||
|
"""
|
||||||
|
Transcribe audio to text.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Transcribed text
|
||||||
|
"""
|
||||||
|
...
|
||||||
|
|
||||||
|
def unload_model(self) -> None:
|
||||||
|
"""Unload model to free memory."""
|
||||||
|
...
|
||||||
|
|
||||||
|
def is_loaded(self) -> bool:
|
||||||
|
"""Check if model is loaded."""
|
||||||
|
...
|
||||||
|
|
||||||
|
|
||||||
|
# Global backend instances
|
||||||
|
_tts_backend: Optional[TTSBackend] = None
|
||||||
|
_stt_backend: Optional[STTBackend] = None
|
||||||
|
|
||||||
|
|
||||||
|
def get_tts_backend() -> TTSBackend:
|
||||||
|
"""
|
||||||
|
Get or create TTS backend instance based on platform.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
TTS backend instance (MLX or PyTorch)
|
||||||
|
"""
|
||||||
|
global _tts_backend
|
||||||
|
|
||||||
|
if _tts_backend is None:
|
||||||
|
backend_type = get_backend_type()
|
||||||
|
|
||||||
|
if backend_type == "mlx":
|
||||||
|
from .mlx_backend import MLXTTSBackend
|
||||||
|
_tts_backend = MLXTTSBackend()
|
||||||
|
else:
|
||||||
|
from .pytorch_backend import PyTorchTTSBackend
|
||||||
|
_tts_backend = PyTorchTTSBackend()
|
||||||
|
|
||||||
|
return _tts_backend
|
||||||
|
|
||||||
|
|
||||||
|
def get_stt_backend() -> STTBackend:
|
||||||
|
"""
|
||||||
|
Get or create STT backend instance based on platform.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
STT backend instance (MLX or PyTorch)
|
||||||
|
"""
|
||||||
|
global _stt_backend
|
||||||
|
|
||||||
|
if _stt_backend is None:
|
||||||
|
backend_type = get_backend_type()
|
||||||
|
|
||||||
|
if backend_type == "mlx":
|
||||||
|
from .mlx_backend import MLXSTTBackend
|
||||||
|
_stt_backend = MLXSTTBackend()
|
||||||
|
else:
|
||||||
|
from .pytorch_backend import PyTorchSTTBackend
|
||||||
|
_stt_backend = PyTorchSTTBackend()
|
||||||
|
|
||||||
|
return _stt_backend
|
||||||
|
|
||||||
|
|
||||||
|
def reset_backends():
|
||||||
|
"""Reset backend instances (useful for testing)."""
|
||||||
|
global _tts_backend, _stt_backend
|
||||||
|
_tts_backend = None
|
||||||
|
_stt_backend = None
|
||||||
@@ -0,0 +1,471 @@
|
|||||||
|
"""
|
||||||
|
MLX backend implementation for TTS and STT using mlx-audio.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from typing import Optional, List, Tuple
|
||||||
|
import asyncio
|
||||||
|
import numpy as np
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from . import TTSBackend, STTBackend
|
||||||
|
from ..utils.cache import get_cache_key, get_cached_voice_prompt, cache_voice_prompt
|
||||||
|
from ..utils.audio import normalize_audio, load_audio
|
||||||
|
from ..utils.progress import get_progress_manager
|
||||||
|
from ..utils.hf_progress import HFProgressTracker, create_hf_progress_callback
|
||||||
|
from ..utils.tasks import get_task_manager
|
||||||
|
|
||||||
|
|
||||||
|
class MLXTTSBackend:
|
||||||
|
"""MLX-based TTS backend using mlx-audio."""
|
||||||
|
|
||||||
|
def __init__(self, model_size: str = "1.7B"):
|
||||||
|
self.model = None
|
||||||
|
self.model_size = model_size
|
||||||
|
self._current_model_size = None
|
||||||
|
|
||||||
|
def is_loaded(self) -> bool:
|
||||||
|
"""Check if model is loaded."""
|
||||||
|
return self.model is not None
|
||||||
|
|
||||||
|
def _get_model_path(self, model_size: str) -> str:
|
||||||
|
"""
|
||||||
|
Get the MLX model path.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
model_size: Model size (1.7B or 0.6B)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
HuggingFace Hub model ID for MLX
|
||||||
|
"""
|
||||||
|
# MLX model mapping
|
||||||
|
mlx_model_map = {
|
||||||
|
"1.7B": "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16",
|
||||||
|
# 0.6B not yet converted to MLX format
|
||||||
|
"0.6B": "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16", # Fallback to 1.7B
|
||||||
|
}
|
||||||
|
|
||||||
|
if model_size not in mlx_model_map:
|
||||||
|
raise ValueError(f"Unknown model size: {model_size}")
|
||||||
|
|
||||||
|
hf_model_id = mlx_model_map[model_size]
|
||||||
|
print(f"Will download MLX model from HuggingFace Hub: {hf_model_id}")
|
||||||
|
|
||||||
|
return hf_model_id
|
||||||
|
|
||||||
|
async def load_model_async(self, model_size: Optional[str] = None):
|
||||||
|
"""
|
||||||
|
Lazy load the MLX TTS model.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
model_size: Model size to load (1.7B or 0.6B)
|
||||||
|
"""
|
||||||
|
if model_size is None:
|
||||||
|
model_size = self.model_size
|
||||||
|
|
||||||
|
# If already loaded with correct size, return
|
||||||
|
if self.model is not None and self._current_model_size == model_size:
|
||||||
|
return
|
||||||
|
|
||||||
|
# Unload existing model if different size requested
|
||||||
|
if self.model is not None and self._current_model_size != model_size:
|
||||||
|
self.unload_model()
|
||||||
|
|
||||||
|
# Run blocking load in thread pool
|
||||||
|
await asyncio.to_thread(self._load_model_sync, model_size)
|
||||||
|
|
||||||
|
# Alias for compatibility
|
||||||
|
load_model = load_model_async
|
||||||
|
|
||||||
|
def _load_model_sync(self, model_size: str):
|
||||||
|
"""Synchronous model loading."""
|
||||||
|
try:
|
||||||
|
from mlx_audio.tts import load
|
||||||
|
|
||||||
|
# Get model path
|
||||||
|
model_path = self._get_model_path(model_size)
|
||||||
|
|
||||||
|
# Set up progress tracking
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
model_name = f"qwen-tts-{model_size}"
|
||||||
|
|
||||||
|
# Start tracking download task
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
task_manager.start_download(model_name)
|
||||||
|
|
||||||
|
print(f"Loading MLX TTS model {model_size}...")
|
||||||
|
|
||||||
|
# Initialize progress state
|
||||||
|
progress_manager.update_progress(
|
||||||
|
model_name=model_name,
|
||||||
|
current=0,
|
||||||
|
total=1,
|
||||||
|
filename="",
|
||||||
|
status="downloading",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Set up progress callback
|
||||||
|
progress_callback = create_hf_progress_callback(model_name, progress_manager)
|
||||||
|
tracker = HFProgressTracker(progress_callback)
|
||||||
|
|
||||||
|
# Use progress tracker during download
|
||||||
|
with tracker.patch_download():
|
||||||
|
# Load MLX model (downloads automatically)
|
||||||
|
self.model = load(model_path)
|
||||||
|
|
||||||
|
self._current_model_size = model_size
|
||||||
|
self.model_size = model_size
|
||||||
|
|
||||||
|
# Mark as complete
|
||||||
|
progress_manager.mark_complete(model_name)
|
||||||
|
task_manager.complete_download(model_name)
|
||||||
|
|
||||||
|
print(f"MLX TTS model {model_size} loaded successfully")
|
||||||
|
|
||||||
|
except ImportError as e:
|
||||||
|
print(f"Error: mlx_audio package not found. Install with: pip install mlx-audio")
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
model_name = f"qwen-tts-{model_size}"
|
||||||
|
progress_manager.mark_error(model_name, str(e))
|
||||||
|
task_manager.error_download(model_name, str(e))
|
||||||
|
raise
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Error loading MLX TTS model: {e}")
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
model_name = f"qwen-tts-{model_size}"
|
||||||
|
progress_manager.mark_error(model_name, str(e))
|
||||||
|
task_manager.error_download(model_name, str(e))
|
||||||
|
raise
|
||||||
|
|
||||||
|
def unload_model(self):
|
||||||
|
"""Unload the model to free memory."""
|
||||||
|
if self.model is not None:
|
||||||
|
del self.model
|
||||||
|
self.model = None
|
||||||
|
self._current_model_size = None
|
||||||
|
print("MLX TTS model unloaded")
|
||||||
|
|
||||||
|
async def create_voice_prompt(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
reference_text: str,
|
||||||
|
use_cache: bool = True,
|
||||||
|
) -> Tuple[dict, bool]:
|
||||||
|
"""
|
||||||
|
Create voice prompt from reference audio.
|
||||||
|
|
||||||
|
MLX backend stores voice prompt as a dict with audio path and text.
|
||||||
|
The actual voice prompt processing happens during generation.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
audio_path: Path to reference audio file
|
||||||
|
reference_text: Transcript of reference audio
|
||||||
|
use_cache: Whether to use cached prompt if available
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (voice_prompt_dict, was_cached)
|
||||||
|
"""
|
||||||
|
await self.load_model_async(None)
|
||||||
|
|
||||||
|
# Check cache if enabled
|
||||||
|
if use_cache:
|
||||||
|
cache_key = get_cache_key(audio_path, reference_text)
|
||||||
|
cached_prompt = get_cached_voice_prompt(cache_key)
|
||||||
|
if cached_prompt is not None:
|
||||||
|
# Return cached prompt (should be dict format)
|
||||||
|
if isinstance(cached_prompt, dict):
|
||||||
|
# Validate that the cached audio file still exists
|
||||||
|
cached_audio_path = cached_prompt.get("ref_audio") or cached_prompt.get("ref_audio_path")
|
||||||
|
if cached_audio_path and Path(cached_audio_path).exists():
|
||||||
|
return cached_prompt, True
|
||||||
|
else:
|
||||||
|
# Cached file no longer exists, invalidate cache
|
||||||
|
print(f"Cached audio file not found: {cached_audio_path}, regenerating prompt")
|
||||||
|
|
||||||
|
# MLX voice prompt format - store audio path and text
|
||||||
|
# The model will process this during generation
|
||||||
|
voice_prompt_items = {
|
||||||
|
"ref_audio": str(audio_path),
|
||||||
|
"ref_text": reference_text,
|
||||||
|
}
|
||||||
|
|
||||||
|
# Cache if enabled
|
||||||
|
if use_cache:
|
||||||
|
cache_key = get_cache_key(audio_path, reference_text)
|
||||||
|
cache_voice_prompt(cache_key, voice_prompt_items)
|
||||||
|
|
||||||
|
return voice_prompt_items, False
|
||||||
|
|
||||||
|
async def combine_voice_prompts(
|
||||||
|
self,
|
||||||
|
audio_paths: List[str],
|
||||||
|
reference_texts: List[str],
|
||||||
|
) -> Tuple[np.ndarray, str]:
|
||||||
|
"""
|
||||||
|
Combine multiple reference samples for better quality.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
audio_paths: List of audio file paths
|
||||||
|
reference_texts: List of reference texts
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (combined_audio, combined_text)
|
||||||
|
"""
|
||||||
|
combined_audio = []
|
||||||
|
|
||||||
|
for audio_path in audio_paths:
|
||||||
|
audio, sr = load_audio(audio_path)
|
||||||
|
audio = normalize_audio(audio)
|
||||||
|
combined_audio.append(audio)
|
||||||
|
|
||||||
|
# Concatenate audio
|
||||||
|
mixed = np.concatenate(combined_audio)
|
||||||
|
mixed = normalize_audio(mixed)
|
||||||
|
|
||||||
|
# Combine texts
|
||||||
|
combined_text = " ".join(reference_texts)
|
||||||
|
|
||||||
|
return mixed, combined_text
|
||||||
|
|
||||||
|
async def generate(
|
||||||
|
self,
|
||||||
|
text: str,
|
||||||
|
voice_prompt: dict,
|
||||||
|
language: str = "en",
|
||||||
|
seed: Optional[int] = None,
|
||||||
|
instruct: Optional[str] = None,
|
||||||
|
) -> Tuple[np.ndarray, int]:
|
||||||
|
"""
|
||||||
|
Generate audio from text using voice prompt.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
text: Text to synthesize
|
||||||
|
voice_prompt: Voice prompt dictionary with ref_audio and ref_text
|
||||||
|
language: Language code (en or zh) - may not be fully supported by MLX
|
||||||
|
seed: Random seed for reproducibility
|
||||||
|
instruct: Natural language instruction (may not be supported by MLX)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (audio_array, sample_rate)
|
||||||
|
"""
|
||||||
|
await self.load_model_async(None)
|
||||||
|
|
||||||
|
print(f"Generating audio for text: {text}")
|
||||||
|
|
||||||
|
def _generate_sync():
|
||||||
|
"""Run synchronous generation in thread pool."""
|
||||||
|
# MLX generate() returns a generator yielding GenerationResult objects
|
||||||
|
audio_chunks = []
|
||||||
|
sample_rate = 24000
|
||||||
|
|
||||||
|
# Set seed if provided (MLX uses numpy random)
|
||||||
|
if seed is not None:
|
||||||
|
import mlx.core as mx
|
||||||
|
np.random.seed(seed)
|
||||||
|
mx.random.seed(seed)
|
||||||
|
|
||||||
|
# Extract voice prompt info
|
||||||
|
ref_audio = voice_prompt.get("ref_audio") or voice_prompt.get("ref_audio_path")
|
||||||
|
ref_text = voice_prompt.get("ref_text", "")
|
||||||
|
|
||||||
|
# Validate that the audio file exists
|
||||||
|
if ref_audio and not Path(ref_audio).exists():
|
||||||
|
print(f"Warning: Audio file not found: {ref_audio}")
|
||||||
|
print("This may be due to a cached voice prompt referencing a deleted temp file.")
|
||||||
|
print("Regenerating without voice prompt.")
|
||||||
|
ref_audio = None
|
||||||
|
|
||||||
|
# Check if model supports voice cloning via generate method
|
||||||
|
# MLX API may support ref_audio parameter directly
|
||||||
|
try:
|
||||||
|
# Try with voice cloning parameters if supported
|
||||||
|
if ref_audio:
|
||||||
|
# Check if generate accepts ref_audio parameter
|
||||||
|
import inspect
|
||||||
|
sig = inspect.signature(self.model.generate)
|
||||||
|
if "ref_audio" in sig.parameters:
|
||||||
|
# Generate with voice cloning
|
||||||
|
for result in self.model.generate(text, ref_audio=ref_audio, ref_text=ref_text):
|
||||||
|
audio_chunks.append(np.array(result.audio))
|
||||||
|
sample_rate = result.sample_rate
|
||||||
|
else:
|
||||||
|
# Fallback: generate without voice cloning
|
||||||
|
for result in self.model.generate(text):
|
||||||
|
audio_chunks.append(np.array(result.audio))
|
||||||
|
sample_rate = result.sample_rate
|
||||||
|
else:
|
||||||
|
# No voice prompt, generate normally
|
||||||
|
for result in self.model.generate(text):
|
||||||
|
audio_chunks.append(np.array(result.audio))
|
||||||
|
sample_rate = result.sample_rate
|
||||||
|
except Exception as e:
|
||||||
|
# If voice cloning fails, try without it
|
||||||
|
print(f"Warning: Voice cloning failed, generating without voice prompt: {e}")
|
||||||
|
for result in self.model.generate(text):
|
||||||
|
audio_chunks.append(np.array(result.audio))
|
||||||
|
sample_rate = result.sample_rate
|
||||||
|
|
||||||
|
# Concatenate all chunks
|
||||||
|
if audio_chunks:
|
||||||
|
audio = np.concatenate([np.asarray(chunk, dtype=np.float32) for chunk in audio_chunks])
|
||||||
|
else:
|
||||||
|
# Fallback: empty audio
|
||||||
|
audio = np.array([], dtype=np.float32)
|
||||||
|
|
||||||
|
return audio, sample_rate
|
||||||
|
|
||||||
|
# Run blocking inference in thread pool
|
||||||
|
audio, sample_rate = await asyncio.to_thread(_generate_sync)
|
||||||
|
|
||||||
|
return audio, sample_rate
|
||||||
|
|
||||||
|
|
||||||
|
class MLXSTTBackend:
|
||||||
|
"""MLX-based STT backend using mlx-audio Whisper."""
|
||||||
|
|
||||||
|
def __init__(self, model_size: str = "base"):
|
||||||
|
self.model = None
|
||||||
|
self.model_size = model_size
|
||||||
|
|
||||||
|
def is_loaded(self) -> bool:
|
||||||
|
"""Check if model is loaded."""
|
||||||
|
return self.model is not None
|
||||||
|
|
||||||
|
async def load_model_async(self, model_size: Optional[str] = None):
|
||||||
|
"""
|
||||||
|
Lazy load the MLX Whisper model.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
model_size: Model size (tiny, base, small, medium, large)
|
||||||
|
"""
|
||||||
|
if model_size is None:
|
||||||
|
model_size = self.model_size
|
||||||
|
|
||||||
|
if self.model is not None and self.model_size == model_size:
|
||||||
|
return
|
||||||
|
|
||||||
|
# Run blocking load in thread pool
|
||||||
|
await asyncio.to_thread(self._load_model_sync, model_size)
|
||||||
|
|
||||||
|
# Alias for compatibility
|
||||||
|
load_model = load_model_async
|
||||||
|
|
||||||
|
def _load_model_sync(self, model_size: str):
|
||||||
|
"""Synchronous model loading."""
|
||||||
|
try:
|
||||||
|
# IMPORTANT: Set up progress tracking BEFORE importing mlx_audio
|
||||||
|
# This ensures tqdm is patched before any HuggingFace Hub imports
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
progress_model_name = f"whisper-{model_size}"
|
||||||
|
|
||||||
|
# Set up progress callback and tracker
|
||||||
|
progress_callback = create_hf_progress_callback(progress_model_name, progress_manager)
|
||||||
|
tracker = HFProgressTracker(progress_callback)
|
||||||
|
|
||||||
|
# Patch tqdm BEFORE importing mlx_audio
|
||||||
|
# This is critical because mlx_audio imports huggingface_hub which imports tqdm
|
||||||
|
print("[DEBUG] Starting tqdm patch BEFORE mlx_audio import")
|
||||||
|
tracker_context = tracker.patch_download()
|
||||||
|
tracker_context.__enter__()
|
||||||
|
print("[DEBUG] tqdm patched, now importing mlx_audio")
|
||||||
|
|
||||||
|
# NOW import mlx_audio - it will use our patched tqdm
|
||||||
|
from mlx_audio.stt import load
|
||||||
|
|
||||||
|
# MLX Whisper uses the standard OpenAI models
|
||||||
|
model_name = f"openai/whisper-{model_size}"
|
||||||
|
|
||||||
|
# Start tracking download task
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
task_manager.start_download(progress_model_name)
|
||||||
|
|
||||||
|
print(f"Loading MLX Whisper model {model_size}...")
|
||||||
|
|
||||||
|
# Initialize progress state
|
||||||
|
progress_manager.update_progress(
|
||||||
|
model_name=progress_model_name,
|
||||||
|
current=0,
|
||||||
|
total=1,
|
||||||
|
filename="",
|
||||||
|
status="downloading",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Load the model (tqdm is already patched from above)
|
||||||
|
try:
|
||||||
|
self.model = load(model_name)
|
||||||
|
finally:
|
||||||
|
# Exit the patch context
|
||||||
|
tracker_context.__exit__(None, None, None)
|
||||||
|
|
||||||
|
self.model_size = model_size
|
||||||
|
|
||||||
|
# Mark as complete
|
||||||
|
progress_manager.mark_complete(progress_model_name)
|
||||||
|
task_manager.complete_download(progress_model_name)
|
||||||
|
|
||||||
|
print(f"MLX Whisper model {model_size} loaded successfully")
|
||||||
|
|
||||||
|
except ImportError as e:
|
||||||
|
print(f"Error: mlx_audio package not found. Install with: pip install mlx-audio")
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
progress_model_name = f"whisper-{model_size}"
|
||||||
|
progress_manager.mark_error(progress_model_name, str(e))
|
||||||
|
task_manager.error_download(progress_model_name, str(e))
|
||||||
|
raise
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Error loading MLX Whisper model: {e}")
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
progress_model_name = f"whisper-{model_size}"
|
||||||
|
progress_manager.mark_error(progress_model_name, str(e))
|
||||||
|
task_manager.error_download(progress_model_name, str(e))
|
||||||
|
raise
|
||||||
|
|
||||||
|
def unload_model(self):
|
||||||
|
"""Unload the model to free memory."""
|
||||||
|
if self.model is not None:
|
||||||
|
del self.model
|
||||||
|
self.model = None
|
||||||
|
print("MLX Whisper model unloaded")
|
||||||
|
|
||||||
|
async def transcribe(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
language: Optional[str] = None,
|
||||||
|
) -> str:
|
||||||
|
"""
|
||||||
|
Transcribe audio to text.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
audio_path: Path to audio file
|
||||||
|
language: Optional language hint (en or zh)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Transcribed text
|
||||||
|
"""
|
||||||
|
await self.load_model_async(None)
|
||||||
|
|
||||||
|
def _transcribe_sync():
|
||||||
|
"""Run synchronous transcription in thread pool."""
|
||||||
|
# MLX Whisper transcription using generate method
|
||||||
|
# The generate method accepts audio path directly
|
||||||
|
decode_options = {}
|
||||||
|
if language:
|
||||||
|
decode_options["language"] = language
|
||||||
|
|
||||||
|
result = self.model.generate(str(audio_path), **decode_options)
|
||||||
|
|
||||||
|
# Extract text from result
|
||||||
|
if isinstance(result, str):
|
||||||
|
return result.strip()
|
||||||
|
elif isinstance(result, dict):
|
||||||
|
return result.get("text", "").strip()
|
||||||
|
elif hasattr(result, "text"):
|
||||||
|
return result.text.strip()
|
||||||
|
else:
|
||||||
|
return str(result).strip()
|
||||||
|
|
||||||
|
# Run blocking transcription in thread pool
|
||||||
|
return await asyncio.to_thread(_transcribe_sync)
|
||||||
@@ -0,0 +1,486 @@
|
|||||||
|
"""
|
||||||
|
PyTorch backend implementation for TTS and STT.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from typing import Optional, List, Tuple
|
||||||
|
import asyncio
|
||||||
|
import torch
|
||||||
|
import numpy as np
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from . import TTSBackend, STTBackend
|
||||||
|
from ..utils.cache import get_cache_key, get_cached_voice_prompt, cache_voice_prompt
|
||||||
|
from ..utils.audio import normalize_audio, load_audio
|
||||||
|
from ..utils.progress import get_progress_manager
|
||||||
|
from ..utils.hf_progress import HFProgressTracker, create_hf_progress_callback
|
||||||
|
from ..utils.tasks import get_task_manager
|
||||||
|
|
||||||
|
|
||||||
|
class PyTorchTTSBackend:
|
||||||
|
"""PyTorch-based TTS backend using Qwen3-TTS."""
|
||||||
|
|
||||||
|
def __init__(self, model_size: str = "1.7B"):
|
||||||
|
self.model = None
|
||||||
|
self.model_size = model_size
|
||||||
|
self.device = self._get_device()
|
||||||
|
self._current_model_size = None
|
||||||
|
|
||||||
|
def _get_device(self) -> str:
|
||||||
|
"""Get the best available device."""
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
return "cuda"
|
||||||
|
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
||||||
|
# MPS can have issues, use CPU for stability
|
||||||
|
return "cpu"
|
||||||
|
return "cpu"
|
||||||
|
|
||||||
|
def is_loaded(self) -> bool:
|
||||||
|
"""Check if model is loaded."""
|
||||||
|
return self.model is not None
|
||||||
|
|
||||||
|
def _get_model_path(self, model_size: str) -> str:
|
||||||
|
"""
|
||||||
|
Get the HuggingFace Hub model ID.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
model_size: Model size (1.7B or 0.6B)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
HuggingFace Hub model ID
|
||||||
|
"""
|
||||||
|
hf_model_map = {
|
||||||
|
"1.7B": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
|
||||||
|
"0.6B": "Qwen/Qwen3-TTS-12Hz-0.6B-Base",
|
||||||
|
}
|
||||||
|
|
||||||
|
if model_size not in hf_model_map:
|
||||||
|
raise ValueError(f"Unknown model size: {model_size}")
|
||||||
|
|
||||||
|
return hf_model_map[model_size]
|
||||||
|
|
||||||
|
async def load_model_async(self, model_size: Optional[str] = None):
|
||||||
|
"""
|
||||||
|
Lazy load the TTS model with automatic downloading from HuggingFace Hub.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
model_size: Model size to load (1.7B or 0.6B)
|
||||||
|
"""
|
||||||
|
if model_size is None:
|
||||||
|
model_size = self.model_size
|
||||||
|
|
||||||
|
# If already loaded with correct size, return
|
||||||
|
if self.model is not None and self._current_model_size == model_size:
|
||||||
|
return
|
||||||
|
|
||||||
|
# Unload existing model if different size requested
|
||||||
|
if self.model is not None and self._current_model_size != model_size:
|
||||||
|
self.unload_model()
|
||||||
|
|
||||||
|
# Run blocking load in thread pool
|
||||||
|
await asyncio.to_thread(self._load_model_sync, model_size)
|
||||||
|
|
||||||
|
# Alias for compatibility
|
||||||
|
load_model = load_model_async
|
||||||
|
|
||||||
|
def _load_model_sync(self, model_size: str):
|
||||||
|
"""Synchronous model loading."""
|
||||||
|
try:
|
||||||
|
# IMPORTANT: Set up progress tracking BEFORE importing qwen_tts
|
||||||
|
# This ensures tqdm is patched before any HuggingFace Hub imports
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
model_name = f"qwen-tts-{model_size}"
|
||||||
|
|
||||||
|
# Set up progress callback and tracker
|
||||||
|
progress_callback = create_hf_progress_callback(model_name, progress_manager)
|
||||||
|
tracker = HFProgressTracker(progress_callback)
|
||||||
|
|
||||||
|
# Patch tqdm BEFORE importing qwen_tts
|
||||||
|
tracker_context = tracker.patch_download()
|
||||||
|
tracker_context.__enter__()
|
||||||
|
|
||||||
|
# NOW import qwen_tts - it will use our patched tqdm
|
||||||
|
from qwen_tts import Qwen3TTSModel
|
||||||
|
|
||||||
|
# Get model path (local or HuggingFace Hub ID)
|
||||||
|
model_path = self._get_model_path(model_size)
|
||||||
|
|
||||||
|
print(f"Loading TTS model {model_size} on {self.device}...")
|
||||||
|
|
||||||
|
# Start tracking download task
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
task_manager.start_download(model_name)
|
||||||
|
|
||||||
|
# Initialize progress state to show download has started
|
||||||
|
progress_manager.update_progress(
|
||||||
|
model_name=model_name,
|
||||||
|
current=0,
|
||||||
|
total=1, # Set to 1 initially, will be updated by callback
|
||||||
|
filename="",
|
||||||
|
status="downloading",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Load the model (tqdm is already patched from above)
|
||||||
|
try:
|
||||||
|
self.model = Qwen3TTSModel.from_pretrained(
|
||||||
|
model_path,
|
||||||
|
device_map=self.device,
|
||||||
|
torch_dtype=torch.float32 if self.device == "cpu" else torch.bfloat16,
|
||||||
|
)
|
||||||
|
finally:
|
||||||
|
# Exit the patch context
|
||||||
|
tracker_context.__exit__(None, None, None)
|
||||||
|
|
||||||
|
# Mark as complete
|
||||||
|
progress_manager.mark_complete(model_name)
|
||||||
|
task_manager.complete_download(model_name)
|
||||||
|
|
||||||
|
self._current_model_size = model_size
|
||||||
|
self.model_size = model_size
|
||||||
|
|
||||||
|
print(f"TTS model {model_size} loaded successfully")
|
||||||
|
|
||||||
|
except ImportError as e:
|
||||||
|
print(f"Error: qwen_tts package not found. Install with: pip install git+https://github.com/QwenLM/Qwen3-TTS.git")
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
model_name = f"qwen-tts-{model_size}"
|
||||||
|
progress_manager.mark_error(model_name, str(e))
|
||||||
|
task_manager.error_download(model_name, str(e))
|
||||||
|
raise
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Error loading TTS model: {e}")
|
||||||
|
print(f"Tip: The model will be automatically downloaded from HuggingFace Hub on first use.")
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
model_name = f"qwen-tts-{model_size}"
|
||||||
|
progress_manager.mark_error(model_name, str(e))
|
||||||
|
task_manager.error_download(model_name, str(e))
|
||||||
|
raise
|
||||||
|
|
||||||
|
def unload_model(self):
|
||||||
|
"""Unload the model to free memory."""
|
||||||
|
if self.model is not None:
|
||||||
|
del self.model
|
||||||
|
self.model = None
|
||||||
|
self._current_model_size = None
|
||||||
|
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
|
||||||
|
print("TTS model unloaded")
|
||||||
|
|
||||||
|
async def create_voice_prompt(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
reference_text: str,
|
||||||
|
use_cache: bool = True,
|
||||||
|
) -> Tuple[dict, bool]:
|
||||||
|
"""
|
||||||
|
Create voice prompt from reference audio.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
audio_path: Path to reference audio file
|
||||||
|
reference_text: Transcript of reference audio
|
||||||
|
use_cache: Whether to use cached prompt if available
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (voice_prompt_dict, was_cached)
|
||||||
|
"""
|
||||||
|
await self.load_model_async(None)
|
||||||
|
|
||||||
|
# Check cache if enabled
|
||||||
|
if use_cache:
|
||||||
|
cache_key = get_cache_key(audio_path, reference_text)
|
||||||
|
cached_prompt = get_cached_voice_prompt(cache_key)
|
||||||
|
if cached_prompt is not None:
|
||||||
|
# Cache stores as torch.Tensor but actual prompt is dict
|
||||||
|
# Convert if needed
|
||||||
|
if isinstance(cached_prompt, dict):
|
||||||
|
# For PyTorch backend, the dict should contain tensors, not file paths
|
||||||
|
# So we can safely return it
|
||||||
|
return cached_prompt, True
|
||||||
|
elif isinstance(cached_prompt, torch.Tensor):
|
||||||
|
# Legacy cache format - convert to dict
|
||||||
|
# This shouldn't happen in practice, but handle it
|
||||||
|
return {"prompt": cached_prompt}, True
|
||||||
|
|
||||||
|
def _create_prompt_sync():
|
||||||
|
"""Run synchronous voice prompt creation in thread pool."""
|
||||||
|
return self.model.create_voice_clone_prompt(
|
||||||
|
ref_audio=str(audio_path),
|
||||||
|
ref_text=reference_text,
|
||||||
|
x_vector_only_mode=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
# Run blocking operation in thread pool
|
||||||
|
voice_prompt_items = await asyncio.to_thread(_create_prompt_sync)
|
||||||
|
|
||||||
|
# Cache if enabled
|
||||||
|
if use_cache:
|
||||||
|
cache_key = get_cache_key(audio_path, reference_text)
|
||||||
|
cache_voice_prompt(cache_key, voice_prompt_items)
|
||||||
|
|
||||||
|
return voice_prompt_items, False
|
||||||
|
|
||||||
|
async def combine_voice_prompts(
|
||||||
|
self,
|
||||||
|
audio_paths: List[str],
|
||||||
|
reference_texts: List[str],
|
||||||
|
) -> Tuple[np.ndarray, str]:
|
||||||
|
"""
|
||||||
|
Combine multiple reference samples for better quality.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
audio_paths: List of audio file paths
|
||||||
|
reference_texts: List of reference texts
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (combined_audio, combined_text)
|
||||||
|
"""
|
||||||
|
combined_audio = []
|
||||||
|
|
||||||
|
for audio_path in audio_paths:
|
||||||
|
audio, sr = load_audio(audio_path)
|
||||||
|
audio = normalize_audio(audio)
|
||||||
|
combined_audio.append(audio)
|
||||||
|
|
||||||
|
# Concatenate audio
|
||||||
|
mixed = np.concatenate(combined_audio)
|
||||||
|
mixed = normalize_audio(mixed)
|
||||||
|
|
||||||
|
# Combine texts
|
||||||
|
combined_text = " ".join(reference_texts)
|
||||||
|
|
||||||
|
return mixed, combined_text
|
||||||
|
|
||||||
|
async def generate(
|
||||||
|
self,
|
||||||
|
text: str,
|
||||||
|
voice_prompt: dict,
|
||||||
|
language: str = "en",
|
||||||
|
seed: Optional[int] = None,
|
||||||
|
instruct: Optional[str] = None,
|
||||||
|
) -> Tuple[np.ndarray, int]:
|
||||||
|
"""
|
||||||
|
Generate audio from text using voice prompt.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
text: Text to synthesize
|
||||||
|
voice_prompt: Voice prompt dictionary from create_voice_prompt
|
||||||
|
language: Language code (en or zh)
|
||||||
|
seed: Random seed for reproducibility
|
||||||
|
instruct: Natural language instruction for speech delivery control
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (audio_array, sample_rate)
|
||||||
|
"""
|
||||||
|
# Load model
|
||||||
|
await self.load_model_async(None)
|
||||||
|
|
||||||
|
def _generate_sync():
|
||||||
|
"""Run synchronous generation in thread pool."""
|
||||||
|
# Set seed if provided
|
||||||
|
if seed is not None:
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
torch.cuda.manual_seed(seed)
|
||||||
|
|
||||||
|
# Generate audio - this is the blocking operation
|
||||||
|
wavs, sample_rate = self.model.generate_voice_clone(
|
||||||
|
text=text,
|
||||||
|
voice_clone_prompt=voice_prompt,
|
||||||
|
instruct=instruct,
|
||||||
|
)
|
||||||
|
return wavs[0], sample_rate
|
||||||
|
|
||||||
|
# Run blocking inference in thread pool to avoid blocking event loop
|
||||||
|
audio, sample_rate = await asyncio.to_thread(_generate_sync)
|
||||||
|
|
||||||
|
return audio, sample_rate
|
||||||
|
|
||||||
|
|
||||||
|
class PyTorchSTTBackend:
|
||||||
|
"""PyTorch-based STT backend using Whisper."""
|
||||||
|
|
||||||
|
def __init__(self, model_size: str = "base"):
|
||||||
|
self.model = None
|
||||||
|
self.processor = None
|
||||||
|
self.model_size = model_size
|
||||||
|
self.device = self._get_device()
|
||||||
|
|
||||||
|
def _get_device(self) -> str:
|
||||||
|
"""Get the best available device."""
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
return "cuda"
|
||||||
|
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
||||||
|
# MPS support for Whisper
|
||||||
|
return "cpu" # Use CPU for stability
|
||||||
|
return "cpu"
|
||||||
|
|
||||||
|
def is_loaded(self) -> bool:
|
||||||
|
"""Check if model is loaded."""
|
||||||
|
return self.model is not None
|
||||||
|
|
||||||
|
async def load_model_async(self, model_size: Optional[str] = None):
|
||||||
|
"""
|
||||||
|
Lazy load the Whisper model.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
model_size: Model size (tiny, base, small, medium, large)
|
||||||
|
"""
|
||||||
|
print(f"[DEBUG] load_model_async called with size: {model_size}")
|
||||||
|
if model_size is None:
|
||||||
|
model_size = self.model_size
|
||||||
|
|
||||||
|
print(f"[DEBUG] Model already loaded? {self.model is not None}, current size: {self.model_size}, requested: {model_size}")
|
||||||
|
if self.model is not None and self.model_size == model_size:
|
||||||
|
print(f"[DEBUG] Early return - model already loaded")
|
||||||
|
return
|
||||||
|
|
||||||
|
print(f"[DEBUG] Calling asyncio.to_thread for _load_model_sync")
|
||||||
|
# Run blocking load in thread pool
|
||||||
|
await asyncio.to_thread(self._load_model_sync, model_size)
|
||||||
|
print(f"[DEBUG] asyncio.to_thread completed")
|
||||||
|
|
||||||
|
# Alias for compatibility
|
||||||
|
load_model = load_model_async
|
||||||
|
|
||||||
|
def _load_model_sync(self, model_size: str):
|
||||||
|
"""Synchronous model loading."""
|
||||||
|
print(f"[DEBUG] _load_model_sync called for Whisper {model_size}")
|
||||||
|
try:
|
||||||
|
# IMPORTANT: Set up progress tracking BEFORE importing transformers
|
||||||
|
# This ensures tqdm is patched before any HuggingFace Hub imports
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
progress_model_name = f"whisper-{model_size}"
|
||||||
|
|
||||||
|
# Set up progress callback and tracker
|
||||||
|
progress_callback = create_hf_progress_callback(progress_model_name, progress_manager)
|
||||||
|
tracker = HFProgressTracker(progress_callback)
|
||||||
|
|
||||||
|
# Patch tqdm BEFORE importing transformers
|
||||||
|
print("[DEBUG] Starting tqdm patch BEFORE transformers import")
|
||||||
|
tracker_context = tracker.patch_download()
|
||||||
|
tracker_context.__enter__()
|
||||||
|
print("[DEBUG] tqdm patched, now importing transformers")
|
||||||
|
|
||||||
|
# NOW import transformers - it will use our patched tqdm
|
||||||
|
from transformers import WhisperProcessor, WhisperForConditionalGeneration
|
||||||
|
|
||||||
|
model_name = f"openai/whisper-{model_size}"
|
||||||
|
print(f"[DEBUG] Model name: {model_name}")
|
||||||
|
|
||||||
|
# Start tracking download task
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
task_manager.start_download(progress_model_name)
|
||||||
|
print(f"[DEBUG] Task manager started download")
|
||||||
|
|
||||||
|
print(f"Loading Whisper model {model_size} on {self.device}...")
|
||||||
|
|
||||||
|
# Initialize progress state to show download has started
|
||||||
|
print(f"[DEBUG] Calling update_progress...")
|
||||||
|
progress_manager.update_progress(
|
||||||
|
model_name=progress_model_name,
|
||||||
|
current=0,
|
||||||
|
total=1, # Set to 1 initially, will be updated by callback
|
||||||
|
filename="",
|
||||||
|
status="downloading",
|
||||||
|
)
|
||||||
|
print(f"[DEBUG] update_progress called, listeners: {len(progress_manager._listeners.get(progress_model_name, []))}")
|
||||||
|
|
||||||
|
# Load models (tqdm is already patched from above)
|
||||||
|
try:
|
||||||
|
self.processor = WhisperProcessor.from_pretrained(model_name)
|
||||||
|
self.model = WhisperForConditionalGeneration.from_pretrained(model_name)
|
||||||
|
finally:
|
||||||
|
# Exit the patch context
|
||||||
|
tracker_context.__exit__(None, None, None)
|
||||||
|
|
||||||
|
self.model.to(self.device)
|
||||||
|
self.model_size = model_size
|
||||||
|
|
||||||
|
# Mark as complete
|
||||||
|
progress_manager.mark_complete(progress_model_name)
|
||||||
|
task_manager.complete_download(progress_model_name)
|
||||||
|
|
||||||
|
print(f"Whisper model {model_size} loaded successfully")
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Error loading Whisper model: {e}")
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
progress_model_name = f"whisper-{model_size}"
|
||||||
|
progress_manager.mark_error(progress_model_name, str(e))
|
||||||
|
task_manager.error_download(progress_model_name, str(e))
|
||||||
|
raise
|
||||||
|
|
||||||
|
def unload_model(self):
|
||||||
|
"""Unload the model to free memory."""
|
||||||
|
if self.model is not None:
|
||||||
|
del self.model
|
||||||
|
del self.processor
|
||||||
|
self.model = None
|
||||||
|
self.processor = None
|
||||||
|
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
|
||||||
|
print("Whisper model unloaded")
|
||||||
|
|
||||||
|
async def transcribe(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
language: Optional[str] = None,
|
||||||
|
) -> str:
|
||||||
|
"""
|
||||||
|
Transcribe audio to text.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
audio_path: Path to audio file
|
||||||
|
language: Optional language hint (en or zh)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Transcribed text
|
||||||
|
"""
|
||||||
|
await self.load_model_async(None)
|
||||||
|
|
||||||
|
def _transcribe_sync():
|
||||||
|
"""Run synchronous transcription in thread pool."""
|
||||||
|
# Load audio
|
||||||
|
audio, sr = load_audio(audio_path, sample_rate=16000)
|
||||||
|
|
||||||
|
# Process audio
|
||||||
|
inputs = self.processor(
|
||||||
|
audio,
|
||||||
|
sampling_rate=16000,
|
||||||
|
return_tensors="pt",
|
||||||
|
)
|
||||||
|
inputs = inputs.to(self.device)
|
||||||
|
|
||||||
|
# Set language if provided
|
||||||
|
forced_decoder_ids = None
|
||||||
|
if language:
|
||||||
|
# Support all languages from frontend: en, zh, ja, ko, de, fr, ru, pt, es, it
|
||||||
|
# Whisper supports these and many more
|
||||||
|
forced_decoder_ids = self.processor.get_decoder_prompt_ids(
|
||||||
|
language=language,
|
||||||
|
task="transcribe",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Generate transcription
|
||||||
|
with torch.no_grad():
|
||||||
|
predicted_ids = self.model.generate(
|
||||||
|
inputs["input_features"],
|
||||||
|
forced_decoder_ids=forced_decoder_ids,
|
||||||
|
)
|
||||||
|
|
||||||
|
# Decode
|
||||||
|
transcription = self.processor.batch_decode(
|
||||||
|
predicted_ids,
|
||||||
|
skip_special_tokens=True,
|
||||||
|
)[0]
|
||||||
|
|
||||||
|
return transcription.strip()
|
||||||
|
|
||||||
|
# Run blocking transcription in thread pool
|
||||||
|
return await asyncio.to_thread(_transcribe_sync)
|
||||||
+38
-8
@@ -4,16 +4,19 @@ PyInstaller build script for creating standalone Python server binary.
|
|||||||
|
|
||||||
import PyInstaller.__main__
|
import PyInstaller.__main__
|
||||||
import os
|
import os
|
||||||
|
import platform
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
def is_apple_silicon():
|
||||||
|
"""Check if running on Apple Silicon."""
|
||||||
|
return platform.system() == "Darwin" and platform.machine() == "arm64"
|
||||||
|
|
||||||
|
|
||||||
def build_server():
|
def build_server():
|
||||||
"""Build Python server as standalone binary."""
|
"""Build Python server as standalone binary."""
|
||||||
backend_dir = Path(__file__).parent
|
backend_dir = Path(__file__).parent
|
||||||
|
|
||||||
# Check for local editable qwen_tts install
|
|
||||||
local_qwen_path = Path.home() / 'Projects' / 'voice' / 'Qwen3-TTS'
|
|
||||||
|
|
||||||
# PyInstaller arguments
|
# PyInstaller arguments
|
||||||
args = [
|
args = [
|
||||||
'server.py', # Use server.py as entry point instead of main.py
|
'server.py', # Use server.py as entry point instead of main.py
|
||||||
@@ -21,12 +24,13 @@ def build_server():
|
|||||||
'--name', 'voicebox-server',
|
'--name', 'voicebox-server',
|
||||||
]
|
]
|
||||||
|
|
||||||
# Add local qwen_tts path if it exists (for editable installs)
|
# Add local qwen_tts path if specified (for editable installs)
|
||||||
if local_qwen_path.exists():
|
qwen_tts_path = os.getenv('QWEN_TTS_PATH')
|
||||||
args.extend(['--paths', str(local_qwen_path)])
|
if qwen_tts_path and Path(qwen_tts_path).exists():
|
||||||
print(f"Using local qwen_tts source from: {local_qwen_path}")
|
args.extend(['--paths', str(qwen_tts_path)])
|
||||||
|
print(f"Using local qwen_tts source from: {qwen_tts_path}")
|
||||||
|
|
||||||
# Add hidden imports
|
# Add common hidden imports
|
||||||
args.extend([
|
args.extend([
|
||||||
'--hidden-import', 'backend',
|
'--hidden-import', 'backend',
|
||||||
'--hidden-import', 'backend.main',
|
'--hidden-import', 'backend.main',
|
||||||
@@ -37,6 +41,9 @@ def build_server():
|
|||||||
'--hidden-import', 'backend.history',
|
'--hidden-import', 'backend.history',
|
||||||
'--hidden-import', 'backend.tts',
|
'--hidden-import', 'backend.tts',
|
||||||
'--hidden-import', 'backend.transcribe',
|
'--hidden-import', 'backend.transcribe',
|
||||||
|
'--hidden-import', 'backend.platform_detect',
|
||||||
|
'--hidden-import', 'backend.backends',
|
||||||
|
'--hidden-import', 'backend.backends.pytorch_backend',
|
||||||
'--hidden-import', 'backend.utils.audio',
|
'--hidden-import', 'backend.utils.audio',
|
||||||
'--hidden-import', 'backend.utils.cache',
|
'--hidden-import', 'backend.utils.cache',
|
||||||
'--hidden-import', 'backend.utils.progress',
|
'--hidden-import', 'backend.utils.progress',
|
||||||
@@ -61,6 +68,29 @@ def build_server():
|
|||||||
# Fix for pkg_resources and jaraco namespace packages
|
# Fix for pkg_resources and jaraco namespace packages
|
||||||
'--hidden-import', 'pkg_resources.extern',
|
'--hidden-import', 'pkg_resources.extern',
|
||||||
'--collect-submodules', 'jaraco',
|
'--collect-submodules', 'jaraco',
|
||||||
|
])
|
||||||
|
|
||||||
|
# Add MLX-specific imports if building on Apple Silicon
|
||||||
|
if is_apple_silicon():
|
||||||
|
print("Building for Apple Silicon - including MLX dependencies")
|
||||||
|
args.extend([
|
||||||
|
'--hidden-import', 'backend.backends.mlx_backend',
|
||||||
|
'--hidden-import', 'mlx',
|
||||||
|
'--hidden-import', 'mlx.core',
|
||||||
|
'--hidden-import', 'mlx.nn',
|
||||||
|
'--hidden-import', 'mlx_audio',
|
||||||
|
'--hidden-import', 'mlx_audio.tts',
|
||||||
|
'--hidden-import', 'mlx_audio.stt',
|
||||||
|
'--collect-submodules', 'mlx',
|
||||||
|
'--collect-submodules', 'mlx_audio',
|
||||||
|
# Collect MLX data files including Metal shader libraries (.metallib)
|
||||||
|
'--collect-data', 'mlx',
|
||||||
|
'--collect-data', 'mlx_audio',
|
||||||
|
])
|
||||||
|
else:
|
||||||
|
print("Building for non-Apple Silicon platform - PyTorch only")
|
||||||
|
|
||||||
|
args.extend([
|
||||||
'--noconfirm',
|
'--noconfirm',
|
||||||
'--clean',
|
'--clean',
|
||||||
])
|
])
|
||||||
|
|||||||
+154
-1
@@ -17,11 +17,12 @@ Base = declarative_base()
|
|||||||
class VoiceProfile(Base):
|
class VoiceProfile(Base):
|
||||||
"""Voice profile database model."""
|
"""Voice profile database model."""
|
||||||
__tablename__ = "profiles"
|
__tablename__ = "profiles"
|
||||||
|
|
||||||
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
|
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
|
||||||
name = Column(String, unique=True, nullable=False)
|
name = Column(String, unique=True, nullable=False)
|
||||||
description = Column(Text)
|
description = Column(Text)
|
||||||
language = Column(String, default="en")
|
language = Column(String, default="en")
|
||||||
|
avatar_path = Column(String, nullable=True)
|
||||||
created_at = Column(DateTime, default=datetime.utcnow)
|
created_at = Column(DateTime, default=datetime.utcnow)
|
||||||
updated_at = Column(DateTime, default=datetime.utcnow, onupdate=datetime.utcnow)
|
updated_at = Column(DateTime, default=datetime.utcnow, onupdate=datetime.utcnow)
|
||||||
|
|
||||||
@@ -51,6 +52,31 @@ class Generation(Base):
|
|||||||
created_at = Column(DateTime, default=datetime.utcnow)
|
created_at = Column(DateTime, default=datetime.utcnow)
|
||||||
|
|
||||||
|
|
||||||
|
class Story(Base):
|
||||||
|
"""Story database model."""
|
||||||
|
__tablename__ = "stories"
|
||||||
|
|
||||||
|
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
|
||||||
|
name = Column(String, nullable=False)
|
||||||
|
description = Column(Text)
|
||||||
|
created_at = Column(DateTime, default=datetime.utcnow)
|
||||||
|
updated_at = Column(DateTime, default=datetime.utcnow, onupdate=datetime.utcnow)
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItem(Base):
|
||||||
|
"""Story item database model (links generations to stories)."""
|
||||||
|
__tablename__ = "story_items"
|
||||||
|
|
||||||
|
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
|
||||||
|
story_id = Column(String, ForeignKey("stories.id"), nullable=False)
|
||||||
|
generation_id = Column(String, ForeignKey("generations.id"), nullable=False)
|
||||||
|
start_time_ms = Column(Integer, nullable=False, default=0) # Milliseconds from story start
|
||||||
|
track = Column(Integer, nullable=False, default=0) # Track number (0 = main track)
|
||||||
|
trim_start_ms = Column(Integer, nullable=False, default=0) # Milliseconds trimmed from start
|
||||||
|
trim_end_ms = Column(Integer, nullable=False, default=0) # Milliseconds trimmed from end
|
||||||
|
created_at = Column(DateTime, default=datetime.utcnow)
|
||||||
|
|
||||||
|
|
||||||
class Project(Base):
|
class Project(Base):
|
||||||
"""Audio studio project database model."""
|
"""Audio studio project database model."""
|
||||||
__tablename__ = "projects"
|
__tablename__ = "projects"
|
||||||
@@ -108,6 +134,10 @@ def init_db():
|
|||||||
)
|
)
|
||||||
|
|
||||||
SessionLocal = sessionmaker(autocommit=False, autoflush=False, bind=engine)
|
SessionLocal = sessionmaker(autocommit=False, autoflush=False, bind=engine)
|
||||||
|
|
||||||
|
# Run migrations before creating tables
|
||||||
|
_run_migrations(engine)
|
||||||
|
|
||||||
Base.metadata.create_all(bind=engine)
|
Base.metadata.create_all(bind=engine)
|
||||||
|
|
||||||
# Create default channel if it doesn't exist
|
# Create default channel if it doesn't exist
|
||||||
@@ -136,6 +166,129 @@ def init_db():
|
|||||||
db.close()
|
db.close()
|
||||||
|
|
||||||
|
|
||||||
|
def _run_migrations(engine):
|
||||||
|
"""Run database migrations."""
|
||||||
|
from sqlalchemy import inspect, text
|
||||||
|
|
||||||
|
inspector = inspect(engine)
|
||||||
|
|
||||||
|
# Check if story_items table exists
|
||||||
|
if 'story_items' not in inspector.get_table_names():
|
||||||
|
return # Table doesn't exist yet, will be created fresh
|
||||||
|
|
||||||
|
# Get columns in story_items table
|
||||||
|
columns = {col['name'] for col in inspector.get_columns('story_items')}
|
||||||
|
|
||||||
|
# Migration: Remove position column and ensure start_time_ms exists
|
||||||
|
# SQLite doesn't support DROP COLUMN easily, so we recreate the table
|
||||||
|
if 'position' in columns:
|
||||||
|
print("Migrating story_items: removing position column, using start_time_ms")
|
||||||
|
|
||||||
|
with engine.connect() as conn:
|
||||||
|
# Check if start_time_ms already exists
|
||||||
|
has_start_time = 'start_time_ms' in columns
|
||||||
|
|
||||||
|
if not has_start_time:
|
||||||
|
# First, add the new column temporarily
|
||||||
|
conn.execute(text("ALTER TABLE story_items ADD COLUMN start_time_ms INTEGER DEFAULT 0"))
|
||||||
|
|
||||||
|
# Calculate timecodes from position ordering
|
||||||
|
result = conn.execute(text("""
|
||||||
|
SELECT si.id, si.story_id, si.position, g.duration
|
||||||
|
FROM story_items si
|
||||||
|
JOIN generations g ON si.generation_id = g.id
|
||||||
|
ORDER BY si.story_id, si.position
|
||||||
|
"""))
|
||||||
|
|
||||||
|
rows = result.fetchall()
|
||||||
|
|
||||||
|
current_story_id = None
|
||||||
|
current_time_ms = 0
|
||||||
|
|
||||||
|
for row in rows:
|
||||||
|
item_id, story_id, position, duration = row
|
||||||
|
|
||||||
|
if story_id != current_story_id:
|
||||||
|
current_story_id = story_id
|
||||||
|
current_time_ms = 0
|
||||||
|
|
||||||
|
conn.execute(
|
||||||
|
text("UPDATE story_items SET start_time_ms = :time WHERE id = :id"),
|
||||||
|
{"time": current_time_ms, "id": item_id}
|
||||||
|
)
|
||||||
|
|
||||||
|
current_time_ms += int(duration * 1000) + 200
|
||||||
|
|
||||||
|
conn.commit()
|
||||||
|
|
||||||
|
# Now recreate the table without the position column
|
||||||
|
# 1. Create new table
|
||||||
|
conn.execute(text("""
|
||||||
|
CREATE TABLE story_items_new (
|
||||||
|
id VARCHAR PRIMARY KEY,
|
||||||
|
story_id VARCHAR NOT NULL,
|
||||||
|
generation_id VARCHAR NOT NULL,
|
||||||
|
start_time_ms INTEGER NOT NULL DEFAULT 0,
|
||||||
|
created_at DATETIME,
|
||||||
|
FOREIGN KEY (story_id) REFERENCES stories(id),
|
||||||
|
FOREIGN KEY (generation_id) REFERENCES generations(id)
|
||||||
|
)
|
||||||
|
"""))
|
||||||
|
|
||||||
|
# 2. Copy data
|
||||||
|
conn.execute(text("""
|
||||||
|
INSERT INTO story_items_new (id, story_id, generation_id, start_time_ms, created_at)
|
||||||
|
SELECT id, story_id, generation_id, start_time_ms, created_at FROM story_items
|
||||||
|
"""))
|
||||||
|
|
||||||
|
# 3. Drop old table
|
||||||
|
conn.execute(text("DROP TABLE story_items"))
|
||||||
|
|
||||||
|
# 4. Rename new table
|
||||||
|
conn.execute(text("ALTER TABLE story_items_new RENAME TO story_items"))
|
||||||
|
|
||||||
|
conn.commit()
|
||||||
|
print("Migrated story_items table to use start_time_ms (removed position column)")
|
||||||
|
|
||||||
|
# Migration: Add track column if it doesn't exist
|
||||||
|
# Re-check columns after potential position migration
|
||||||
|
columns = {col['name'] for col in inspector.get_columns('story_items')}
|
||||||
|
if 'track' not in columns:
|
||||||
|
print("Migrating story_items: adding track column")
|
||||||
|
with engine.connect() as conn:
|
||||||
|
conn.execute(text("ALTER TABLE story_items ADD COLUMN track INTEGER NOT NULL DEFAULT 0"))
|
||||||
|
conn.commit()
|
||||||
|
print("Added track column to story_items")
|
||||||
|
|
||||||
|
# Migration: Add trim columns if they don't exist
|
||||||
|
# Re-check columns after potential track migration
|
||||||
|
columns = {col['name'] for col in inspector.get_columns('story_items')}
|
||||||
|
if 'trim_start_ms' not in columns:
|
||||||
|
print("Migrating story_items: adding trim_start_ms column")
|
||||||
|
with engine.connect() as conn:
|
||||||
|
conn.execute(text("ALTER TABLE story_items ADD COLUMN trim_start_ms INTEGER NOT NULL DEFAULT 0"))
|
||||||
|
conn.commit()
|
||||||
|
print("Added trim_start_ms column to story_items")
|
||||||
|
|
||||||
|
columns = {col['name'] for col in inspector.get_columns('story_items')}
|
||||||
|
if 'trim_end_ms' not in columns:
|
||||||
|
print("Migrating story_items: adding trim_end_ms column")
|
||||||
|
with engine.connect() as conn:
|
||||||
|
conn.execute(text("ALTER TABLE story_items ADD COLUMN trim_end_ms INTEGER NOT NULL DEFAULT 0"))
|
||||||
|
conn.commit()
|
||||||
|
print("Added trim_end_ms column to story_items")
|
||||||
|
|
||||||
|
# Migration: Add avatar_path to profiles table
|
||||||
|
if 'profiles' in inspector.get_table_names():
|
||||||
|
columns = {col['name'] for col in inspector.get_columns('profiles')}
|
||||||
|
if 'avatar_path' not in columns:
|
||||||
|
print("Migrating profiles: adding avatar_path column")
|
||||||
|
with engine.connect() as conn:
|
||||||
|
conn.execute(text("ALTER TABLE profiles ADD COLUMN avatar_path VARCHAR"))
|
||||||
|
conn.commit()
|
||||||
|
print("Added avatar_path column to profiles")
|
||||||
|
|
||||||
|
|
||||||
def get_db():
|
def get_db():
|
||||||
"""Get database session (generator for dependency injection)."""
|
"""Get database session (generator for dependency injection)."""
|
||||||
db = SessionLocal()
|
db = SessionLocal()
|
||||||
|
|||||||
@@ -75,6 +75,16 @@ def export_profile_to_zip(profile_id: str, db: Session) -> bytes:
|
|||||||
zip_buffer = io.BytesIO()
|
zip_buffer = io.BytesIO()
|
||||||
|
|
||||||
with zipfile.ZipFile(zip_buffer, 'w', zipfile.ZIP_DEFLATED) as zip_file:
|
with zipfile.ZipFile(zip_buffer, 'w', zipfile.ZIP_DEFLATED) as zip_file:
|
||||||
|
# Check if profile has avatar
|
||||||
|
has_avatar = False
|
||||||
|
if profile.avatar_path:
|
||||||
|
avatar_path = Path(profile.avatar_path)
|
||||||
|
if avatar_path.exists():
|
||||||
|
has_avatar = True
|
||||||
|
# Add avatar to ZIP root with original extension
|
||||||
|
avatar_ext = avatar_path.suffix
|
||||||
|
zip_file.write(avatar_path, f"avatar{avatar_ext}")
|
||||||
|
|
||||||
# Create manifest.json
|
# Create manifest.json
|
||||||
manifest = {
|
manifest = {
|
||||||
"version": "1.0",
|
"version": "1.0",
|
||||||
@@ -82,30 +92,31 @@ def export_profile_to_zip(profile_id: str, db: Session) -> bytes:
|
|||||||
"name": profile.name,
|
"name": profile.name,
|
||||||
"description": profile.description,
|
"description": profile.description,
|
||||||
"language": profile.language,
|
"language": profile.language,
|
||||||
}
|
},
|
||||||
|
"has_avatar": has_avatar,
|
||||||
}
|
}
|
||||||
zip_file.writestr("manifest.json", json.dumps(manifest, indent=2))
|
zip_file.writestr("manifest.json", json.dumps(manifest, indent=2))
|
||||||
|
|
||||||
# Create samples.json mapping
|
# Create samples.json mapping
|
||||||
samples_data = {}
|
samples_data = {}
|
||||||
profile_dir = _get_profiles_dir() / profile_id
|
profile_dir = _get_profiles_dir() / profile_id
|
||||||
|
|
||||||
for sample in samples:
|
for sample in samples:
|
||||||
# Get filename from audio_path (should be {sample_id}.wav)
|
# Get filename from audio_path (should be {sample_id}.wav)
|
||||||
audio_path = Path(sample.audio_path)
|
audio_path = Path(sample.audio_path)
|
||||||
filename = audio_path.name
|
filename = audio_path.name
|
||||||
|
|
||||||
# Read audio file
|
# Read audio file
|
||||||
if not audio_path.exists():
|
if not audio_path.exists():
|
||||||
raise ValueError(f"Audio file not found: {audio_path}")
|
raise ValueError(f"Audio file not found: {audio_path}")
|
||||||
|
|
||||||
# Add to samples directory in ZIP
|
# Add to samples directory in ZIP
|
||||||
zip_path = f"samples/{filename}"
|
zip_path = f"samples/{filename}"
|
||||||
zip_file.write(audio_path, zip_path)
|
zip_file.write(audio_path, zip_path)
|
||||||
|
|
||||||
# Map filename to reference text
|
# Map filename to reference text
|
||||||
samples_data[filename] = sample.reference_text
|
samples_data[filename] = sample.reference_text
|
||||||
|
|
||||||
zip_file.writestr("samples.json", json.dumps(samples_data, indent=2))
|
zip_file.writestr("samples.json", json.dumps(samples_data, indent=2))
|
||||||
|
|
||||||
zip_buffer.seek(0)
|
zip_buffer.seek(0)
|
||||||
@@ -168,11 +179,31 @@ async def import_profile_from_zip(file_bytes: bytes, db: Session) -> VoiceProfil
|
|||||||
)
|
)
|
||||||
|
|
||||||
profile = await create_profile(profile_create, db)
|
profile = await create_profile(profile_create, db)
|
||||||
|
|
||||||
# Extract and add samples
|
# Extract and add samples
|
||||||
profile_dir = _get_profiles_dir() / profile.id
|
profile_dir = _get_profiles_dir() / profile.id
|
||||||
profile_dir.mkdir(parents=True, exist_ok=True)
|
profile_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# Handle avatar if present
|
||||||
|
avatar_files = [f for f in namelist if f.startswith("avatar.")]
|
||||||
|
if avatar_files:
|
||||||
|
try:
|
||||||
|
avatar_file = avatar_files[0]
|
||||||
|
# Extract to temporary file
|
||||||
|
import tempfile
|
||||||
|
with tempfile.NamedTemporaryFile(suffix=Path(avatar_file).suffix, delete=False) as tmp:
|
||||||
|
tmp.write(zip_file.read(avatar_file))
|
||||||
|
tmp_path = tmp.name
|
||||||
|
|
||||||
|
try:
|
||||||
|
from .profiles import upload_avatar
|
||||||
|
await upload_avatar(profile.id, tmp_path, db)
|
||||||
|
finally:
|
||||||
|
Path(tmp_path).unlink(missing_ok=True)
|
||||||
|
except Exception as e:
|
||||||
|
# Avatar import is optional - continue even if it fails
|
||||||
|
pass
|
||||||
|
|
||||||
for filename, reference_text in samples_data.items():
|
for filename, reference_text in samples_data.items():
|
||||||
# Validate filename
|
# Validate filename
|
||||||
if not filename.endswith('.wav'):
|
if not filename.endswith('.wav'):
|
||||||
|
|||||||
+466
-40
@@ -11,6 +11,7 @@ from fastapi.staticfiles import StaticFiles
|
|||||||
from sqlalchemy.orm import Session
|
from sqlalchemy.orm import Session
|
||||||
from typing import List, Optional
|
from typing import List, Optional
|
||||||
from datetime import datetime
|
from datetime import datetime
|
||||||
|
import asyncio
|
||||||
import uvicorn
|
import uvicorn
|
||||||
import argparse
|
import argparse
|
||||||
import torch
|
import torch
|
||||||
@@ -18,16 +19,21 @@ import tempfile
|
|||||||
import io
|
import io
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
import uuid
|
import uuid
|
||||||
|
import asyncio
|
||||||
|
import signal
|
||||||
|
import os
|
||||||
|
|
||||||
from . import database, models, profiles, history, tts, transcribe, config, export_import, channels
|
from . import database, models, profiles, history, tts, transcribe, config, export_import, channels, stories, __version__
|
||||||
from .database import get_db, Generation as DBGeneration, VoiceProfile as DBVoiceProfile
|
from .database import get_db, Generation as DBGeneration, VoiceProfile as DBVoiceProfile
|
||||||
from .utils.progress import get_progress_manager
|
from .utils.progress import get_progress_manager
|
||||||
from .utils.tasks import get_task_manager
|
from .utils.tasks import get_task_manager
|
||||||
|
from .utils.cache import clear_voice_prompt_cache
|
||||||
|
from .platform_detect import get_backend_type
|
||||||
|
|
||||||
app = FastAPI(
|
app = FastAPI(
|
||||||
title="voicebox API",
|
title="voicebox API",
|
||||||
description="Production-quality Qwen3-TTS voice cloning API",
|
description="Production-quality Qwen3-TTS voice cloning API",
|
||||||
version="0.1.0",
|
version=__version__,
|
||||||
)
|
)
|
||||||
|
|
||||||
# CORS middleware
|
# CORS middleware
|
||||||
@@ -47,23 +53,43 @@ app.add_middleware(
|
|||||||
@app.get("/")
|
@app.get("/")
|
||||||
async def root():
|
async def root():
|
||||||
"""Root endpoint."""
|
"""Root endpoint."""
|
||||||
return {"message": "voicebox API", "version": "0.1.4"}
|
return {"message": "voicebox API", "version": __version__}
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/shutdown")
|
||||||
|
async def shutdown():
|
||||||
|
"""Gracefully shutdown the server."""
|
||||||
|
async def shutdown_async():
|
||||||
|
await asyncio.sleep(0.1) # Give response time to send
|
||||||
|
os.kill(os.getpid(), signal.SIGTERM)
|
||||||
|
|
||||||
|
asyncio.create_task(shutdown_async())
|
||||||
|
return {"message": "Shutting down..."}
|
||||||
|
|
||||||
|
|
||||||
@app.get("/health", response_model=models.HealthResponse)
|
@app.get("/health", response_model=models.HealthResponse)
|
||||||
async def health():
|
async def health():
|
||||||
"""Health check endpoint."""
|
"""Health check endpoint."""
|
||||||
from huggingface_hub import hf_hub_download
|
from huggingface_hub import hf_hub_download, constants as hf_constants
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
import os
|
import os
|
||||||
|
|
||||||
tts_model = tts.get_tts_model()
|
tts_model = tts.get_tts_model()
|
||||||
|
backend_type = get_backend_type()
|
||||||
|
|
||||||
# Check for GPU availability (CUDA or MPS)
|
# Check for GPU availability (CUDA or MPS)
|
||||||
has_cuda = torch.cuda.is_available()
|
has_cuda = torch.cuda.is_available()
|
||||||
has_mps = hasattr(torch.backends, 'mps') and torch.backends.mps.is_available()
|
has_mps = hasattr(torch.backends, 'mps') and torch.backends.mps.is_available()
|
||||||
gpu_available = has_cuda or has_mps
|
gpu_available = has_cuda or has_mps
|
||||||
|
|
||||||
|
gpu_type = None
|
||||||
|
if has_cuda:
|
||||||
|
gpu_type = f"CUDA ({torch.cuda.get_device_name(0)})"
|
||||||
|
elif has_mps:
|
||||||
|
gpu_type = "MPS (Apple Silicon)"
|
||||||
|
elif backend_type == "mlx":
|
||||||
|
gpu_type = "Metal (Apple Silicon via MLX)"
|
||||||
|
|
||||||
vram_used = None
|
vram_used = None
|
||||||
if has_cuda:
|
if has_cuda:
|
||||||
vram_used = torch.cuda.memory_allocated() / 1024 / 1024 # MB
|
vram_used = torch.cuda.memory_allocated() / 1024 / 1024 # MB
|
||||||
@@ -90,7 +116,11 @@ async def health():
|
|||||||
model_downloaded = None
|
model_downloaded = None
|
||||||
try:
|
try:
|
||||||
# Check if the default model (1.7B) is cached
|
# Check if the default model (1.7B) is cached
|
||||||
default_model_id = "Qwen/Qwen3-TTS-12Hz-1.7B-Base"
|
# Use different model IDs based on backend
|
||||||
|
if backend_type == "mlx":
|
||||||
|
default_model_id = "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16"
|
||||||
|
else:
|
||||||
|
default_model_id = "Qwen/Qwen3-TTS-12Hz-1.7B-Base"
|
||||||
|
|
||||||
# Method 1: Try scan_cache_dir if available
|
# Method 1: Try scan_cache_dir if available
|
||||||
try:
|
try:
|
||||||
@@ -101,15 +131,16 @@ async def health():
|
|||||||
model_downloaded = True
|
model_downloaded = True
|
||||||
break
|
break
|
||||||
except (ImportError, Exception):
|
except (ImportError, Exception):
|
||||||
# Method 2: Check cache directory
|
# Method 2: Check cache directory (using HuggingFace's OS-specific cache location)
|
||||||
cache_dir = os.path.expanduser("~/.cache/huggingface/hub")
|
cache_dir = hf_constants.HF_HUB_CACHE
|
||||||
repo_cache = Path(cache_dir) / "models--" + default_model_id.replace("/", "--")
|
repo_cache = Path(cache_dir) / ("models--" + default_model_id.replace("/", "--"))
|
||||||
if repo_cache.exists():
|
if repo_cache.exists():
|
||||||
has_model_files = (
|
has_model_files = (
|
||||||
any(repo_cache.rglob("*.bin")) or
|
any(repo_cache.rglob("*.bin")) or
|
||||||
any(repo_cache.rglob("*.safetensors")) or
|
any(repo_cache.rglob("*.safetensors")) or
|
||||||
any(repo_cache.rglob("*.pt")) or
|
any(repo_cache.rglob("*.pt")) or
|
||||||
any(repo_cache.rglob("*.pth"))
|
any(repo_cache.rglob("*.pth")) or
|
||||||
|
any(repo_cache.rglob("*.npz")) # MLX models may use npz
|
||||||
)
|
)
|
||||||
model_downloaded = has_model_files
|
model_downloaded = has_model_files
|
||||||
except Exception:
|
except Exception:
|
||||||
@@ -121,7 +152,9 @@ async def health():
|
|||||||
model_downloaded=model_downloaded,
|
model_downloaded=model_downloaded,
|
||||||
model_size=model_size,
|
model_size=model_size,
|
||||||
gpu_available=gpu_available,
|
gpu_available=gpu_available,
|
||||||
|
gpu_type=gpu_type,
|
||||||
vram_used_mb=vram_used,
|
vram_used_mb=vram_used,
|
||||||
|
backend_type=backend_type,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -261,6 +294,74 @@ async def delete_profile_sample(
|
|||||||
return {"message": "Sample deleted successfully"}
|
return {"message": "Sample deleted successfully"}
|
||||||
|
|
||||||
|
|
||||||
|
@app.put("/profiles/samples/{sample_id}", response_model=models.ProfileSampleResponse)
|
||||||
|
async def update_profile_sample(
|
||||||
|
sample_id: str,
|
||||||
|
data: models.ProfileSampleUpdate,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Update a profile sample's reference text."""
|
||||||
|
sample = await profiles.update_profile_sample(sample_id, data.reference_text, db)
|
||||||
|
if not sample:
|
||||||
|
raise HTTPException(status_code=404, detail="Sample not found")
|
||||||
|
return sample
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/profiles/{profile_id}/avatar", response_model=models.VoiceProfileResponse)
|
||||||
|
async def upload_profile_avatar(
|
||||||
|
profile_id: str,
|
||||||
|
file: UploadFile = File(...),
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Upload or update avatar image for a profile."""
|
||||||
|
# Save uploaded file to temp location
|
||||||
|
with tempfile.NamedTemporaryFile(delete=False, suffix=Path(file.filename).suffix) as tmp:
|
||||||
|
content = await file.read()
|
||||||
|
tmp.write(content)
|
||||||
|
tmp_path = tmp.name
|
||||||
|
|
||||||
|
try:
|
||||||
|
profile = await profiles.upload_avatar(profile_id, tmp_path, db)
|
||||||
|
return profile
|
||||||
|
except ValueError as e:
|
||||||
|
raise HTTPException(status_code=400, detail=str(e))
|
||||||
|
finally:
|
||||||
|
# Clean up temp file
|
||||||
|
Path(tmp_path).unlink(missing_ok=True)
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/profiles/{profile_id}/avatar")
|
||||||
|
async def get_profile_avatar(
|
||||||
|
profile_id: str,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Get avatar image for a profile."""
|
||||||
|
profile = await profiles.get_profile(profile_id, db)
|
||||||
|
if not profile:
|
||||||
|
raise HTTPException(status_code=404, detail="Profile not found")
|
||||||
|
|
||||||
|
if not profile.avatar_path:
|
||||||
|
raise HTTPException(status_code=404, detail="No avatar found for this profile")
|
||||||
|
|
||||||
|
avatar_path = Path(profile.avatar_path)
|
||||||
|
if not avatar_path.exists():
|
||||||
|
raise HTTPException(status_code=404, detail="Avatar file not found")
|
||||||
|
|
||||||
|
return FileResponse(avatar_path)
|
||||||
|
|
||||||
|
|
||||||
|
@app.delete("/profiles/{profile_id}/avatar")
|
||||||
|
async def delete_profile_avatar(
|
||||||
|
profile_id: str,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Delete avatar image for a profile."""
|
||||||
|
success = await profiles.delete_avatar(profile_id, db)
|
||||||
|
if not success:
|
||||||
|
raise HTTPException(status_code=404, detail="Profile not found or no avatar to delete")
|
||||||
|
return {"message": "Avatar deleted successfully"}
|
||||||
|
|
||||||
|
|
||||||
@app.get("/profiles/{profile_id}/export")
|
@app.get("/profiles/{profile_id}/export")
|
||||||
async def export_profile(
|
async def export_profile(
|
||||||
profile_id: str,
|
profile_id: str,
|
||||||
@@ -451,6 +552,36 @@ async def generate_speech(
|
|||||||
tts_model = tts.get_tts_model()
|
tts_model = tts.get_tts_model()
|
||||||
# Load the requested model size if different from current (async to not block)
|
# Load the requested model size if different from current (async to not block)
|
||||||
model_size = data.model_size or "1.7B"
|
model_size = data.model_size or "1.7B"
|
||||||
|
|
||||||
|
# Check if model needs to be downloaded first
|
||||||
|
model_path = tts_model._get_model_path(model_size)
|
||||||
|
if model_path.startswith("Qwen/"):
|
||||||
|
# Model not cached - check if it exists remotely or needs download
|
||||||
|
from huggingface_hub import constants as hf_constants
|
||||||
|
repo_cache = Path(hf_constants.HF_HUB_CACHE) / ("models--" + model_path.replace("/", "--"))
|
||||||
|
if not repo_cache.exists():
|
||||||
|
# Start download in background
|
||||||
|
model_name = f"qwen-tts-{model_size}"
|
||||||
|
|
||||||
|
async def download_model_background():
|
||||||
|
try:
|
||||||
|
await tts_model.load_model_async(model_size)
|
||||||
|
except Exception as e:
|
||||||
|
task_manager.error_download(model_name, str(e))
|
||||||
|
|
||||||
|
task_manager.start_download(model_name)
|
||||||
|
asyncio.create_task(download_model_background())
|
||||||
|
|
||||||
|
# Return 202 Accepted with download info
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=202,
|
||||||
|
detail={
|
||||||
|
"message": f"Model {model_size} is being downloaded. Please wait and try again.",
|
||||||
|
"model_name": model_name,
|
||||||
|
"downloading": True
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
await tts_model.load_model_async(model_size)
|
await tts_model.load_model_async(model_size)
|
||||||
audio, sample_rate = await tts_model.generate(
|
audio, sample_rate = await tts_model.generate(
|
||||||
data.text,
|
data.text,
|
||||||
@@ -684,6 +815,37 @@ async def transcribe_audio(
|
|||||||
|
|
||||||
# Transcribe
|
# Transcribe
|
||||||
whisper_model = transcribe.get_whisper_model()
|
whisper_model = transcribe.get_whisper_model()
|
||||||
|
|
||||||
|
# Check if Whisper model is downloaded (uses default size "base")
|
||||||
|
model_size = whisper_model.model_size
|
||||||
|
model_name = f"openai/whisper-{model_size}"
|
||||||
|
|
||||||
|
# Check if model is cached
|
||||||
|
from huggingface_hub import constants as hf_constants
|
||||||
|
repo_cache = Path(hf_constants.HF_HUB_CACHE) / ("models--" + model_name.replace("/", "--"))
|
||||||
|
if not repo_cache.exists():
|
||||||
|
# Start download in background
|
||||||
|
progress_model_name = f"whisper-{model_size}"
|
||||||
|
|
||||||
|
async def download_whisper_background():
|
||||||
|
try:
|
||||||
|
await whisper_model.load_model_async(model_size)
|
||||||
|
except Exception as e:
|
||||||
|
get_task_manager().error_download(progress_model_name, str(e))
|
||||||
|
|
||||||
|
get_task_manager().start_download(progress_model_name)
|
||||||
|
asyncio.create_task(download_whisper_background())
|
||||||
|
|
||||||
|
# Return 202 Accepted
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=202,
|
||||||
|
detail={
|
||||||
|
"message": f"Whisper model {model_size} is being downloaded. Please wait and try again.",
|
||||||
|
"model_name": progress_model_name,
|
||||||
|
"downloading": True
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
text = await whisper_model.transcribe(tmp_path, language)
|
text = await whisper_model.transcribe(tmp_path, language)
|
||||||
|
|
||||||
return models.TranscriptionResponse(
|
return models.TranscriptionResponse(
|
||||||
@@ -698,6 +860,209 @@ async def transcribe_audio(
|
|||||||
Path(tmp_path).unlink(missing_ok=True)
|
Path(tmp_path).unlink(missing_ok=True)
|
||||||
|
|
||||||
|
|
||||||
|
# ============================================
|
||||||
|
# STORY ENDPOINTS
|
||||||
|
# ============================================
|
||||||
|
|
||||||
|
@app.get("/stories", response_model=List[models.StoryResponse])
|
||||||
|
async def list_stories(db: Session = Depends(get_db)):
|
||||||
|
"""List all stories."""
|
||||||
|
return await stories.list_stories(db)
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/stories", response_model=models.StoryResponse)
|
||||||
|
async def create_story(
|
||||||
|
data: models.StoryCreate,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Create a new story."""
|
||||||
|
try:
|
||||||
|
return await stories.create_story(data, db)
|
||||||
|
except Exception as e:
|
||||||
|
raise HTTPException(status_code=400, detail=str(e))
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/stories/{story_id}", response_model=models.StoryDetailResponse)
|
||||||
|
async def get_story(
|
||||||
|
story_id: str,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Get a story with all its items."""
|
||||||
|
story = await stories.get_story(story_id, db)
|
||||||
|
if not story:
|
||||||
|
raise HTTPException(status_code=404, detail="Story not found")
|
||||||
|
return story
|
||||||
|
|
||||||
|
|
||||||
|
@app.put("/stories/{story_id}", response_model=models.StoryResponse)
|
||||||
|
async def update_story(
|
||||||
|
story_id: str,
|
||||||
|
data: models.StoryCreate,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Update a story."""
|
||||||
|
story = await stories.update_story(story_id, data, db)
|
||||||
|
if not story:
|
||||||
|
raise HTTPException(status_code=404, detail="Story not found")
|
||||||
|
return story
|
||||||
|
|
||||||
|
|
||||||
|
@app.delete("/stories/{story_id}")
|
||||||
|
async def delete_story(
|
||||||
|
story_id: str,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Delete a story."""
|
||||||
|
success = await stories.delete_story(story_id, db)
|
||||||
|
if not success:
|
||||||
|
raise HTTPException(status_code=404, detail="Story not found")
|
||||||
|
return {"message": "Story deleted successfully"}
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/stories/{story_id}/items", response_model=models.StoryItemDetail)
|
||||||
|
async def add_story_item(
|
||||||
|
story_id: str,
|
||||||
|
data: models.StoryItemCreate,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Add a generation to a story."""
|
||||||
|
item = await stories.add_item_to_story(story_id, data, db)
|
||||||
|
if not item:
|
||||||
|
raise HTTPException(status_code=404, detail="Story or generation not found")
|
||||||
|
return item
|
||||||
|
|
||||||
|
|
||||||
|
@app.delete("/stories/{story_id}/items/{item_id}")
|
||||||
|
async def remove_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Remove a story item from a story."""
|
||||||
|
success = await stories.remove_item_from_story(story_id, item_id, db)
|
||||||
|
if not success:
|
||||||
|
raise HTTPException(status_code=404, detail="Story item not found")
|
||||||
|
return {"message": "Item removed successfully"}
|
||||||
|
|
||||||
|
|
||||||
|
@app.put("/stories/{story_id}/items/times")
|
||||||
|
async def update_story_item_times(
|
||||||
|
story_id: str,
|
||||||
|
data: models.StoryItemBatchUpdate,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Update story item timecodes."""
|
||||||
|
success = await stories.update_story_item_times(story_id, data, db)
|
||||||
|
if not success:
|
||||||
|
raise HTTPException(status_code=400, detail="Invalid timecode update request")
|
||||||
|
return {"message": "Item timecodes updated successfully"}
|
||||||
|
|
||||||
|
|
||||||
|
@app.put("/stories/{story_id}/items/reorder", response_model=List[models.StoryItemDetail])
|
||||||
|
async def reorder_story_items(
|
||||||
|
story_id: str,
|
||||||
|
data: models.StoryItemReorder,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Reorder story items and recalculate timecodes."""
|
||||||
|
items = await stories.reorder_story_items(story_id, data.generation_ids, db)
|
||||||
|
if items is None:
|
||||||
|
raise HTTPException(status_code=400, detail="Invalid reorder request - ensure all generation IDs belong to this story")
|
||||||
|
return items
|
||||||
|
|
||||||
|
|
||||||
|
@app.put("/stories/{story_id}/items/{item_id}/move", response_model=models.StoryItemDetail)
|
||||||
|
async def move_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: models.StoryItemMove,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Move a story item (update position and/or track)."""
|
||||||
|
item = await stories.move_story_item(story_id, item_id, data, db)
|
||||||
|
if item is None:
|
||||||
|
raise HTTPException(status_code=404, detail="Story item not found")
|
||||||
|
return item
|
||||||
|
|
||||||
|
|
||||||
|
@app.put("/stories/{story_id}/items/{item_id}/trim", response_model=models.StoryItemDetail)
|
||||||
|
async def trim_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: models.StoryItemTrim,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Trim a story item (update trim_start_ms and trim_end_ms)."""
|
||||||
|
item = await stories.trim_story_item(story_id, item_id, data, db)
|
||||||
|
if item is None:
|
||||||
|
raise HTTPException(status_code=404, detail="Story item not found or invalid trim values")
|
||||||
|
return item
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/stories/{story_id}/items/{item_id}/split", response_model=List[models.StoryItemDetail])
|
||||||
|
async def split_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: models.StoryItemSplit,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Split a story item at a given time, creating two clips."""
|
||||||
|
items = await stories.split_story_item(story_id, item_id, data, db)
|
||||||
|
if items is None:
|
||||||
|
raise HTTPException(status_code=404, detail="Story item not found or invalid split point")
|
||||||
|
return items
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/stories/{story_id}/items/{item_id}/duplicate", response_model=models.StoryItemDetail)
|
||||||
|
async def duplicate_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Duplicate a story item, creating a copy with all properties."""
|
||||||
|
item = await stories.duplicate_story_item(story_id, item_id, db)
|
||||||
|
if item is None:
|
||||||
|
raise HTTPException(status_code=404, detail="Story item not found")
|
||||||
|
return item
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/stories/{story_id}/export-audio")
|
||||||
|
async def export_story_audio(
|
||||||
|
story_id: str,
|
||||||
|
db: Session = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Export story as single mixed audio file with timecode-based mixing."""
|
||||||
|
try:
|
||||||
|
# Get story to create filename
|
||||||
|
story = db.query(database.Story).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
raise HTTPException(status_code=404, detail="Story not found")
|
||||||
|
|
||||||
|
# Export audio
|
||||||
|
audio_bytes = await stories.export_story_audio(story_id, db)
|
||||||
|
if not audio_bytes:
|
||||||
|
raise HTTPException(status_code=400, detail="Story has no audio items")
|
||||||
|
|
||||||
|
# Create safe filename
|
||||||
|
safe_name = "".join(c for c in story.name if c.isalnum() or c in (' ', '-', '_')).strip()
|
||||||
|
if not safe_name:
|
||||||
|
safe_name = "story"
|
||||||
|
filename = f"{safe_name}.wav"
|
||||||
|
|
||||||
|
# Return as streaming response
|
||||||
|
return StreamingResponse(
|
||||||
|
io.BytesIO(audio_bytes),
|
||||||
|
media_type="audio/wav",
|
||||||
|
headers={
|
||||||
|
"Content-Disposition": f'attachment; filename="{filename}"'
|
||||||
|
}
|
||||||
|
)
|
||||||
|
except HTTPException:
|
||||||
|
raise
|
||||||
|
except Exception as e:
|
||||||
|
raise HTTPException(status_code=500, detail=str(e))
|
||||||
|
|
||||||
|
|
||||||
# ============================================
|
# ============================================
|
||||||
# FILE SERVING
|
# FILE SERVING
|
||||||
# ============================================
|
# ============================================
|
||||||
@@ -791,10 +1156,12 @@ async def get_model_progress(model_name: str):
|
|||||||
@app.get("/models/status", response_model=models.ModelStatusListResponse)
|
@app.get("/models/status", response_model=models.ModelStatusListResponse)
|
||||||
async def get_model_status():
|
async def get_model_status():
|
||||||
"""Get status of all available models."""
|
"""Get status of all available models."""
|
||||||
from huggingface_hub import hf_hub_download
|
from huggingface_hub import hf_hub_download, constants as hf_constants
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
import os
|
import os
|
||||||
|
|
||||||
|
backend_type = get_backend_type()
|
||||||
|
|
||||||
# Try to import scan_cache_dir (might not be available in older versions)
|
# Try to import scan_cache_dir (might not be available in older versions)
|
||||||
try:
|
try:
|
||||||
from huggingface_hub import scan_cache_dir
|
from huggingface_hub import scan_cache_dir
|
||||||
@@ -806,7 +1173,7 @@ async def get_model_status():
|
|||||||
"""Check if TTS model is loaded with specific size."""
|
"""Check if TTS model is loaded with specific size."""
|
||||||
try:
|
try:
|
||||||
tts_model = tts.get_tts_model()
|
tts_model = tts.get_tts_model()
|
||||||
return tts_model.is_loaded() and tts_model.model_size == model_size
|
return tts_model.is_loaded() and getattr(tts_model, 'model_size', None) == model_size
|
||||||
except Exception:
|
except Exception:
|
||||||
return False
|
return False
|
||||||
|
|
||||||
@@ -814,50 +1181,66 @@ async def get_model_status():
|
|||||||
"""Check if Whisper model is loaded with specific size."""
|
"""Check if Whisper model is loaded with specific size."""
|
||||||
try:
|
try:
|
||||||
whisper_model = transcribe.get_whisper_model()
|
whisper_model = transcribe.get_whisper_model()
|
||||||
return whisper_model.is_loaded() and whisper_model.model_size == model_size
|
return whisper_model.is_loaded() and getattr(whisper_model, 'model_size', None) == model_size
|
||||||
except Exception:
|
except Exception:
|
||||||
return False
|
return False
|
||||||
|
|
||||||
|
# Use backend-specific model IDs
|
||||||
|
if backend_type == "mlx":
|
||||||
|
tts_1_7b_id = "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16"
|
||||||
|
tts_0_6b_id = "mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16" # Fallback to 1.7B
|
||||||
|
whisper_base_id = "mlx-community/whisper-base"
|
||||||
|
whisper_small_id = "mlx-community/whisper-small"
|
||||||
|
whisper_medium_id = "mlx-community/whisper-medium"
|
||||||
|
whisper_large_id = "mlx-community/whisper-large"
|
||||||
|
else:
|
||||||
|
tts_1_7b_id = "Qwen/Qwen3-TTS-12Hz-1.7B-Base"
|
||||||
|
tts_0_6b_id = "Qwen/Qwen3-TTS-12Hz-0.6B-Base"
|
||||||
|
whisper_base_id = "openai/whisper-base"
|
||||||
|
whisper_small_id = "openai/whisper-small"
|
||||||
|
whisper_medium_id = "openai/whisper-medium"
|
||||||
|
whisper_large_id = "openai/whisper-large"
|
||||||
|
|
||||||
model_configs = [
|
model_configs = [
|
||||||
{
|
{
|
||||||
"model_name": "qwen-tts-1.7B",
|
"model_name": "qwen-tts-1.7B",
|
||||||
"display_name": "Qwen TTS 1.7B",
|
"display_name": "Qwen TTS 1.7B",
|
||||||
"hf_repo_id": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
|
"hf_repo_id": tts_1_7b_id,
|
||||||
"model_size": "1.7B",
|
"model_size": "1.7B",
|
||||||
"check_loaded": lambda: check_tts_loaded("1.7B"),
|
"check_loaded": lambda: check_tts_loaded("1.7B"),
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"model_name": "qwen-tts-0.6B",
|
"model_name": "qwen-tts-0.6B",
|
||||||
"display_name": "Qwen TTS 0.6B",
|
"display_name": "Qwen TTS 0.6B",
|
||||||
"hf_repo_id": "Qwen/Qwen3-TTS-12Hz-0.6B-Base",
|
"hf_repo_id": tts_0_6b_id,
|
||||||
"model_size": "0.6B",
|
"model_size": "0.6B",
|
||||||
"check_loaded": lambda: check_tts_loaded("0.6B"),
|
"check_loaded": lambda: check_tts_loaded("0.6B"),
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"model_name": "whisper-base",
|
"model_name": "whisper-base",
|
||||||
"display_name": "Whisper Base",
|
"display_name": "Whisper Base",
|
||||||
"hf_repo_id": "openai/whisper-base",
|
"hf_repo_id": whisper_base_id,
|
||||||
"model_size": "base",
|
"model_size": "base",
|
||||||
"check_loaded": lambda: check_whisper_loaded("base"),
|
"check_loaded": lambda: check_whisper_loaded("base"),
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"model_name": "whisper-small",
|
"model_name": "whisper-small",
|
||||||
"display_name": "Whisper Small",
|
"display_name": "Whisper Small",
|
||||||
"hf_repo_id": "openai/whisper-small",
|
"hf_repo_id": whisper_small_id,
|
||||||
"model_size": "small",
|
"model_size": "small",
|
||||||
"check_loaded": lambda: check_whisper_loaded("small"),
|
"check_loaded": lambda: check_whisper_loaded("small"),
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"model_name": "whisper-medium",
|
"model_name": "whisper-medium",
|
||||||
"display_name": "Whisper Medium",
|
"display_name": "Whisper Medium",
|
||||||
"hf_repo_id": "openai/whisper-medium",
|
"hf_repo_id": whisper_medium_id,
|
||||||
"model_size": "medium",
|
"model_size": "medium",
|
||||||
"check_loaded": lambda: check_whisper_loaded("medium"),
|
"check_loaded": lambda: check_whisper_loaded("medium"),
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"model_name": "whisper-large",
|
"model_name": "whisper-large",
|
||||||
"display_name": "Whisper Large",
|
"display_name": "Whisper Large",
|
||||||
"hf_repo_id": "openai/whisper-large",
|
"hf_repo_id": whisper_large_id,
|
||||||
"model_size": "large",
|
"model_size": "large",
|
||||||
"check_loaded": lambda: check_whisper_loaded("large"),
|
"check_loaded": lambda: check_whisper_loaded("large"),
|
||||||
},
|
},
|
||||||
@@ -894,19 +1277,21 @@ async def get_model_status():
|
|||||||
pass
|
pass
|
||||||
break
|
break
|
||||||
|
|
||||||
# Method 2: Fallback to checking cache directory directly
|
# Method 2: Fallback to checking cache directory directly (using HuggingFace's OS-specific cache location)
|
||||||
if not downloaded:
|
if not downloaded:
|
||||||
try:
|
try:
|
||||||
cache_dir = os.path.expanduser("~/.cache/huggingface/hub")
|
cache_dir = hf_constants.HF_HUB_CACHE
|
||||||
repo_cache = Path(cache_dir) / "models--" + config["hf_repo_id"].replace("/", "--")
|
repo_cache = Path(cache_dir) / ("models--" + config["hf_repo_id"].replace("/", "--"))
|
||||||
|
|
||||||
if repo_cache.exists():
|
if repo_cache.exists():
|
||||||
# Check for model files (bin, safetensors, or other common model files)
|
# Check for model files (bin, safetensors, or other common model files)
|
||||||
|
# MLX models may use .npz or .safetensors
|
||||||
has_model_files = (
|
has_model_files = (
|
||||||
any(repo_cache.rglob("*.bin")) or
|
any(repo_cache.rglob("*.bin")) or
|
||||||
any(repo_cache.rglob("*.safetensors")) or
|
any(repo_cache.rglob("*.safetensors")) or
|
||||||
any(repo_cache.rglob("*.pt")) or
|
any(repo_cache.rglob("*.pt")) or
|
||||||
any(repo_cache.rglob("*.pth")) or
|
any(repo_cache.rglob("*.pth")) or
|
||||||
|
any(repo_cache.rglob("*.npz")) or
|
||||||
any(repo_cache.rglob("model.safetensors.index.json")) or
|
any(repo_cache.rglob("model.safetensors.index.json")) or
|
||||||
any(repo_cache.rglob("pytorch_model.bin.index.json"))
|
any(repo_cache.rglob("pytorch_model.bin.index.json"))
|
||||||
)
|
)
|
||||||
@@ -1006,22 +1391,26 @@ async def trigger_model_download(request: models.ModelDownloadRequest):
|
|||||||
|
|
||||||
config = model_configs[request.model_name]
|
config = model_configs[request.model_name]
|
||||||
|
|
||||||
try:
|
async def download_in_background():
|
||||||
# Start tracking download
|
"""Download model in background without blocking the HTTP request."""
|
||||||
task_manager.start_download(request.model_name)
|
try:
|
||||||
|
# Call the load function (which may be async)
|
||||||
# Trigger download by loading the model (which will download if not cached)
|
result = config["load_func"]()
|
||||||
# Run in background to avoid blocking
|
# If it's a coroutine, await it
|
||||||
await asyncio.to_thread(config["load_func"])
|
if asyncio.iscoroutine(result):
|
||||||
|
await result
|
||||||
# Mark download as complete
|
task_manager.complete_download(request.model_name)
|
||||||
task_manager.complete_download(request.model_name)
|
except Exception as e:
|
||||||
|
task_manager.error_download(request.model_name, str(e))
|
||||||
return {"message": f"Model {request.model_name} download started"}
|
|
||||||
except Exception as e:
|
# Start tracking download
|
||||||
# Mark download as failed
|
task_manager.start_download(request.model_name)
|
||||||
task_manager.error_download(request.model_name, str(e))
|
|
||||||
raise HTTPException(status_code=500, detail=str(e))
|
# Start download in background task (don't await)
|
||||||
|
asyncio.create_task(download_in_background())
|
||||||
|
|
||||||
|
# Return immediately - frontend should poll progress endpoint
|
||||||
|
return {"message": f"Model {request.model_name} download started"}
|
||||||
|
|
||||||
|
|
||||||
@app.delete("/models/{model_name}")
|
@app.delete("/models/{model_name}")
|
||||||
@@ -1029,6 +1418,7 @@ async def delete_model(model_name: str):
|
|||||||
"""Delete a downloaded model from the HuggingFace cache."""
|
"""Delete a downloaded model from the HuggingFace cache."""
|
||||||
import shutil
|
import shutil
|
||||||
import os
|
import os
|
||||||
|
from huggingface_hub import constants as hf_constants
|
||||||
|
|
||||||
# Map model names to HuggingFace repo IDs
|
# Map model names to HuggingFace repo IDs
|
||||||
model_configs = {
|
model_configs = {
|
||||||
@@ -1081,8 +1471,8 @@ async def delete_model(model_name: str):
|
|||||||
if whisper_model.is_loaded() and whisper_model.model_size == config["model_size"]:
|
if whisper_model.is_loaded() and whisper_model.model_size == config["model_size"]:
|
||||||
transcribe.unload_whisper_model()
|
transcribe.unload_whisper_model()
|
||||||
|
|
||||||
# Find and delete the cache directory
|
# Find and delete the cache directory (using HuggingFace's OS-specific cache location)
|
||||||
cache_dir = os.path.expanduser("~/.cache/huggingface/hub")
|
cache_dir = hf_constants.HF_HUB_CACHE
|
||||||
repo_cache_dir = Path(cache_dir) / ("models--" + hf_repo_id.replace("/", "--"))
|
repo_cache_dir = Path(cache_dir) / ("models--" + hf_repo_id.replace("/", "--"))
|
||||||
|
|
||||||
# Check if the cache directory exists
|
# Check if the cache directory exists
|
||||||
@@ -1106,6 +1496,19 @@ async def delete_model(model_name: str):
|
|||||||
raise HTTPException(status_code=500, detail=f"Failed to delete model: {str(e)}")
|
raise HTTPException(status_code=500, detail=f"Failed to delete model: {str(e)}")
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/cache/clear")
|
||||||
|
async def clear_cache():
|
||||||
|
"""Clear all voice prompt caches (memory and disk)."""
|
||||||
|
try:
|
||||||
|
deleted_count = clear_voice_prompt_cache()
|
||||||
|
return {
|
||||||
|
"message": f"Voice prompt cache cleared successfully",
|
||||||
|
"files_deleted": deleted_count,
|
||||||
|
}
|
||||||
|
except Exception as e:
|
||||||
|
raise HTTPException(status_code=500, detail=f"Failed to clear cache: {str(e)}")
|
||||||
|
|
||||||
|
|
||||||
# ============================================
|
# ============================================
|
||||||
# TASK MANAGEMENT
|
# TASK MANAGEMENT
|
||||||
# ============================================
|
# ============================================
|
||||||
@@ -1178,10 +1581,13 @@ async def get_active_tasks():
|
|||||||
|
|
||||||
def _get_gpu_status() -> str:
|
def _get_gpu_status() -> str:
|
||||||
"""Get GPU availability status."""
|
"""Get GPU availability status."""
|
||||||
|
backend_type = get_backend_type()
|
||||||
if torch.cuda.is_available():
|
if torch.cuda.is_available():
|
||||||
return f"CUDA ({torch.cuda.get_device_name(0)})"
|
return f"CUDA ({torch.cuda.get_device_name(0)})"
|
||||||
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
||||||
return "MPS (Apple Silicon)"
|
return "MPS (Apple Silicon)"
|
||||||
|
elif backend_type == "mlx":
|
||||||
|
return "Metal (Apple Silicon via MLX)"
|
||||||
return "None (CPU only)"
|
return "None (CPU only)"
|
||||||
|
|
||||||
|
|
||||||
@@ -1191,8 +1597,28 @@ async def startup_event():
|
|||||||
print("voicebox API starting up...")
|
print("voicebox API starting up...")
|
||||||
database.init_db()
|
database.init_db()
|
||||||
print(f"Database initialized at {database._db_path}")
|
print(f"Database initialized at {database._db_path}")
|
||||||
|
backend_type = get_backend_type()
|
||||||
|
print(f"Backend: {backend_type.upper()}")
|
||||||
print(f"GPU available: {_get_gpu_status()}")
|
print(f"GPU available: {_get_gpu_status()}")
|
||||||
|
|
||||||
|
# Initialize progress manager with main event loop for thread-safe operations
|
||||||
|
try:
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
progress_manager._set_main_loop(asyncio.get_running_loop())
|
||||||
|
print("Progress manager initialized with event loop")
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Warning: Could not initialize progress manager event loop: {e}")
|
||||||
|
|
||||||
|
# Ensure HuggingFace cache directory exists
|
||||||
|
try:
|
||||||
|
from huggingface_hub import constants as hf_constants
|
||||||
|
cache_dir = Path(hf_constants.HF_HUB_CACHE)
|
||||||
|
cache_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
print(f"HuggingFace cache directory: {cache_dir}")
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Warning: Could not create HuggingFace cache directory: {e}")
|
||||||
|
print("Model downloads may fail. Please ensure the directory exists and has write permissions.")
|
||||||
|
|
||||||
|
|
||||||
@app.on_event("shutdown")
|
@app.on_event("shutdown")
|
||||||
async def shutdown_event():
|
async def shutdown_event():
|
||||||
|
|||||||
@@ -20,6 +20,7 @@ class VoiceProfileResponse(BaseModel):
|
|||||||
name: str
|
name: str
|
||||||
description: Optional[str]
|
description: Optional[str]
|
||||||
language: str
|
language: str
|
||||||
|
avatar_path: Optional[str] = None
|
||||||
created_at: datetime
|
created_at: datetime
|
||||||
updated_at: datetime
|
updated_at: datetime
|
||||||
|
|
||||||
@@ -32,6 +33,11 @@ class ProfileSampleCreate(BaseModel):
|
|||||||
reference_text: str = Field(..., min_length=1, max_length=1000)
|
reference_text: str = Field(..., min_length=1, max_length=1000)
|
||||||
|
|
||||||
|
|
||||||
|
class ProfileSampleUpdate(BaseModel):
|
||||||
|
"""Request model for updating a profile sample."""
|
||||||
|
reference_text: str = Field(..., min_length=1, max_length=1000)
|
||||||
|
|
||||||
|
|
||||||
class ProfileSampleResponse(BaseModel):
|
class ProfileSampleResponse(BaseModel):
|
||||||
"""Response model for profile sample."""
|
"""Response model for profile sample."""
|
||||||
id: str
|
id: str
|
||||||
@@ -118,7 +124,9 @@ class HealthResponse(BaseModel):
|
|||||||
model_downloaded: Optional[bool] = None # Whether model is cached/downloaded
|
model_downloaded: Optional[bool] = None # Whether model is cached/downloaded
|
||||||
model_size: Optional[str] = None # Current model size if loaded
|
model_size: Optional[str] = None # Current model size if loaded
|
||||||
gpu_available: bool
|
gpu_available: bool
|
||||||
|
gpu_type: Optional[str] = None # GPU type (CUDA, MPS, or None)
|
||||||
vram_used_mb: Optional[float] = None
|
vram_used_mb: Optional[float] = None
|
||||||
|
backend_type: Optional[str] = None # Backend type (mlx or pytorch)
|
||||||
|
|
||||||
|
|
||||||
class ModelStatus(BaseModel):
|
class ModelStatus(BaseModel):
|
||||||
@@ -193,3 +201,100 @@ class ChannelVoiceAssignment(BaseModel):
|
|||||||
class ProfileChannelAssignment(BaseModel):
|
class ProfileChannelAssignment(BaseModel):
|
||||||
"""Request model for assigning channels to a profile."""
|
"""Request model for assigning channels to a profile."""
|
||||||
channel_ids: List[str]
|
channel_ids: List[str]
|
||||||
|
|
||||||
|
|
||||||
|
class StoryCreate(BaseModel):
|
||||||
|
"""Request model for creating a story."""
|
||||||
|
name: str = Field(..., min_length=1, max_length=100)
|
||||||
|
description: Optional[str] = Field(None, max_length=500)
|
||||||
|
|
||||||
|
|
||||||
|
class StoryResponse(BaseModel):
|
||||||
|
"""Response model for story (list view)."""
|
||||||
|
id: str
|
||||||
|
name: str
|
||||||
|
description: Optional[str]
|
||||||
|
created_at: datetime
|
||||||
|
updated_at: datetime
|
||||||
|
item_count: int = 0
|
||||||
|
|
||||||
|
class Config:
|
||||||
|
from_attributes = True
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemDetail(BaseModel):
|
||||||
|
"""Detail model for story item with generation info."""
|
||||||
|
id: str
|
||||||
|
story_id: str
|
||||||
|
generation_id: str
|
||||||
|
start_time_ms: int
|
||||||
|
track: int = 0
|
||||||
|
trim_start_ms: int = 0
|
||||||
|
trim_end_ms: int = 0
|
||||||
|
created_at: datetime
|
||||||
|
# Generation details
|
||||||
|
profile_id: str
|
||||||
|
profile_name: str
|
||||||
|
text: str
|
||||||
|
language: str
|
||||||
|
audio_path: str
|
||||||
|
duration: float
|
||||||
|
seed: Optional[int]
|
||||||
|
instruct: Optional[str]
|
||||||
|
generation_created_at: datetime
|
||||||
|
|
||||||
|
class Config:
|
||||||
|
from_attributes = True
|
||||||
|
|
||||||
|
|
||||||
|
class StoryDetailResponse(BaseModel):
|
||||||
|
"""Response model for story with items."""
|
||||||
|
id: str
|
||||||
|
name: str
|
||||||
|
description: Optional[str]
|
||||||
|
created_at: datetime
|
||||||
|
updated_at: datetime
|
||||||
|
items: List[StoryItemDetail] = []
|
||||||
|
|
||||||
|
class Config:
|
||||||
|
from_attributes = True
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemCreate(BaseModel):
|
||||||
|
"""Request model for adding a generation to a story."""
|
||||||
|
generation_id: str
|
||||||
|
start_time_ms: Optional[int] = None # If not provided, will be calculated automatically
|
||||||
|
track: Optional[int] = 0 # Track number (0 = main track)
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemUpdateTime(BaseModel):
|
||||||
|
"""Request model for updating a story item's timecode."""
|
||||||
|
generation_id: str
|
||||||
|
start_time_ms: int = Field(..., ge=0)
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemBatchUpdate(BaseModel):
|
||||||
|
"""Request model for batch updating story item timecodes."""
|
||||||
|
updates: List[StoryItemUpdateTime]
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemReorder(BaseModel):
|
||||||
|
"""Request model for reordering story items."""
|
||||||
|
generation_ids: List[str] = Field(..., min_length=1)
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemMove(BaseModel):
|
||||||
|
"""Request model for moving a story item (position and/or track)."""
|
||||||
|
start_time_ms: int = Field(..., ge=0)
|
||||||
|
track: int = 0
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemTrim(BaseModel):
|
||||||
|
"""Request model for trimming a story item."""
|
||||||
|
trim_start_ms: int = Field(..., ge=0)
|
||||||
|
trim_end_ms: int = Field(..., ge=0)
|
||||||
|
|
||||||
|
|
||||||
|
class StoryItemSplit(BaseModel):
|
||||||
|
"""Request model for splitting a story item."""
|
||||||
|
split_time_ms: int = Field(..., ge=0) # Time within the clip to split at (relative to clip start)
|
||||||
|
|||||||
@@ -0,0 +1,33 @@
|
|||||||
|
"""
|
||||||
|
Platform detection for backend selection.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import platform
|
||||||
|
from typing import Literal
|
||||||
|
|
||||||
|
|
||||||
|
def is_apple_silicon() -> bool:
|
||||||
|
"""
|
||||||
|
Check if running on Apple Silicon (arm64 macOS).
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
True if on Apple Silicon, False otherwise
|
||||||
|
"""
|
||||||
|
return platform.system() == "Darwin" and platform.machine() == "arm64"
|
||||||
|
|
||||||
|
|
||||||
|
def get_backend_type() -> Literal["mlx", "pytorch"]:
|
||||||
|
"""
|
||||||
|
Detect the best backend for the current platform.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
"mlx" on Apple Silicon (if MLX is available), "pytorch" otherwise
|
||||||
|
"""
|
||||||
|
if is_apple_silicon():
|
||||||
|
try:
|
||||||
|
import mlx
|
||||||
|
return "mlx"
|
||||||
|
except ImportError:
|
||||||
|
# MLX not installed, fallback to PyTorch
|
||||||
|
return "pytorch"
|
||||||
|
return "pytorch"
|
||||||
+172
-22
@@ -21,6 +21,8 @@ from .database import (
|
|||||||
ProfileSample as DBProfileSample,
|
ProfileSample as DBProfileSample,
|
||||||
)
|
)
|
||||||
from .utils.audio import validate_reference_audio, load_audio, save_audio
|
from .utils.audio import validate_reference_audio, load_audio, save_audio
|
||||||
|
from .utils.images import validate_image, process_avatar
|
||||||
|
from .utils.cache import _get_cache_dir, clear_profile_cache
|
||||||
from .tts import get_tts_model
|
from .tts import get_tts_model
|
||||||
from . import config
|
from . import config
|
||||||
|
|
||||||
@@ -119,6 +121,10 @@ async def add_profile_sample(
|
|||||||
db.commit()
|
db.commit()
|
||||||
db.refresh(db_sample)
|
db.refresh(db_sample)
|
||||||
|
|
||||||
|
# Invalidate combined audio cache for this profile
|
||||||
|
# Since a new sample was added, any cached combined audio is now stale
|
||||||
|
clear_profile_cache(profile_id)
|
||||||
|
|
||||||
return ProfileSampleResponse.model_validate(db_sample)
|
return ProfileSampleResponse.model_validate(db_sample)
|
||||||
|
|
||||||
|
|
||||||
@@ -240,6 +246,9 @@ async def delete_profile(
|
|||||||
if profile_dir.exists():
|
if profile_dir.exists():
|
||||||
shutil.rmtree(profile_dir)
|
shutil.rmtree(profile_dir)
|
||||||
|
|
||||||
|
# Clean up combined audio cache files for this profile
|
||||||
|
clear_profile_cache(profile_id)
|
||||||
|
|
||||||
return True
|
return True
|
||||||
|
|
||||||
|
|
||||||
@@ -261,6 +270,9 @@ async def delete_profile_sample(
|
|||||||
if not sample:
|
if not sample:
|
||||||
return False
|
return False
|
||||||
|
|
||||||
|
# Store profile_id before deleting
|
||||||
|
profile_id = sample.profile_id
|
||||||
|
|
||||||
# Delete audio file
|
# Delete audio file
|
||||||
audio_path = Path(sample.audio_path)
|
audio_path = Path(sample.audio_path)
|
||||||
if audio_path.exists():
|
if audio_path.exists():
|
||||||
@@ -270,9 +282,47 @@ async def delete_profile_sample(
|
|||||||
db.delete(sample)
|
db.delete(sample)
|
||||||
db.commit()
|
db.commit()
|
||||||
|
|
||||||
|
# Invalidate combined audio cache for this profile
|
||||||
|
# Since the sample set changed, any cached combined audio is now stale
|
||||||
|
clear_profile_cache(profile_id)
|
||||||
|
|
||||||
return True
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
async def update_profile_sample(
|
||||||
|
sample_id: str,
|
||||||
|
reference_text: str,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[ProfileSampleResponse]:
|
||||||
|
"""
|
||||||
|
Update a profile sample's reference text.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
sample_id: Sample ID
|
||||||
|
reference_text: Updated reference text
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Updated sample or None if not found
|
||||||
|
"""
|
||||||
|
sample = db.query(DBProfileSample).filter_by(id=sample_id).first()
|
||||||
|
if not sample:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Store profile_id before updating
|
||||||
|
profile_id = sample.profile_id
|
||||||
|
|
||||||
|
sample.reference_text = reference_text
|
||||||
|
db.commit()
|
||||||
|
db.refresh(sample)
|
||||||
|
|
||||||
|
# Invalidate combined audio cache for this profile
|
||||||
|
# Since the reference text changed, cache keys and combined text are now stale
|
||||||
|
clear_profile_cache(profile_id)
|
||||||
|
|
||||||
|
return ProfileSampleResponse.model_validate(sample)
|
||||||
|
|
||||||
|
|
||||||
async def create_voice_prompt_for_profile(
|
async def create_voice_prompt_for_profile(
|
||||||
profile_id: str,
|
profile_id: str,
|
||||||
db: Session,
|
db: Session,
|
||||||
@@ -280,23 +330,23 @@ async def create_voice_prompt_for_profile(
|
|||||||
) -> dict:
|
) -> dict:
|
||||||
"""
|
"""
|
||||||
Create a combined voice prompt from all samples in a profile.
|
Create a combined voice prompt from all samples in a profile.
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
profile_id: Profile ID
|
profile_id: Profile ID
|
||||||
db: Database session
|
db: Database session
|
||||||
use_cache: Whether to use cached prompts
|
use_cache: Whether to use cached prompts
|
||||||
|
|
||||||
Returns:
|
Returns:
|
||||||
Voice prompt dictionary
|
Voice prompt dictionary
|
||||||
"""
|
"""
|
||||||
# Get all samples for profile
|
# Get all samples for profile
|
||||||
samples = db.query(DBProfileSample).filter_by(profile_id=profile_id).all()
|
samples = db.query(DBProfileSample).filter_by(profile_id=profile_id).all()
|
||||||
|
|
||||||
if not samples:
|
if not samples:
|
||||||
raise ValueError(f"No samples found for profile {profile_id}")
|
raise ValueError(f"No samples found for profile {profile_id}")
|
||||||
|
|
||||||
tts_model = get_tts_model()
|
tts_model = get_tts_model()
|
||||||
|
|
||||||
if len(samples) == 1:
|
if len(samples) == 1:
|
||||||
# Single sample - use directly
|
# Single sample - use directly
|
||||||
sample = samples[0]
|
sample = samples[0]
|
||||||
@@ -310,27 +360,127 @@ async def create_voice_prompt_for_profile(
|
|||||||
# Multiple samples - combine them
|
# Multiple samples - combine them
|
||||||
audio_paths = [s.audio_path for s in samples]
|
audio_paths = [s.audio_path for s in samples]
|
||||||
reference_texts = [s.reference_text for s in samples]
|
reference_texts = [s.reference_text for s in samples]
|
||||||
|
|
||||||
# Combine audio
|
# Combine audio
|
||||||
combined_audio, combined_text = await tts_model.combine_voice_prompts(
|
combined_audio, combined_text = await tts_model.combine_voice_prompts(
|
||||||
audio_paths,
|
audio_paths,
|
||||||
reference_texts,
|
reference_texts,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# Save combined audio to cache directory (persistent)
|
||||||
|
# Create a hash of sample IDs to identify this specific combination
|
||||||
|
import hashlib
|
||||||
|
sample_ids_str = "-".join(sorted([s.id for s in samples]))
|
||||||
|
combination_hash = hashlib.md5(sample_ids_str.encode()).hexdigest()[:12]
|
||||||
|
|
||||||
# Save combined audio temporarily
|
# Store in cache directory
|
||||||
import tempfile
|
cache_dir = _get_cache_dir()
|
||||||
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp:
|
cache_dir.mkdir(parents=True, exist_ok=True)
|
||||||
save_audio(combined_audio, tmp.name, 24000)
|
combined_path = cache_dir / f"combined_{profile_id}_{combination_hash}.wav"
|
||||||
tmp_path = tmp.name
|
|
||||||
|
|
||||||
try:
|
# Save combined audio
|
||||||
# Create prompt from combined audio
|
save_audio(combined_audio, str(combined_path), 24000)
|
||||||
voice_prompt, _ = await tts_model.create_voice_prompt(
|
|
||||||
tmp_path,
|
# Create prompt from combined audio
|
||||||
combined_text,
|
voice_prompt, _ = await tts_model.create_voice_prompt(
|
||||||
use_cache=use_cache,
|
str(combined_path),
|
||||||
)
|
combined_text,
|
||||||
return voice_prompt
|
use_cache=use_cache,
|
||||||
finally:
|
)
|
||||||
# Clean up temp file
|
return voice_prompt
|
||||||
Path(tmp_path).unlink(missing_ok=True)
|
|
||||||
|
|
||||||
|
async def upload_avatar(
|
||||||
|
profile_id: str,
|
||||||
|
image_path: str,
|
||||||
|
db: Session,
|
||||||
|
) -> VoiceProfileResponse:
|
||||||
|
"""
|
||||||
|
Upload and process avatar image for a profile.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
profile_id: Profile ID
|
||||||
|
image_path: Path to uploaded image file
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Updated profile
|
||||||
|
"""
|
||||||
|
# Validate profile exists
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=profile_id).first()
|
||||||
|
if not profile:
|
||||||
|
raise ValueError(f"Profile {profile_id} not found")
|
||||||
|
|
||||||
|
# Validate image
|
||||||
|
is_valid, error_msg = validate_image(image_path)
|
||||||
|
if not is_valid:
|
||||||
|
raise ValueError(error_msg)
|
||||||
|
|
||||||
|
# Delete existing avatar if present
|
||||||
|
if profile.avatar_path:
|
||||||
|
old_avatar = Path(profile.avatar_path)
|
||||||
|
if old_avatar.exists():
|
||||||
|
old_avatar.unlink()
|
||||||
|
|
||||||
|
# Determine file extension from uploaded file
|
||||||
|
from PIL import Image
|
||||||
|
with Image.open(image_path) as img:
|
||||||
|
# Normalize JPEG variants (MPO is multi-picture format from some cameras)
|
||||||
|
img_format = img.format
|
||||||
|
if img_format in ('MPO', 'JPG'):
|
||||||
|
img_format = 'JPEG'
|
||||||
|
|
||||||
|
ext_map = {
|
||||||
|
'PNG': '.png',
|
||||||
|
'JPEG': '.jpg',
|
||||||
|
'WEBP': '.webp'
|
||||||
|
}
|
||||||
|
ext = ext_map.get(img_format, '.png')
|
||||||
|
|
||||||
|
# Save processed image to profile directory
|
||||||
|
profile_dir = _get_profiles_dir() / profile_id
|
||||||
|
profile_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
output_path = profile_dir / f"avatar{ext}"
|
||||||
|
|
||||||
|
process_avatar(image_path, str(output_path))
|
||||||
|
|
||||||
|
# Update database
|
||||||
|
profile.avatar_path = str(output_path)
|
||||||
|
profile.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
db.refresh(profile)
|
||||||
|
|
||||||
|
return VoiceProfileResponse.model_validate(profile)
|
||||||
|
|
||||||
|
|
||||||
|
async def delete_avatar(
|
||||||
|
profile_id: str,
|
||||||
|
db: Session,
|
||||||
|
) -> bool:
|
||||||
|
"""
|
||||||
|
Delete avatar image for a profile.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
profile_id: Profile ID
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
True if deleted, False if not found or no avatar
|
||||||
|
"""
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=profile_id).first()
|
||||||
|
if not profile or not profile.avatar_path:
|
||||||
|
return False
|
||||||
|
|
||||||
|
# Delete avatar file
|
||||||
|
avatar_path = Path(profile.avatar_path)
|
||||||
|
if avatar_path.exists():
|
||||||
|
avatar_path.unlink()
|
||||||
|
|
||||||
|
# Update database
|
||||||
|
profile.avatar_path = None
|
||||||
|
profile.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
|
||||||
|
return True
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
# MLX-specific dependencies (Apple Silicon only)
|
||||||
|
# These should only be installed on aarch64-apple-darwin platforms
|
||||||
|
|
||||||
|
mlx>=0.30.0
|
||||||
|
mlx-audio>=0.3.1
|
||||||
@@ -21,3 +21,4 @@ numpy>=1.24.0
|
|||||||
|
|
||||||
# Utilities
|
# Utilities
|
||||||
python-multipart>=0.0.6
|
python-multipart>=0.0.6
|
||||||
|
Pillow>=10.0.0
|
||||||
|
|||||||
@@ -0,0 +1,972 @@
|
|||||||
|
"""
|
||||||
|
Story management module.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from typing import List, Optional
|
||||||
|
from datetime import datetime
|
||||||
|
import uuid
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
from sqlalchemy.orm import Session
|
||||||
|
from sqlalchemy import func
|
||||||
|
|
||||||
|
from .models import (
|
||||||
|
StoryCreate,
|
||||||
|
StoryResponse,
|
||||||
|
StoryDetailResponse,
|
||||||
|
StoryItemDetail,
|
||||||
|
StoryItemCreate,
|
||||||
|
StoryItemBatchUpdate,
|
||||||
|
StoryItemMove,
|
||||||
|
StoryItemTrim,
|
||||||
|
StoryItemSplit,
|
||||||
|
)
|
||||||
|
from .database import Story as DBStory, StoryItem as DBStoryItem, Generation as DBGeneration, VoiceProfile as DBVoiceProfile
|
||||||
|
from .utils.audio import load_audio, save_audio
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
|
||||||
|
async def create_story(
|
||||||
|
data: StoryCreate,
|
||||||
|
db: Session,
|
||||||
|
) -> StoryResponse:
|
||||||
|
"""
|
||||||
|
Create a new story.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
data: Story creation data
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Created story
|
||||||
|
"""
|
||||||
|
db_story = DBStory(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
name=data.name,
|
||||||
|
description=data.description,
|
||||||
|
created_at=datetime.utcnow(),
|
||||||
|
updated_at=datetime.utcnow(),
|
||||||
|
)
|
||||||
|
|
||||||
|
db.add(db_story)
|
||||||
|
db.commit()
|
||||||
|
db.refresh(db_story)
|
||||||
|
|
||||||
|
# Get item count
|
||||||
|
item_count = db.query(func.count(DBStoryItem.id)).filter(
|
||||||
|
DBStoryItem.story_id == db_story.id
|
||||||
|
).scalar()
|
||||||
|
|
||||||
|
response = StoryResponse.model_validate(db_story)
|
||||||
|
response.item_count = item_count
|
||||||
|
return response
|
||||||
|
|
||||||
|
|
||||||
|
async def list_stories(
|
||||||
|
db: Session,
|
||||||
|
) -> List[StoryResponse]:
|
||||||
|
"""
|
||||||
|
List all stories.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List of stories with item counts
|
||||||
|
"""
|
||||||
|
stories = db.query(DBStory).order_by(DBStory.updated_at.desc()).all()
|
||||||
|
|
||||||
|
result = []
|
||||||
|
for story in stories:
|
||||||
|
item_count = db.query(func.count(DBStoryItem.id)).filter(
|
||||||
|
DBStoryItem.story_id == story.id
|
||||||
|
).scalar()
|
||||||
|
|
||||||
|
response = StoryResponse.model_validate(story)
|
||||||
|
response.item_count = item_count
|
||||||
|
result.append(response)
|
||||||
|
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
async def get_story(
|
||||||
|
story_id: str,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[StoryDetailResponse]:
|
||||||
|
"""
|
||||||
|
Get a story with all its items.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Story with items or None if not found
|
||||||
|
"""
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get all items ordered by start_time_ms
|
||||||
|
items = db.query(
|
||||||
|
DBStoryItem,
|
||||||
|
DBGeneration,
|
||||||
|
DBVoiceProfile.name.label('profile_name')
|
||||||
|
).join(
|
||||||
|
DBGeneration,
|
||||||
|
DBStoryItem.generation_id == DBGeneration.id
|
||||||
|
).join(
|
||||||
|
DBVoiceProfile,
|
||||||
|
DBGeneration.profile_id == DBVoiceProfile.id
|
||||||
|
).filter(
|
||||||
|
DBStoryItem.story_id == story_id
|
||||||
|
).order_by(DBStoryItem.start_time_ms).all()
|
||||||
|
|
||||||
|
# Build item details
|
||||||
|
item_details = []
|
||||||
|
for item, generation, profile_name in items:
|
||||||
|
item_detail = StoryItemDetail(
|
||||||
|
id=item.id,
|
||||||
|
story_id=item.story_id,
|
||||||
|
generation_id=item.generation_id,
|
||||||
|
start_time_ms=item.start_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=getattr(item, 'trim_start_ms', 0),
|
||||||
|
trim_end_ms=getattr(item, 'trim_end_ms', 0),
|
||||||
|
created_at=item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile_name,
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
item_details.append(item_detail)
|
||||||
|
|
||||||
|
response = StoryDetailResponse.model_validate(story)
|
||||||
|
response.items = item_details
|
||||||
|
return response
|
||||||
|
|
||||||
|
|
||||||
|
async def update_story(
|
||||||
|
story_id: str,
|
||||||
|
data: StoryCreate,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[StoryResponse]:
|
||||||
|
"""
|
||||||
|
Update a story.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
data: Update data
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Updated story or None if not found
|
||||||
|
"""
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
return None
|
||||||
|
|
||||||
|
story.name = data.name
|
||||||
|
story.description = data.description
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
db.refresh(story)
|
||||||
|
|
||||||
|
# Get item count
|
||||||
|
item_count = db.query(func.count(DBStoryItem.id)).filter(
|
||||||
|
DBStoryItem.story_id == story.id
|
||||||
|
).scalar()
|
||||||
|
|
||||||
|
response = StoryResponse.model_validate(story)
|
||||||
|
response.item_count = item_count
|
||||||
|
return response
|
||||||
|
|
||||||
|
|
||||||
|
async def delete_story(
|
||||||
|
story_id: str,
|
||||||
|
db: Session,
|
||||||
|
) -> bool:
|
||||||
|
"""
|
||||||
|
Delete a story and all its items.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
True if deleted, False if not found
|
||||||
|
"""
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
return False
|
||||||
|
|
||||||
|
# Delete all items
|
||||||
|
db.query(DBStoryItem).filter_by(story_id=story_id).delete()
|
||||||
|
|
||||||
|
# Delete story
|
||||||
|
db.delete(story)
|
||||||
|
db.commit()
|
||||||
|
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
async def add_item_to_story(
|
||||||
|
story_id: str,
|
||||||
|
data: StoryItemCreate,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[StoryItemDetail]:
|
||||||
|
"""
|
||||||
|
Add a generation to a story.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
data: Item creation data
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Created item detail or None if story/generation not found
|
||||||
|
"""
|
||||||
|
# Verify story exists
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Verify generation exists
|
||||||
|
generation = db.query(DBGeneration).filter_by(id=data.generation_id).first()
|
||||||
|
if not generation:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Check if generation is already in story
|
||||||
|
existing = db.query(DBStoryItem).filter_by(
|
||||||
|
story_id=story_id,
|
||||||
|
generation_id=data.generation_id
|
||||||
|
).first()
|
||||||
|
if existing:
|
||||||
|
# Return existing item
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=generation.profile_id).first()
|
||||||
|
return StoryItemDetail(
|
||||||
|
id=existing.id,
|
||||||
|
story_id=existing.story_id,
|
||||||
|
generation_id=existing.generation_id,
|
||||||
|
start_time_ms=existing.start_time_ms,
|
||||||
|
track=existing.track,
|
||||||
|
trim_start_ms=getattr(existing, 'trim_start_ms', 0),
|
||||||
|
trim_end_ms=getattr(existing, 'trim_end_ms', 0),
|
||||||
|
created_at=existing.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile.name if profile else "Unknown",
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
# Calculate start_time_ms if not provided
|
||||||
|
if data.start_time_ms is not None:
|
||||||
|
start_time_ms = data.start_time_ms
|
||||||
|
else:
|
||||||
|
# Find the maximum end time (start_time_ms + duration_ms) of existing items
|
||||||
|
existing_items = db.query(
|
||||||
|
DBStoryItem,
|
||||||
|
DBGeneration
|
||||||
|
).join(
|
||||||
|
DBGeneration,
|
||||||
|
DBStoryItem.generation_id == DBGeneration.id
|
||||||
|
).filter(
|
||||||
|
DBStoryItem.story_id == story_id
|
||||||
|
).all()
|
||||||
|
|
||||||
|
if not existing_items:
|
||||||
|
# First item starts at 0
|
||||||
|
start_time_ms = 0
|
||||||
|
else:
|
||||||
|
max_end_time_ms = 0
|
||||||
|
for item, gen in existing_items:
|
||||||
|
item_end_ms = item.start_time_ms + int(gen.duration * 1000)
|
||||||
|
max_end_time_ms = max(max_end_time_ms, item_end_ms)
|
||||||
|
|
||||||
|
# Add 200ms gap after the last item
|
||||||
|
start_time_ms = max_end_time_ms + 200
|
||||||
|
|
||||||
|
# Get track from data or default to 0
|
||||||
|
track = data.track if data.track is not None else 0
|
||||||
|
|
||||||
|
# Create item
|
||||||
|
item = DBStoryItem(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
story_id=story_id,
|
||||||
|
generation_id=data.generation_id,
|
||||||
|
start_time_ms=start_time_ms,
|
||||||
|
track=track,
|
||||||
|
created_at=datetime.utcnow(),
|
||||||
|
)
|
||||||
|
|
||||||
|
db.add(item)
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
db.refresh(item)
|
||||||
|
|
||||||
|
# Get profile name
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=generation.profile_id).first()
|
||||||
|
|
||||||
|
return StoryItemDetail(
|
||||||
|
id=item.id,
|
||||||
|
story_id=item.story_id,
|
||||||
|
generation_id=item.generation_id,
|
||||||
|
start_time_ms=item.start_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=getattr(item, 'trim_start_ms', 0),
|
||||||
|
trim_end_ms=getattr(item, 'trim_end_ms', 0),
|
||||||
|
created_at=item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile.name if profile else "Unknown",
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def move_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: StoryItemMove,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[StoryItemDetail]:
|
||||||
|
"""
|
||||||
|
Move a story item (update position and/or track).
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
item_id: Story item ID
|
||||||
|
data: New position and track data
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Updated item detail or None if not found
|
||||||
|
"""
|
||||||
|
# Get the item
|
||||||
|
item = db.query(DBStoryItem).filter_by(
|
||||||
|
id=item_id,
|
||||||
|
story_id=story_id,
|
||||||
|
).first()
|
||||||
|
if not item:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get the generation
|
||||||
|
generation = db.query(DBGeneration).filter_by(id=item.generation_id).first()
|
||||||
|
if not generation:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Update position and track
|
||||||
|
item.start_time_ms = data.start_time_ms
|
||||||
|
item.track = data.track
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if story:
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
db.refresh(item)
|
||||||
|
|
||||||
|
# Get profile name
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=generation.profile_id).first()
|
||||||
|
|
||||||
|
return StoryItemDetail(
|
||||||
|
id=item.id,
|
||||||
|
story_id=item.story_id,
|
||||||
|
generation_id=item.generation_id,
|
||||||
|
start_time_ms=item.start_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=getattr(item, 'trim_start_ms', 0),
|
||||||
|
trim_end_ms=getattr(item, 'trim_end_ms', 0),
|
||||||
|
created_at=item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile.name if profile else "Unknown",
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def remove_item_from_story(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
db: Session,
|
||||||
|
) -> bool:
|
||||||
|
"""
|
||||||
|
Remove a story item from a story.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
item_id: Story item ID to remove
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
True if removed, False if not found
|
||||||
|
"""
|
||||||
|
item = db.query(DBStoryItem).filter_by(
|
||||||
|
id=item_id,
|
||||||
|
story_id=story_id,
|
||||||
|
).first()
|
||||||
|
if not item:
|
||||||
|
return False
|
||||||
|
|
||||||
|
# Delete item
|
||||||
|
db.delete(item)
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if story:
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
async def trim_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: StoryItemTrim,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[StoryItemDetail]:
|
||||||
|
"""
|
||||||
|
Trim a story item (update trim_start_ms and trim_end_ms).
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
item_id: Story item ID
|
||||||
|
data: Trim data (trim_start_ms, trim_end_ms)
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Updated item detail or None if not found
|
||||||
|
"""
|
||||||
|
# Get the item
|
||||||
|
item = db.query(DBStoryItem).filter_by(
|
||||||
|
id=item_id,
|
||||||
|
story_id=story_id,
|
||||||
|
).first()
|
||||||
|
if not item:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get the generation
|
||||||
|
generation = db.query(DBGeneration).filter_by(id=item.generation_id).first()
|
||||||
|
if not generation:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Validate trim values don't exceed duration
|
||||||
|
max_duration_ms = int(generation.duration * 1000)
|
||||||
|
if data.trim_start_ms + data.trim_end_ms >= max_duration_ms:
|
||||||
|
return None # Invalid trim - would result in zero or negative duration
|
||||||
|
|
||||||
|
# Update trim values
|
||||||
|
item.trim_start_ms = data.trim_start_ms
|
||||||
|
item.trim_end_ms = data.trim_end_ms
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if story:
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
db.refresh(item)
|
||||||
|
|
||||||
|
# Get profile name
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=generation.profile_id).first()
|
||||||
|
|
||||||
|
return StoryItemDetail(
|
||||||
|
id=item.id,
|
||||||
|
story_id=item.story_id,
|
||||||
|
generation_id=item.generation_id,
|
||||||
|
start_time_ms=item.start_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=item.trim_start_ms,
|
||||||
|
trim_end_ms=item.trim_end_ms,
|
||||||
|
created_at=item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile.name if profile else "Unknown",
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def split_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: StoryItemSplit,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[List[StoryItemDetail]]:
|
||||||
|
"""
|
||||||
|
Split a story item at a given time, creating two clips.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
item_id: Story item ID to split
|
||||||
|
data: Split data (split_time_ms - time within clip to split at)
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List of two updated item details (original and new) or None if not found/invalid
|
||||||
|
"""
|
||||||
|
# Get the item
|
||||||
|
item = db.query(DBStoryItem).filter_by(
|
||||||
|
id=item_id,
|
||||||
|
story_id=story_id,
|
||||||
|
).first()
|
||||||
|
if not item:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get the generation
|
||||||
|
generation = db.query(DBGeneration).filter_by(id=item.generation_id).first()
|
||||||
|
if not generation:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Calculate effective duration and validate split point
|
||||||
|
current_trim_start = getattr(item, 'trim_start_ms', 0)
|
||||||
|
current_trim_end = getattr(item, 'trim_end_ms', 0)
|
||||||
|
original_duration_ms = int(generation.duration * 1000)
|
||||||
|
effective_duration_ms = original_duration_ms - current_trim_start - current_trim_end
|
||||||
|
|
||||||
|
# Validate split_time_ms is within the effective duration
|
||||||
|
if data.split_time_ms <= 0 or data.split_time_ms >= effective_duration_ms:
|
||||||
|
return None # Invalid split point
|
||||||
|
|
||||||
|
# Calculate the absolute time in the original audio where we're splitting
|
||||||
|
absolute_split_ms = current_trim_start + data.split_time_ms
|
||||||
|
|
||||||
|
# Update original clip: trim from the end
|
||||||
|
item.trim_end_ms = original_duration_ms - absolute_split_ms
|
||||||
|
|
||||||
|
# Create new clip: starts after the split, trimmed from the start
|
||||||
|
new_item = DBStoryItem(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
story_id=story_id,
|
||||||
|
generation_id=item.generation_id, # Same generation, different trim
|
||||||
|
start_time_ms=item.start_time_ms + data.split_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=absolute_split_ms,
|
||||||
|
trim_end_ms=current_trim_end,
|
||||||
|
created_at=datetime.utcnow(),
|
||||||
|
)
|
||||||
|
|
||||||
|
db.add(new_item)
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if story:
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
db.refresh(item)
|
||||||
|
db.refresh(new_item)
|
||||||
|
|
||||||
|
# Get profile name
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=generation.profile_id).first()
|
||||||
|
profile_name = profile.name if profile else "Unknown"
|
||||||
|
|
||||||
|
# Build response items
|
||||||
|
original_item_detail = StoryItemDetail(
|
||||||
|
id=item.id,
|
||||||
|
story_id=item.story_id,
|
||||||
|
generation_id=item.generation_id,
|
||||||
|
start_time_ms=item.start_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=item.trim_start_ms,
|
||||||
|
trim_end_ms=item.trim_end_ms,
|
||||||
|
created_at=item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile_name,
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
new_item_detail = StoryItemDetail(
|
||||||
|
id=new_item.id,
|
||||||
|
story_id=new_item.story_id,
|
||||||
|
generation_id=new_item.generation_id,
|
||||||
|
start_time_ms=new_item.start_time_ms,
|
||||||
|
track=new_item.track,
|
||||||
|
trim_start_ms=new_item.trim_start_ms,
|
||||||
|
trim_end_ms=new_item.trim_end_ms,
|
||||||
|
created_at=new_item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile_name,
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
return [original_item_detail, new_item_detail]
|
||||||
|
|
||||||
|
|
||||||
|
async def duplicate_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[StoryItemDetail]:
|
||||||
|
"""
|
||||||
|
Duplicate a story item, creating a copy with all properties.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
item_id: Story item ID to duplicate
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
New item detail or None if not found
|
||||||
|
"""
|
||||||
|
# Get the original item
|
||||||
|
original_item = db.query(DBStoryItem).filter_by(
|
||||||
|
id=item_id,
|
||||||
|
story_id=story_id,
|
||||||
|
).first()
|
||||||
|
if not original_item:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get the generation
|
||||||
|
generation = db.query(DBGeneration).filter_by(id=original_item.generation_id).first()
|
||||||
|
if not generation:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Calculate effective duration
|
||||||
|
current_trim_start = getattr(original_item, 'trim_start_ms', 0)
|
||||||
|
current_trim_end = getattr(original_item, 'trim_end_ms', 0)
|
||||||
|
original_duration_ms = int(generation.duration * 1000)
|
||||||
|
effective_duration_ms = original_duration_ms - current_trim_start - current_trim_end
|
||||||
|
|
||||||
|
# Create duplicate item - place it right after the original
|
||||||
|
new_item = DBStoryItem(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
story_id=story_id,
|
||||||
|
generation_id=original_item.generation_id, # Same generation as original
|
||||||
|
start_time_ms=original_item.start_time_ms + effective_duration_ms + 200, # 200ms gap
|
||||||
|
track=original_item.track,
|
||||||
|
trim_start_ms=current_trim_start,
|
||||||
|
trim_end_ms=current_trim_end,
|
||||||
|
created_at=datetime.utcnow(),
|
||||||
|
)
|
||||||
|
|
||||||
|
db.add(new_item)
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if story:
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
db.refresh(new_item)
|
||||||
|
|
||||||
|
# Get profile name
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=generation.profile_id).first()
|
||||||
|
|
||||||
|
return StoryItemDetail(
|
||||||
|
id=new_item.id,
|
||||||
|
story_id=new_item.story_id,
|
||||||
|
generation_id=new_item.generation_id,
|
||||||
|
start_time_ms=new_item.start_time_ms,
|
||||||
|
track=new_item.track,
|
||||||
|
trim_start_ms=new_item.trim_start_ms,
|
||||||
|
trim_end_ms=new_item.trim_end_ms,
|
||||||
|
created_at=new_item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile.name if profile else "Unknown",
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def update_story_item_times(
|
||||||
|
story_id: str,
|
||||||
|
data: StoryItemBatchUpdate,
|
||||||
|
db: Session,
|
||||||
|
) -> bool:
|
||||||
|
"""
|
||||||
|
Update story item timecodes.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
data: Batch update data with timecodes
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
True if updated, False if story not found or invalid
|
||||||
|
"""
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
return False
|
||||||
|
|
||||||
|
# Get all items for this story
|
||||||
|
items = db.query(DBStoryItem).filter_by(story_id=story_id).all()
|
||||||
|
item_map = {item.generation_id: item for item in items}
|
||||||
|
|
||||||
|
# Verify all generation IDs belong to this story and update timecodes
|
||||||
|
for update in data.updates:
|
||||||
|
if update.generation_id not in item_map:
|
||||||
|
return False
|
||||||
|
item_map[update.generation_id].start_time_ms = update.start_time_ms
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
async def reorder_story_items(
|
||||||
|
story_id: str,
|
||||||
|
generation_ids: List[str],
|
||||||
|
db: Session,
|
||||||
|
gap_ms: int = 200,
|
||||||
|
) -> Optional[List[StoryItemDetail]]:
|
||||||
|
"""
|
||||||
|
Reorder story items and recalculate timecodes.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
generation_ids: List of generation IDs in the desired order
|
||||||
|
db: Database session
|
||||||
|
gap_ms: Gap in milliseconds between items (default 200ms)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Updated list of story items with new timecodes, or None if invalid
|
||||||
|
"""
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get all items for this story with their generation data
|
||||||
|
items_with_gen = db.query(
|
||||||
|
DBStoryItem,
|
||||||
|
DBGeneration,
|
||||||
|
DBVoiceProfile.name.label('profile_name')
|
||||||
|
).join(
|
||||||
|
DBGeneration,
|
||||||
|
DBStoryItem.generation_id == DBGeneration.id
|
||||||
|
).join(
|
||||||
|
DBVoiceProfile,
|
||||||
|
DBGeneration.profile_id == DBVoiceProfile.id
|
||||||
|
).filter(
|
||||||
|
DBStoryItem.story_id == story_id
|
||||||
|
).all()
|
||||||
|
|
||||||
|
# Create maps for quick lookup
|
||||||
|
item_map = {item.generation_id: (item, gen, profile_name) for item, gen, profile_name in items_with_gen}
|
||||||
|
|
||||||
|
# Verify all generation IDs belong to this story
|
||||||
|
if set(generation_ids) != set(item_map.keys()):
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Recalculate timecodes based on new order
|
||||||
|
current_time_ms = 0
|
||||||
|
updated_items = []
|
||||||
|
|
||||||
|
for gen_id in generation_ids:
|
||||||
|
item, generation, profile_name = item_map[gen_id]
|
||||||
|
|
||||||
|
# Update the item's start time
|
||||||
|
item.start_time_ms = current_time_ms
|
||||||
|
|
||||||
|
# Calculate the duration in ms
|
||||||
|
duration_ms = int(generation.duration * 1000)
|
||||||
|
|
||||||
|
# Move to next position (current end + gap)
|
||||||
|
current_time_ms += duration_ms + gap_ms
|
||||||
|
|
||||||
|
# Build the response item
|
||||||
|
updated_items.append(StoryItemDetail(
|
||||||
|
id=item.id,
|
||||||
|
story_id=item.story_id,
|
||||||
|
generation_id=item.generation_id,
|
||||||
|
start_time_ms=item.start_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=getattr(item, 'trim_start_ms', 0),
|
||||||
|
trim_end_ms=getattr(item, 'trim_end_ms', 0),
|
||||||
|
created_at=item.created_at,
|
||||||
|
profile_id=generation.profile_id,
|
||||||
|
profile_name=profile_name,
|
||||||
|
text=generation.text,
|
||||||
|
language=generation.language,
|
||||||
|
audio_path=generation.audio_path,
|
||||||
|
duration=generation.duration,
|
||||||
|
seed=generation.seed,
|
||||||
|
instruct=generation.instruct,
|
||||||
|
generation_created_at=generation.created_at,
|
||||||
|
))
|
||||||
|
|
||||||
|
# Update story updated_at
|
||||||
|
story.updated_at = datetime.utcnow()
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
return updated_items
|
||||||
|
|
||||||
|
|
||||||
|
async def export_story_audio(
|
||||||
|
story_id: str,
|
||||||
|
db: Session,
|
||||||
|
) -> Optional[bytes]:
|
||||||
|
"""
|
||||||
|
Export story as single mixed audio file with timecode-based mixing.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
story_id: Story ID
|
||||||
|
db: Database session
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Audio file bytes or None if story not found
|
||||||
|
"""
|
||||||
|
story = db.query(DBStory).filter_by(id=story_id).first()
|
||||||
|
if not story:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Get all items ordered by start_time_ms
|
||||||
|
items = db.query(
|
||||||
|
DBStoryItem,
|
||||||
|
DBGeneration
|
||||||
|
).join(
|
||||||
|
DBGeneration,
|
||||||
|
DBStoryItem.generation_id == DBGeneration.id
|
||||||
|
).filter(
|
||||||
|
DBStoryItem.story_id == story_id
|
||||||
|
).order_by(DBStoryItem.start_time_ms).all()
|
||||||
|
|
||||||
|
if not items:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Load all audio files and calculate total duration
|
||||||
|
audio_data = []
|
||||||
|
sample_rate = 24000 # Default sample rate
|
||||||
|
|
||||||
|
for item, generation in items:
|
||||||
|
audio_path = Path(generation.audio_path)
|
||||||
|
if not audio_path.exists():
|
||||||
|
continue
|
||||||
|
|
||||||
|
try:
|
||||||
|
audio, sr = load_audio(str(audio_path), sample_rate=sample_rate)
|
||||||
|
sample_rate = sr # Use actual sample rate from first file
|
||||||
|
|
||||||
|
# Get trim values
|
||||||
|
trim_start_ms = getattr(item, 'trim_start_ms', 0)
|
||||||
|
trim_end_ms = getattr(item, 'trim_end_ms', 0)
|
||||||
|
|
||||||
|
# Calculate effective duration
|
||||||
|
original_duration_ms = int(generation.duration * 1000)
|
||||||
|
effective_duration_ms = original_duration_ms - trim_start_ms - trim_end_ms
|
||||||
|
|
||||||
|
# Slice audio based on trim values
|
||||||
|
trim_start_sample = int((trim_start_ms / 1000.0) * sample_rate)
|
||||||
|
trim_end_sample = int((trim_end_ms / 1000.0) * sample_rate)
|
||||||
|
|
||||||
|
# Extract the trimmed portion
|
||||||
|
if trim_end_ms > 0:
|
||||||
|
trimmed_audio = audio[trim_start_sample:-trim_end_sample] if trim_end_sample > 0 else audio[trim_start_sample:]
|
||||||
|
else:
|
||||||
|
trimmed_audio = audio[trim_start_sample:]
|
||||||
|
|
||||||
|
# Store audio with its timecode info
|
||||||
|
start_time_ms = item.start_time_ms
|
||||||
|
|
||||||
|
audio_data.append({
|
||||||
|
'audio': trimmed_audio,
|
||||||
|
'start_time_ms': start_time_ms,
|
||||||
|
'duration_ms': effective_duration_ms,
|
||||||
|
})
|
||||||
|
except Exception:
|
||||||
|
# Skip files that can't be loaded
|
||||||
|
continue
|
||||||
|
|
||||||
|
if not audio_data:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# Calculate total duration: max(start_time_ms + duration_ms)
|
||||||
|
max_end_time_ms = max(
|
||||||
|
(data['start_time_ms'] + data['duration_ms'] for data in audio_data),
|
||||||
|
default=0
|
||||||
|
)
|
||||||
|
|
||||||
|
# Convert to samples
|
||||||
|
total_samples = int((max_end_time_ms / 1000.0) * sample_rate)
|
||||||
|
|
||||||
|
# Create output buffer initialized to zeros
|
||||||
|
final_audio = np.zeros(total_samples, dtype=np.float32)
|
||||||
|
|
||||||
|
# Mix each audio segment at its timecode position
|
||||||
|
for data in audio_data:
|
||||||
|
audio = data['audio']
|
||||||
|
start_time_ms = data['start_time_ms']
|
||||||
|
|
||||||
|
# Calculate start sample index
|
||||||
|
start_sample = int((start_time_ms / 1000.0) * sample_rate)
|
||||||
|
|
||||||
|
# Ensure we don't exceed buffer bounds
|
||||||
|
audio_length = len(audio)
|
||||||
|
end_sample = min(start_sample + audio_length, total_samples)
|
||||||
|
|
||||||
|
if start_sample < total_samples:
|
||||||
|
# Trim audio if it extends beyond buffer
|
||||||
|
audio_to_mix = audio[:end_sample - start_sample]
|
||||||
|
|
||||||
|
# Mix: add audio to existing buffer (overlapping audio will sum)
|
||||||
|
# Normalize to prevent clipping (simple approach: divide by max)
|
||||||
|
final_audio[start_sample:end_sample] += audio_to_mix
|
||||||
|
|
||||||
|
# Normalize to prevent clipping
|
||||||
|
max_val = np.abs(final_audio).max()
|
||||||
|
if max_val > 1.0:
|
||||||
|
final_audio = final_audio / max_val
|
||||||
|
|
||||||
|
# Save to temporary file
|
||||||
|
with tempfile.NamedTemporaryFile(suffix='.wav', delete=False) as tmp:
|
||||||
|
tmp_path = tmp.name
|
||||||
|
|
||||||
|
try:
|
||||||
|
save_audio(final_audio, tmp_path, sample_rate)
|
||||||
|
|
||||||
|
# Read file bytes
|
||||||
|
with open(tmp_path, 'rb') as f:
|
||||||
|
audio_bytes = f.read()
|
||||||
|
|
||||||
|
return audio_bytes
|
||||||
|
finally:
|
||||||
|
# Clean up temp file
|
||||||
|
Path(tmp_path).unlink(missing_ok=True)
|
||||||
+12
-264
@@ -1,274 +1,22 @@
|
|||||||
"""
|
"""
|
||||||
Whisper ASR module for transcription.
|
STT (Speech-to-Text) module - delegates to backend abstraction layer.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from typing import Optional, List, Dict
|
from typing import Optional
|
||||||
import asyncio
|
from .backends import get_stt_backend, STTBackend
|
||||||
import torch
|
|
||||||
import numpy as np
|
|
||||||
from pathlib import Path
|
|
||||||
from .utils.progress import get_progress_manager
|
|
||||||
from .utils.hf_progress import HFProgressTracker, create_hf_progress_callback
|
|
||||||
from .utils.tasks import get_task_manager
|
|
||||||
|
|
||||||
|
|
||||||
class WhisperModel:
|
def get_whisper_model() -> STTBackend:
|
||||||
"""Manages Whisper model loading and transcription."""
|
"""
|
||||||
|
Get STT backend instance (MLX or PyTorch based on platform).
|
||||||
|
|
||||||
def __init__(self, model_size: str = "base"):
|
Returns:
|
||||||
self.model = None
|
STT backend instance
|
||||||
self.processor = None
|
"""
|
||||||
self.model_size = model_size
|
return get_stt_backend()
|
||||||
self.device = self._get_device()
|
|
||||||
|
|
||||||
def _get_device(self) -> str:
|
|
||||||
"""Get the best available device."""
|
|
||||||
if torch.cuda.is_available():
|
|
||||||
return "cuda"
|
|
||||||
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
|
||||||
# MPS support for Whisper
|
|
||||||
return "cpu" # Use CPU for stability
|
|
||||||
return "cpu"
|
|
||||||
|
|
||||||
def is_loaded(self) -> bool:
|
|
||||||
"""Check if model is loaded."""
|
|
||||||
return self.model is not None
|
|
||||||
|
|
||||||
def load_model(self, model_size: Optional[str] = None):
|
|
||||||
"""
|
|
||||||
Lazy load the Whisper model.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
model_size: Model size (tiny, base, small, medium, large)
|
|
||||||
"""
|
|
||||||
if model_size is None:
|
|
||||||
model_size = self.model_size
|
|
||||||
|
|
||||||
if self.model is not None and self.model_size == model_size:
|
|
||||||
return
|
|
||||||
|
|
||||||
try:
|
|
||||||
from transformers import WhisperProcessor, WhisperForConditionalGeneration
|
|
||||||
|
|
||||||
model_name = f"openai/whisper-{model_size}"
|
|
||||||
|
|
||||||
# Set up progress tracking
|
|
||||||
progress_manager = get_progress_manager()
|
|
||||||
progress_model_name = f"whisper-{model_size}"
|
|
||||||
|
|
||||||
# Start tracking download task
|
|
||||||
task_manager = get_task_manager()
|
|
||||||
task_manager.start_download(progress_model_name)
|
|
||||||
|
|
||||||
print(f"Loading Whisper model {model_size} on {self.device}...")
|
|
||||||
|
|
||||||
# Initialize progress state to show download has started
|
|
||||||
progress_manager.update_progress(
|
|
||||||
model_name=progress_model_name,
|
|
||||||
current=0,
|
|
||||||
total=1, # Set to 1 initially, will be updated by callback
|
|
||||||
filename="",
|
|
||||||
status="downloading",
|
|
||||||
)
|
|
||||||
|
|
||||||
# Set up progress callback
|
|
||||||
progress_callback = create_hf_progress_callback(progress_model_name, progress_manager)
|
|
||||||
tracker = HFProgressTracker(progress_callback)
|
|
||||||
|
|
||||||
# Use progress tracker during download
|
|
||||||
with tracker.patch_download():
|
|
||||||
self.processor = WhisperProcessor.from_pretrained(model_name)
|
|
||||||
self.model = WhisperForConditionalGeneration.from_pretrained(model_name)
|
|
||||||
|
|
||||||
self.model.to(self.device)
|
|
||||||
self.model_size = model_size
|
|
||||||
|
|
||||||
# Mark as complete
|
|
||||||
progress_manager.mark_complete(progress_model_name)
|
|
||||||
task_manager.complete_download(progress_model_name)
|
|
||||||
|
|
||||||
print(f"Whisper model {model_size} loaded successfully")
|
|
||||||
|
|
||||||
except Exception as e:
|
|
||||||
print(f"Error loading Whisper model: {e}")
|
|
||||||
progress_manager = get_progress_manager()
|
|
||||||
task_manager = get_task_manager()
|
|
||||||
progress_model_name = f"whisper-{model_size}"
|
|
||||||
progress_manager.mark_error(progress_model_name, str(e))
|
|
||||||
task_manager.error_download(progress_model_name, str(e))
|
|
||||||
raise
|
|
||||||
|
|
||||||
async def load_model_async(self, model_size: Optional[str] = None):
|
|
||||||
"""
|
|
||||||
Async version of load_model that runs in thread pool.
|
|
||||||
|
|
||||||
This prevents blocking the event loop during model loading.
|
|
||||||
"""
|
|
||||||
if model_size is None:
|
|
||||||
model_size = self.model_size
|
|
||||||
|
|
||||||
# If already loaded with correct size, return immediately
|
|
||||||
if self.model is not None and self.model_size == model_size:
|
|
||||||
return
|
|
||||||
|
|
||||||
# Run the blocking load operation in a thread pool
|
|
||||||
await asyncio.to_thread(self.load_model, model_size)
|
|
||||||
|
|
||||||
def unload_model(self):
|
|
||||||
"""Unload the model to free memory."""
|
|
||||||
if self.model is not None:
|
|
||||||
del self.model
|
|
||||||
del self.processor
|
|
||||||
self.model = None
|
|
||||||
self.processor = None
|
|
||||||
|
|
||||||
if torch.cuda.is_available():
|
|
||||||
torch.cuda.empty_cache()
|
|
||||||
|
|
||||||
print("Whisper model unloaded")
|
|
||||||
|
|
||||||
async def transcribe(
|
|
||||||
self,
|
|
||||||
audio_path: str,
|
|
||||||
language: Optional[str] = None,
|
|
||||||
) -> str:
|
|
||||||
"""
|
|
||||||
Transcribe audio to text.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
audio_path: Path to audio file
|
|
||||||
language: Optional language hint (en or zh)
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Transcribed text
|
|
||||||
"""
|
|
||||||
await self.load_model_async()
|
|
||||||
|
|
||||||
from .utils.audio import load_audio
|
|
||||||
|
|
||||||
def _transcribe_sync():
|
|
||||||
"""Run synchronous transcription in thread pool."""
|
|
||||||
# Load audio
|
|
||||||
audio, sr = load_audio(audio_path, sample_rate=16000)
|
|
||||||
|
|
||||||
# Process audio
|
|
||||||
inputs = self.processor(
|
|
||||||
audio,
|
|
||||||
sampling_rate=16000,
|
|
||||||
return_tensors="pt",
|
|
||||||
)
|
|
||||||
inputs = inputs.to(self.device)
|
|
||||||
|
|
||||||
# Set language if provided
|
|
||||||
forced_decoder_ids = None
|
|
||||||
if language:
|
|
||||||
lang_code = "en" if language == "en" else "zh"
|
|
||||||
forced_decoder_ids = self.processor.get_decoder_prompt_ids(
|
|
||||||
language=lang_code,
|
|
||||||
task="transcribe",
|
|
||||||
)
|
|
||||||
|
|
||||||
# Generate transcription
|
|
||||||
with torch.no_grad():
|
|
||||||
predicted_ids = self.model.generate(
|
|
||||||
inputs["input_features"],
|
|
||||||
forced_decoder_ids=forced_decoder_ids,
|
|
||||||
)
|
|
||||||
|
|
||||||
# Decode
|
|
||||||
transcription = self.processor.batch_decode(
|
|
||||||
predicted_ids,
|
|
||||||
skip_special_tokens=True,
|
|
||||||
)[0]
|
|
||||||
|
|
||||||
return transcription.strip()
|
|
||||||
|
|
||||||
# Run blocking transcription in thread pool
|
|
||||||
return await asyncio.to_thread(_transcribe_sync)
|
|
||||||
|
|
||||||
async def transcribe_with_timestamps(
|
|
||||||
self,
|
|
||||||
audio_path: str,
|
|
||||||
language: Optional[str] = None,
|
|
||||||
) -> List[Dict[str, any]]:
|
|
||||||
"""
|
|
||||||
Transcribe audio with word-level timestamps.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
audio_path: Path to audio file
|
|
||||||
language: Optional language hint
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
List of word segments with timestamps
|
|
||||||
"""
|
|
||||||
await self.load_model_async()
|
|
||||||
|
|
||||||
from .utils.audio import load_audio
|
|
||||||
|
|
||||||
def _transcribe_timestamps_sync():
|
|
||||||
"""Run synchronous transcription with timestamps in thread pool."""
|
|
||||||
# Load audio
|
|
||||||
audio, sr = load_audio(audio_path, sample_rate=16000)
|
|
||||||
|
|
||||||
# Process audio
|
|
||||||
inputs = self.processor(
|
|
||||||
audio,
|
|
||||||
sampling_rate=16000,
|
|
||||||
return_tensors="pt",
|
|
||||||
)
|
|
||||||
inputs = inputs.to(self.device)
|
|
||||||
|
|
||||||
# Set language if provided
|
|
||||||
forced_decoder_ids = None
|
|
||||||
if language:
|
|
||||||
lang_code = "en" if language == "en" else "zh"
|
|
||||||
forced_decoder_ids = self.processor.get_decoder_prompt_ids(
|
|
||||||
language=lang_code,
|
|
||||||
task="transcribe",
|
|
||||||
)
|
|
||||||
|
|
||||||
# Generate with timestamps
|
|
||||||
with torch.no_grad():
|
|
||||||
predicted_ids = self.model.generate(
|
|
||||||
inputs["input_features"],
|
|
||||||
forced_decoder_ids=forced_decoder_ids,
|
|
||||||
return_timestamps=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
# Parse timestamps (simplified - would need more robust parsing)
|
|
||||||
# For now, return basic transcription
|
|
||||||
# TODO: Implement proper timestamp parsing
|
|
||||||
transcription = self.processor.batch_decode(
|
|
||||||
predicted_ids,
|
|
||||||
skip_special_tokens=True,
|
|
||||||
)[0]
|
|
||||||
|
|
||||||
return [
|
|
||||||
{
|
|
||||||
"text": transcription,
|
|
||||||
"start": 0.0,
|
|
||||||
"end": len(audio) / sr,
|
|
||||||
}
|
|
||||||
]
|
|
||||||
|
|
||||||
# Run blocking transcription in thread pool
|
|
||||||
return await asyncio.to_thread(_transcribe_timestamps_sync)
|
|
||||||
|
|
||||||
|
|
||||||
# Global model instance
|
|
||||||
_whisper_model: Optional[WhisperModel] = None
|
|
||||||
|
|
||||||
|
|
||||||
def get_whisper_model() -> WhisperModel:
|
|
||||||
"""Get or create Whisper model instance."""
|
|
||||||
global _whisper_model
|
|
||||||
if _whisper_model is None:
|
|
||||||
_whisper_model = WhisperModel()
|
|
||||||
return _whisper_model
|
|
||||||
|
|
||||||
|
|
||||||
def unload_whisper_model():
|
def unload_whisper_model():
|
||||||
"""Unload Whisper model to free memory."""
|
"""Unload Whisper model to free memory."""
|
||||||
global _whisper_model
|
backend = get_stt_backend()
|
||||||
if _whisper_model is not None:
|
backend.unload_model()
|
||||||
_whisper_model.unload_model()
|
|
||||||
|
|||||||
+20
-355
@@ -1,372 +1,37 @@
|
|||||||
"""
|
"""
|
||||||
TTS inference module using Qwen3-TTS.
|
TTS inference module - delegates to backend abstraction layer.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from typing import Optional, List, Tuple
|
from typing import Optional
|
||||||
import asyncio
|
|
||||||
import torch
|
|
||||||
import numpy as np
|
import numpy as np
|
||||||
import io
|
import io
|
||||||
import soundfile as sf
|
import soundfile as sf
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
from .utils.cache import get_cache_key, get_cached_voice_prompt, cache_voice_prompt
|
from .backends import get_tts_backend, TTSBackend
|
||||||
from .utils.audio import normalize_audio
|
|
||||||
from .utils.progress import get_progress_manager
|
|
||||||
from .utils.hf_progress import HFProgressTracker, create_hf_progress_callback
|
|
||||||
from .utils.tasks import get_task_manager
|
|
||||||
from . import config
|
|
||||||
|
|
||||||
|
|
||||||
class TTSModel:
|
def get_tts_model() -> TTSBackend:
|
||||||
"""Manages Qwen3-TTS model loading and inference."""
|
"""
|
||||||
|
Get TTS backend instance (MLX or PyTorch based on platform).
|
||||||
|
|
||||||
def __init__(self, model_size: str = "1.7B"):
|
Returns:
|
||||||
self.model = None
|
TTS backend instance
|
||||||
self.model_size = model_size
|
"""
|
||||||
self.device = self._get_device()
|
return get_tts_backend()
|
||||||
self._current_model_size = None
|
|
||||||
|
|
||||||
def _get_device(self) -> str:
|
|
||||||
"""Get the best available device."""
|
|
||||||
if torch.cuda.is_available():
|
|
||||||
return "cuda"
|
|
||||||
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
|
|
||||||
# MPS can have issues, use CPU for stability
|
|
||||||
return "cpu"
|
|
||||||
return "cpu"
|
|
||||||
|
|
||||||
def is_loaded(self) -> bool:
|
|
||||||
"""Check if model is loaded."""
|
|
||||||
return self.model is not None
|
|
||||||
|
|
||||||
def _get_model_path(self, model_size: str) -> str:
|
|
||||||
"""
|
|
||||||
Get the model path, downloading from HuggingFace Hub if needed.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
model_size: Model size (1.7B or 0.6B)
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Path to model (either local or HuggingFace Hub ID)
|
|
||||||
"""
|
|
||||||
# HuggingFace Hub model IDs
|
|
||||||
hf_model_map = {
|
|
||||||
"1.7B": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
|
|
||||||
"0.6B": "Qwen/Qwen3-TTS-12Hz-0.6B-Base",
|
|
||||||
}
|
|
||||||
|
|
||||||
# Local directory names (for backwards compatibility)
|
|
||||||
local_model_map = {
|
|
||||||
"1.7B": "Qwen--Qwen3-TTS-12Hz-1.7B-Base",
|
|
||||||
"0.6B": "Qwen--Qwen3-TTS-12Hz-0.6B-Base",
|
|
||||||
}
|
|
||||||
|
|
||||||
if model_size not in hf_model_map:
|
|
||||||
raise ValueError(f"Unknown model size: {model_size}")
|
|
||||||
|
|
||||||
# Check if model exists locally (backwards compatibility)
|
|
||||||
local_path = config.get_models_dir() / local_model_map[model_size]
|
|
||||||
if local_path.exists():
|
|
||||||
print(f"Found local model at {local_path}")
|
|
||||||
return str(local_path)
|
|
||||||
|
|
||||||
# Use HuggingFace Hub model ID (will auto-download)
|
|
||||||
hf_model_id = hf_model_map[model_size]
|
|
||||||
print(f"Will download model from HuggingFace Hub: {hf_model_id}")
|
|
||||||
|
|
||||||
return hf_model_id
|
|
||||||
|
|
||||||
def load_model(self, model_size: Optional[str] = None):
|
|
||||||
"""
|
|
||||||
Lazy load the TTS model with automatic downloading from HuggingFace Hub.
|
|
||||||
|
|
||||||
The model will be automatically downloaded on first use and cached locally.
|
|
||||||
This works similar to how Whisper models are loaded.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
model_size: Model size to load (1.7B or 0.6B)
|
|
||||||
"""
|
|
||||||
if model_size is None:
|
|
||||||
model_size = self.model_size
|
|
||||||
|
|
||||||
# If already loaded with correct size, return
|
|
||||||
if self.model is not None and self._current_model_size == model_size:
|
|
||||||
return
|
|
||||||
|
|
||||||
# Unload existing model if different size requested
|
|
||||||
if self.model is not None and self._current_model_size != model_size:
|
|
||||||
self.unload_model()
|
|
||||||
|
|
||||||
try:
|
|
||||||
from qwen_tts import Qwen3TTSModel
|
|
||||||
|
|
||||||
# Get model path (local or HuggingFace Hub ID)
|
|
||||||
model_path = self._get_model_path(model_size)
|
|
||||||
|
|
||||||
# Set up progress tracking
|
|
||||||
progress_manager = get_progress_manager()
|
|
||||||
model_name = f"qwen-tts-{model_size}"
|
|
||||||
|
|
||||||
# Check if model is being downloaded from HuggingFace Hub
|
|
||||||
if model_path.startswith("Qwen/"):
|
|
||||||
print(f"Loading TTS model {model_size} on {self.device}...")
|
|
||||||
|
|
||||||
# Start tracking download task
|
|
||||||
task_manager = get_task_manager()
|
|
||||||
task_manager.start_download(model_name)
|
|
||||||
|
|
||||||
# Initialize progress state to show download has started
|
|
||||||
progress_manager.update_progress(
|
|
||||||
model_name=model_name,
|
|
||||||
current=0,
|
|
||||||
total=1, # Set to 1 initially, will be updated by callback
|
|
||||||
filename="",
|
|
||||||
status="downloading",
|
|
||||||
)
|
|
||||||
|
|
||||||
# Set up progress callback
|
|
||||||
progress_callback = create_hf_progress_callback(model_name, progress_manager)
|
|
||||||
tracker = HFProgressTracker(progress_callback)
|
|
||||||
|
|
||||||
# Use progress tracker during download
|
|
||||||
with tracker.patch_download():
|
|
||||||
# Load the model - downloads will happen automatically with progress tracking
|
|
||||||
self.model = Qwen3TTSModel.from_pretrained(
|
|
||||||
model_path,
|
|
||||||
device_map=self.device,
|
|
||||||
torch_dtype=torch.float32 if self.device == "cpu" else torch.bfloat16,
|
|
||||||
)
|
|
||||||
|
|
||||||
# Mark as complete
|
|
||||||
progress_manager.mark_complete(model_name)
|
|
||||||
task_manager.complete_download(model_name)
|
|
||||||
else:
|
|
||||||
# Local model, no download needed
|
|
||||||
print(f"Loading TTS model {model_size} on {self.device}...")
|
|
||||||
self.model = Qwen3TTSModel.from_pretrained(
|
|
||||||
model_path,
|
|
||||||
device_map=self.device,
|
|
||||||
torch_dtype=torch.float32 if self.device == "cpu" else torch.bfloat16,
|
|
||||||
)
|
|
||||||
|
|
||||||
self._current_model_size = model_size
|
|
||||||
self.model_size = model_size
|
|
||||||
|
|
||||||
print(f"TTS model {model_size} loaded successfully")
|
|
||||||
|
|
||||||
except ImportError as e:
|
|
||||||
print(f"Error: qwen_tts package not found. Install with: pip install git+https://github.com/QwenLM/Qwen3-TTS.git")
|
|
||||||
progress_manager = get_progress_manager()
|
|
||||||
task_manager = get_task_manager()
|
|
||||||
model_name = f"qwen-tts-{model_size}"
|
|
||||||
progress_manager.mark_error(model_name, str(e))
|
|
||||||
task_manager.error_download(model_name, str(e))
|
|
||||||
raise
|
|
||||||
except Exception as e:
|
|
||||||
print(f"Error loading TTS model: {e}")
|
|
||||||
print(f"Tip: The model will be automatically downloaded from HuggingFace Hub on first use.")
|
|
||||||
progress_manager = get_progress_manager()
|
|
||||||
task_manager = get_task_manager()
|
|
||||||
model_name = f"qwen-tts-{model_size}"
|
|
||||||
progress_manager.mark_error(model_name, str(e))
|
|
||||||
task_manager.error_download(model_name, str(e))
|
|
||||||
raise
|
|
||||||
|
|
||||||
async def load_model_async(self, model_size: Optional[str] = None):
|
|
||||||
"""
|
|
||||||
Async version of load_model that runs in thread pool.
|
|
||||||
|
|
||||||
This prevents blocking the event loop during model loading.
|
|
||||||
"""
|
|
||||||
if model_size is None:
|
|
||||||
model_size = self.model_size
|
|
||||||
|
|
||||||
# If already loaded with correct size, return immediately
|
|
||||||
if self.model is not None and self._current_model_size == model_size:
|
|
||||||
return
|
|
||||||
|
|
||||||
# Run the blocking load operation in a thread pool
|
|
||||||
await asyncio.to_thread(self.load_model, model_size)
|
|
||||||
|
|
||||||
def unload_model(self):
|
|
||||||
"""Unload the model to free memory."""
|
|
||||||
if self.model is not None:
|
|
||||||
del self.model
|
|
||||||
self.model = None
|
|
||||||
self._current_model_size = None
|
|
||||||
|
|
||||||
if torch.cuda.is_available():
|
|
||||||
torch.cuda.empty_cache()
|
|
||||||
|
|
||||||
print("TTS model unloaded")
|
|
||||||
|
|
||||||
async def create_voice_prompt(
|
|
||||||
self,
|
|
||||||
audio_path: str,
|
|
||||||
reference_text: str,
|
|
||||||
use_cache: bool = True,
|
|
||||||
) -> Tuple[dict, bool]:
|
|
||||||
"""
|
|
||||||
Create voice prompt from reference audio.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
audio_path: Path to reference audio file
|
|
||||||
reference_text: Transcript of reference audio
|
|
||||||
use_cache: Whether to use cached prompt if available
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Tuple of (voice_prompt_dict, was_cached)
|
|
||||||
"""
|
|
||||||
await self.load_model_async()
|
|
||||||
|
|
||||||
# Check cache if enabled
|
|
||||||
if use_cache:
|
|
||||||
cache_key = get_cache_key(audio_path, reference_text)
|
|
||||||
cached_prompt = get_cached_voice_prompt(cache_key)
|
|
||||||
if cached_prompt is not None:
|
|
||||||
return cached_prompt, True
|
|
||||||
|
|
||||||
def _create_prompt_sync():
|
|
||||||
"""Run synchronous voice prompt creation in thread pool."""
|
|
||||||
return self.model.create_voice_clone_prompt(
|
|
||||||
ref_audio=str(audio_path),
|
|
||||||
ref_text=reference_text,
|
|
||||||
x_vector_only_mode=False,
|
|
||||||
)
|
|
||||||
|
|
||||||
# Run blocking operation in thread pool
|
|
||||||
voice_prompt_items = await asyncio.to_thread(_create_prompt_sync)
|
|
||||||
|
|
||||||
# Cache if enabled
|
|
||||||
if use_cache:
|
|
||||||
cache_voice_prompt(cache_key, voice_prompt_items)
|
|
||||||
|
|
||||||
return voice_prompt_items, False
|
|
||||||
|
|
||||||
async def combine_voice_prompts(
|
|
||||||
self,
|
|
||||||
audio_paths: List[str],
|
|
||||||
reference_texts: List[str],
|
|
||||||
) -> Tuple[np.ndarray, str]:
|
|
||||||
"""
|
|
||||||
Combine multiple reference samples for better quality.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
audio_paths: List of audio file paths
|
|
||||||
reference_texts: List of reference texts
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Tuple of (combined_audio, combined_text)
|
|
||||||
"""
|
|
||||||
from .utils.audio import load_audio
|
|
||||||
|
|
||||||
combined_audio = []
|
|
||||||
|
|
||||||
for audio_path in audio_paths:
|
|
||||||
audio, sr = load_audio(audio_path)
|
|
||||||
audio = normalize_audio(audio)
|
|
||||||
combined_audio.append(audio)
|
|
||||||
|
|
||||||
# Concatenate audio
|
|
||||||
mixed = np.concatenate(combined_audio)
|
|
||||||
mixed = normalize_audio(mixed)
|
|
||||||
|
|
||||||
# Combine texts
|
|
||||||
combined_text = " ".join(reference_texts)
|
|
||||||
|
|
||||||
return mixed, combined_text
|
|
||||||
|
|
||||||
async def generate(
|
|
||||||
self,
|
|
||||||
text: str,
|
|
||||||
voice_prompt: dict,
|
|
||||||
language: str = "en",
|
|
||||||
seed: Optional[int] = None,
|
|
||||||
instruct: Optional[str] = None,
|
|
||||||
) -> Tuple[np.ndarray, int]:
|
|
||||||
"""
|
|
||||||
Generate audio from text using voice prompt.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
text: Text to synthesize
|
|
||||||
voice_prompt: Voice prompt dictionary from create_voice_prompt
|
|
||||||
language: Language code (en or zh)
|
|
||||||
seed: Random seed for reproducibility
|
|
||||||
instruct: Natural language instruction for speech delivery control
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Tuple of (audio_array, sample_rate)
|
|
||||||
"""
|
|
||||||
# Load model (already handles async via to_thread if needed)
|
|
||||||
await self.load_model_async()
|
|
||||||
|
|
||||||
def _generate_sync():
|
|
||||||
"""Run synchronous generation in thread pool."""
|
|
||||||
# Set seed if provided
|
|
||||||
if seed is not None:
|
|
||||||
torch.manual_seed(seed)
|
|
||||||
if torch.cuda.is_available():
|
|
||||||
torch.cuda.manual_seed(seed)
|
|
||||||
|
|
||||||
# Generate audio - this is the blocking operation
|
|
||||||
wavs, sample_rate = self.model.generate_voice_clone(
|
|
||||||
text=text,
|
|
||||||
voice_clone_prompt=voice_prompt,
|
|
||||||
instruct=instruct,
|
|
||||||
)
|
|
||||||
return wavs[0], sample_rate
|
|
||||||
|
|
||||||
# Run blocking inference in thread pool to avoid blocking event loop
|
|
||||||
audio, sample_rate = await asyncio.to_thread(_generate_sync)
|
|
||||||
|
|
||||||
return audio, sample_rate
|
|
||||||
|
|
||||||
async def generate_from_reference(
|
|
||||||
self,
|
|
||||||
text: str,
|
|
||||||
audio_path: str,
|
|
||||||
reference_text: str,
|
|
||||||
language: str = "en",
|
|
||||||
seed: Optional[int] = None,
|
|
||||||
) -> Tuple[np.ndarray, int]:
|
|
||||||
"""
|
|
||||||
Generate audio directly from reference (convenience method).
|
|
||||||
|
|
||||||
Args:
|
|
||||||
text: Text to synthesize
|
|
||||||
audio_path: Path to reference audio
|
|
||||||
reference_text: Transcript of reference audio
|
|
||||||
language: Language code
|
|
||||||
seed: Random seed
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Tuple of (audio_array, sample_rate)
|
|
||||||
"""
|
|
||||||
# Create voice prompt (with caching)
|
|
||||||
voice_prompt, _ = await self.create_voice_prompt(audio_path, reference_text)
|
|
||||||
|
|
||||||
# Generate
|
|
||||||
return await self.generate(text, voice_prompt, language, seed)
|
|
||||||
|
|
||||||
|
|
||||||
# Global model instance
|
|
||||||
_tts_model: Optional[TTSModel] = None
|
|
||||||
|
|
||||||
|
|
||||||
def get_tts_model() -> TTSModel:
|
|
||||||
"""Get or create TTS model instance."""
|
|
||||||
global _tts_model
|
|
||||||
if _tts_model is None:
|
|
||||||
_tts_model = TTSModel()
|
|
||||||
return _tts_model
|
|
||||||
|
|
||||||
|
|
||||||
def unload_tts_model():
|
def unload_tts_model():
|
||||||
"""Unload TTS model to free memory."""
|
"""Unload TTS model to free memory."""
|
||||||
global _tts_model
|
backend = get_tts_backend()
|
||||||
if _tts_model is not None:
|
backend.unload_model()
|
||||||
_tts_model.unload_model()
|
|
||||||
|
|
||||||
|
def audio_to_wav_bytes(audio: np.ndarray, sample_rate: int) -> bytes:
|
||||||
|
"""Convert audio array to WAV bytes."""
|
||||||
|
buffer = io.BytesIO()
|
||||||
|
sf.write(buffer, audio, sample_rate, format="WAV")
|
||||||
|
buffer.seek(0)
|
||||||
|
return buffer.read()
|
||||||
|
|
||||||
|
|
||||||
def audio_to_wav_bytes(audio: np.ndarray, sample_rate: int) -> bytes:
|
def audio_to_wav_bytes(audio: np.ndarray, sample_rate: int) -> bytes:
|
||||||
|
|||||||
+68
-8
@@ -5,7 +5,7 @@ Voice prompt caching utilities.
|
|||||||
import hashlib
|
import hashlib
|
||||||
import torch
|
import torch
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Optional
|
from typing import Optional, Union, Dict, Any
|
||||||
|
|
||||||
from .. import config
|
from .. import config
|
||||||
|
|
||||||
@@ -15,8 +15,8 @@ def _get_cache_dir() -> Path:
|
|||||||
return config.get_cache_dir()
|
return config.get_cache_dir()
|
||||||
|
|
||||||
|
|
||||||
# In-memory cache
|
# In-memory cache - can store dict (voice prompt) or tensor (legacy)
|
||||||
_memory_cache: dict[str, torch.Tensor] = {}
|
_memory_cache: dict[str, Union[torch.Tensor, Dict[str, Any]]] = {}
|
||||||
|
|
||||||
|
|
||||||
def get_cache_key(audio_path: str, reference_text: str) -> str:
|
def get_cache_key(audio_path: str, reference_text: str) -> str:
|
||||||
@@ -43,7 +43,7 @@ def get_cache_key(audio_path: str, reference_text: str) -> str:
|
|||||||
|
|
||||||
def get_cached_voice_prompt(
|
def get_cached_voice_prompt(
|
||||||
cache_key: str,
|
cache_key: str,
|
||||||
) -> Optional[torch.Tensor]:
|
) -> Optional[Union[torch.Tensor, Dict[str, Any]]]:
|
||||||
"""
|
"""
|
||||||
Get cached voice prompt if available.
|
Get cached voice prompt if available.
|
||||||
|
|
||||||
@@ -51,7 +51,7 @@ def get_cached_voice_prompt(
|
|||||||
cache_key: Cache key
|
cache_key: Cache key
|
||||||
|
|
||||||
Returns:
|
Returns:
|
||||||
Cached voice prompt tensor or None
|
Cached voice prompt (dict or tensor) or None
|
||||||
"""
|
"""
|
||||||
# Check in-memory cache
|
# Check in-memory cache
|
||||||
if cache_key in _memory_cache:
|
if cache_key in _memory_cache:
|
||||||
@@ -73,18 +73,78 @@ def get_cached_voice_prompt(
|
|||||||
|
|
||||||
def cache_voice_prompt(
|
def cache_voice_prompt(
|
||||||
cache_key: str,
|
cache_key: str,
|
||||||
voice_prompt: torch.Tensor,
|
voice_prompt: Union[torch.Tensor, Dict[str, Any]],
|
||||||
) -> None:
|
) -> None:
|
||||||
"""
|
"""
|
||||||
Cache voice prompt to memory and disk.
|
Cache voice prompt to memory and disk.
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
cache_key: Cache key
|
cache_key: Cache key
|
||||||
voice_prompt: Voice prompt tensor
|
voice_prompt: Voice prompt (dict or tensor)
|
||||||
"""
|
"""
|
||||||
# Store in memory
|
# Store in memory
|
||||||
_memory_cache[cache_key] = voice_prompt
|
_memory_cache[cache_key] = voice_prompt
|
||||||
|
|
||||||
# Store on disk
|
# Store on disk (torch.save can handle both dicts and tensors)
|
||||||
cache_file = _get_cache_dir() / f"{cache_key}.prompt"
|
cache_file = _get_cache_dir() / f"{cache_key}.prompt"
|
||||||
torch.save(voice_prompt, cache_file)
|
torch.save(voice_prompt, cache_file)
|
||||||
|
|
||||||
|
|
||||||
|
def clear_voice_prompt_cache() -> int:
|
||||||
|
"""
|
||||||
|
Clear all voice prompt caches (memory and disk).
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Number of cache files deleted
|
||||||
|
"""
|
||||||
|
# Clear memory cache
|
||||||
|
_memory_cache.clear()
|
||||||
|
|
||||||
|
# Clear disk cache
|
||||||
|
cache_dir = _get_cache_dir()
|
||||||
|
deleted_count = 0
|
||||||
|
|
||||||
|
if cache_dir.exists():
|
||||||
|
# Delete prompt cache files
|
||||||
|
for cache_file in cache_dir.glob("*.prompt"):
|
||||||
|
try:
|
||||||
|
cache_file.unlink()
|
||||||
|
deleted_count += 1
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Failed to delete cache file {cache_file}: {e}")
|
||||||
|
|
||||||
|
# Delete combined audio files
|
||||||
|
for audio_file in cache_dir.glob("combined_*.wav"):
|
||||||
|
try:
|
||||||
|
audio_file.unlink()
|
||||||
|
deleted_count += 1
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Failed to delete combined audio file {audio_file}: {e}")
|
||||||
|
|
||||||
|
return deleted_count
|
||||||
|
|
||||||
|
|
||||||
|
def clear_profile_cache(profile_id: str) -> int:
|
||||||
|
"""
|
||||||
|
Clear cache files for a specific profile.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
profile_id: Profile ID
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Number of cache files deleted
|
||||||
|
"""
|
||||||
|
cache_dir = _get_cache_dir()
|
||||||
|
deleted_count = 0
|
||||||
|
|
||||||
|
if cache_dir.exists():
|
||||||
|
# Delete combined audio files for this profile
|
||||||
|
pattern = f"combined_{profile_id}_*.wav"
|
||||||
|
for audio_file in cache_dir.glob(pattern):
|
||||||
|
try:
|
||||||
|
audio_file.unlink()
|
||||||
|
deleted_count += 1
|
||||||
|
except Exception as e:
|
||||||
|
print(f"Failed to delete combined audio file {audio_file}: {e}")
|
||||||
|
|
||||||
|
return deleted_count
|
||||||
|
|||||||
@@ -29,8 +29,9 @@ class HFProgressTracker:
|
|||||||
|
|
||||||
class TrackedTqdm(original_tqdm):
|
class TrackedTqdm(original_tqdm):
|
||||||
"""A tqdm subclass that reports progress to our tracker."""
|
"""A tqdm subclass that reports progress to our tracker."""
|
||||||
|
|
||||||
def __init__(self, *args, **kwargs):
|
def __init__(self, *args, **kwargs):
|
||||||
|
print(f"[DEBUG TrackedTqdm] __init__ called with desc: {kwargs.get('desc', '')}")
|
||||||
# Extract filename from desc before passing to parent
|
# Extract filename from desc before passing to parent
|
||||||
desc = kwargs.get("desc", "")
|
desc = kwargs.get("desc", "")
|
||||||
if not desc and args:
|
if not desc and args:
|
||||||
@@ -79,8 +80,9 @@ class HFProgressTracker:
|
|||||||
}
|
}
|
||||||
|
|
||||||
def update(self, n=1):
|
def update(self, n=1):
|
||||||
|
print(f"[DEBUG TrackedTqdm] update called with n={n}")
|
||||||
result = super().update(n)
|
result = super().update(n)
|
||||||
|
|
||||||
# Report progress
|
# Report progress
|
||||||
with tracker._lock:
|
with tracker._lock:
|
||||||
if id(self) in tracker._active_tqdms:
|
if id(self) in tracker._active_tqdms:
|
||||||
@@ -118,11 +120,13 @@ class HFProgressTracker:
|
|||||||
@contextmanager
|
@contextmanager
|
||||||
def patch_download(self):
|
def patch_download(self):
|
||||||
"""Context manager to patch tqdm for progress tracking."""
|
"""Context manager to patch tqdm for progress tracking."""
|
||||||
|
print("[DEBUG HFProgressTracker] patch_download called")
|
||||||
try:
|
try:
|
||||||
import tqdm as tqdm_module
|
import tqdm as tqdm_module
|
||||||
|
|
||||||
# Store original tqdm class
|
# Store original tqdm class
|
||||||
self._original_tqdm_class = tqdm_module.tqdm
|
self._original_tqdm_class = tqdm_module.tqdm
|
||||||
|
print(f"[DEBUG HFProgressTracker] Original tqdm class: {self._original_tqdm_class}")
|
||||||
|
|
||||||
# Reset totals
|
# Reset totals
|
||||||
with self._lock:
|
with self._lock:
|
||||||
@@ -135,18 +139,22 @@ class HFProgressTracker:
|
|||||||
|
|
||||||
# Create our tracked tqdm class
|
# Create our tracked tqdm class
|
||||||
tracked_tqdm = self._create_tracked_tqdm_class()
|
tracked_tqdm = self._create_tracked_tqdm_class()
|
||||||
|
print(f"[DEBUG HFProgressTracker] Created TrackedTqdm class: {tracked_tqdm}")
|
||||||
|
|
||||||
# Patch tqdm.tqdm
|
# Patch tqdm.tqdm
|
||||||
tqdm_module.tqdm = tracked_tqdm
|
tqdm_module.tqdm = tracked_tqdm
|
||||||
|
print(f"[DEBUG HFProgressTracker] Patched tqdm.tqdm")
|
||||||
|
|
||||||
# Also patch tqdm.auto.tqdm if it exists (used by huggingface_hub)
|
# Also patch tqdm.auto.tqdm if it exists (used by huggingface_hub)
|
||||||
self._original_tqdm_auto = None
|
self._original_tqdm_auto = None
|
||||||
if hasattr(tqdm_module, "auto") and hasattr(tqdm_module.auto, "tqdm"):
|
if hasattr(tqdm_module, "auto") and hasattr(tqdm_module.auto, "tqdm"):
|
||||||
self._original_tqdm_auto = tqdm_module.auto.tqdm
|
self._original_tqdm_auto = tqdm_module.auto.tqdm
|
||||||
tqdm_module.auto.tqdm = tracked_tqdm
|
tqdm_module.auto.tqdm = tracked_tqdm
|
||||||
|
print(f"[DEBUG HFProgressTracker] Patched tqdm.auto.tqdm")
|
||||||
|
|
||||||
# Patch in sys.modules to catch already-imported references
|
# Patch in sys.modules to catch already-imported references
|
||||||
self._patched_modules = {}
|
self._patched_modules = {}
|
||||||
|
patched_count = 0
|
||||||
for module_name in list(sys.modules.keys()):
|
for module_name in list(sys.modules.keys()):
|
||||||
if "huggingface" in module_name or module_name.startswith("tqdm"):
|
if "huggingface" in module_name or module_name.startswith("tqdm"):
|
||||||
try:
|
try:
|
||||||
@@ -159,8 +167,11 @@ class HFProgressTracker:
|
|||||||
):
|
):
|
||||||
self._patched_modules[module_name] = attr
|
self._patched_modules[module_name] = attr
|
||||||
setattr(module, "tqdm", tracked_tqdm)
|
setattr(module, "tqdm", tracked_tqdm)
|
||||||
|
patched_count += 1
|
||||||
|
print(f"[DEBUG HFProgressTracker] Patched {module_name}.tqdm")
|
||||||
except (AttributeError, TypeError):
|
except (AttributeError, TypeError):
|
||||||
pass
|
pass
|
||||||
|
print(f"[DEBUG HFProgressTracker] Patched {patched_count} modules in sys.modules")
|
||||||
|
|
||||||
yield
|
yield
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,114 @@
|
|||||||
|
"""Image processing utilities for avatar uploads."""
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Optional, Tuple
|
||||||
|
from PIL import Image
|
||||||
|
|
||||||
|
# JPEG can be reported as 'JPEG' or 'MPO' (for multi-picture format from some cameras)
|
||||||
|
ALLOWED_FORMATS = {'PNG', 'JPEG', 'WEBP', 'MPO', 'JPG'}
|
||||||
|
MAX_SIZE = 512
|
||||||
|
MAX_FILE_SIZE = 5 * 1024 * 1024 # 5MB
|
||||||
|
|
||||||
|
|
||||||
|
def validate_image(file_path: str) -> Tuple[bool, Optional[str]]:
|
||||||
|
"""
|
||||||
|
Validate image format and file size.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
file_path: Path to image file
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Tuple of (is_valid, error_message)
|
||||||
|
"""
|
||||||
|
path = Path(file_path)
|
||||||
|
|
||||||
|
# Check file size
|
||||||
|
if path.stat().st_size > MAX_FILE_SIZE:
|
||||||
|
return False, f"File size exceeds maximum of {MAX_FILE_SIZE // (1024 * 1024)}MB"
|
||||||
|
|
||||||
|
try:
|
||||||
|
with Image.open(file_path) as img:
|
||||||
|
# Verify the image can be loaded
|
||||||
|
img.load()
|
||||||
|
|
||||||
|
# Check format (normalize JPEG variants)
|
||||||
|
img_format = img.format
|
||||||
|
if img_format in ('MPO', 'JPG'):
|
||||||
|
img_format = 'JPEG'
|
||||||
|
|
||||||
|
if img_format not in {'PNG', 'JPEG', 'WEBP'}:
|
||||||
|
return False, f"Invalid format '{img_format}'. Allowed formats: PNG, JPEG, WEBP"
|
||||||
|
|
||||||
|
return True, None
|
||||||
|
except Exception as e:
|
||||||
|
return False, f"Invalid image file: {str(e)}"
|
||||||
|
|
||||||
|
|
||||||
|
def process_avatar(input_path: str, output_path: str, max_size: int = MAX_SIZE) -> None:
|
||||||
|
"""
|
||||||
|
Process avatar image: resize and optimize.
|
||||||
|
|
||||||
|
Resizes image to fit within max_size x max_size while maintaining aspect ratio.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
input_path: Path to input image
|
||||||
|
output_path: Path to save processed image
|
||||||
|
max_size: Maximum width or height in pixels
|
||||||
|
"""
|
||||||
|
with Image.open(input_path) as img:
|
||||||
|
# Handle EXIF orientation for JPEG images
|
||||||
|
try:
|
||||||
|
from PIL import ExifTags
|
||||||
|
for orientation in ExifTags.TAGS.keys():
|
||||||
|
if ExifTags.TAGS[orientation] == 'Orientation':
|
||||||
|
break
|
||||||
|
exif = img._getexif()
|
||||||
|
if exif is not None:
|
||||||
|
orientation_value = exif.get(orientation)
|
||||||
|
if orientation_value == 3:
|
||||||
|
img = img.rotate(180, expand=True)
|
||||||
|
elif orientation_value == 6:
|
||||||
|
img = img.rotate(270, expand=True)
|
||||||
|
elif orientation_value == 8:
|
||||||
|
img = img.rotate(90, expand=True)
|
||||||
|
except (AttributeError, KeyError, IndexError, TypeError):
|
||||||
|
# No EXIF data or orientation tag
|
||||||
|
pass
|
||||||
|
|
||||||
|
# Convert to RGB if necessary (handles RGBA, P, CMYK, etc.)
|
||||||
|
if img.mode not in ('RGB', 'L'):
|
||||||
|
if img.mode == 'RGBA':
|
||||||
|
# Create white background for RGBA images
|
||||||
|
background = Image.new('RGB', img.size, (255, 255, 255))
|
||||||
|
background.paste(img, mask=img.split()[3]) # Use alpha channel as mask
|
||||||
|
img = background
|
||||||
|
elif img.mode == 'CMYK':
|
||||||
|
# Convert CMYK to RGB
|
||||||
|
img = img.convert('RGB')
|
||||||
|
elif img.mode == 'P':
|
||||||
|
# Convert palette mode to RGB
|
||||||
|
img = img.convert('RGB')
|
||||||
|
else:
|
||||||
|
img = img.convert('RGB')
|
||||||
|
|
||||||
|
# Calculate new size maintaining aspect ratio
|
||||||
|
img.thumbnail((max_size, max_size), Image.Resampling.LANCZOS)
|
||||||
|
|
||||||
|
# Determine output format from extension
|
||||||
|
output_ext = Path(output_path).suffix.lower()
|
||||||
|
|
||||||
|
format_map = {
|
||||||
|
'.png': 'PNG',
|
||||||
|
'.jpeg': 'JPEG',
|
||||||
|
'.jpg': 'JPEG',
|
||||||
|
'.webp': 'WEBP'
|
||||||
|
}
|
||||||
|
|
||||||
|
output_format = format_map.get(output_ext, 'PNG')
|
||||||
|
|
||||||
|
# Save with optimization
|
||||||
|
save_kwargs = {'optimize': True}
|
||||||
|
if output_format == 'JPEG':
|
||||||
|
save_kwargs['quality'] = 90
|
||||||
|
|
||||||
|
img.save(output_path, format=output_format, **save_kwargs)
|
||||||
+157
-48
@@ -6,16 +6,55 @@ from typing import Optional, Callable, Dict, List
|
|||||||
from fastapi.responses import StreamingResponse
|
from fastapi.responses import StreamingResponse
|
||||||
import asyncio
|
import asyncio
|
||||||
import json
|
import json
|
||||||
|
import threading
|
||||||
from datetime import datetime
|
from datetime import datetime
|
||||||
|
|
||||||
|
|
||||||
class ProgressManager:
|
class ProgressManager:
|
||||||
"""Manages download progress for multiple models."""
|
"""Manages download progress for multiple models.
|
||||||
|
|
||||||
|
Thread-safe: can be called from background threads (e.g., via asyncio.to_thread).
|
||||||
|
"""
|
||||||
|
|
||||||
def __init__(self):
|
def __init__(self):
|
||||||
self._progress: Dict[str, Dict] = {}
|
self._progress: Dict[str, Dict] = {}
|
||||||
self._listeners: Dict[str, list] = {}
|
self._listeners: Dict[str, list] = {}
|
||||||
|
self._lock = threading.Lock() # Thread-safe lock for progress dict
|
||||||
|
self._main_loop: Optional[asyncio.AbstractEventLoop] = None
|
||||||
|
|
||||||
|
def _set_main_loop(self, loop: asyncio.AbstractEventLoop):
|
||||||
|
"""Set the main event loop for thread-safe operations."""
|
||||||
|
self._main_loop = loop
|
||||||
|
|
||||||
|
def _notify_listeners_threadsafe(self, model_name: str, progress_data: Dict):
|
||||||
|
"""Notify listeners in a thread-safe manner."""
|
||||||
|
import logging
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
if model_name not in self._listeners:
|
||||||
|
return
|
||||||
|
|
||||||
|
for queue in self._listeners[model_name]:
|
||||||
|
try:
|
||||||
|
# Check if we're in the main event loop thread
|
||||||
|
try:
|
||||||
|
running_loop = asyncio.get_running_loop()
|
||||||
|
# We're in an async context, can use put_nowait directly
|
||||||
|
queue.put_nowait(progress_data.copy())
|
||||||
|
except RuntimeError:
|
||||||
|
# Not in async context (running in background thread)
|
||||||
|
# Use call_soon_threadsafe to safely put on queue
|
||||||
|
if self._main_loop and self._main_loop.is_running():
|
||||||
|
self._main_loop.call_soon_threadsafe(
|
||||||
|
lambda q=queue, d=progress_data.copy(): q.put_nowait(d) if not q.full() else None
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
logger.debug(f"No main loop available for {model_name}, skipping notification")
|
||||||
|
except asyncio.QueueFull:
|
||||||
|
logger.warning(f"Queue full for {model_name}, dropping update")
|
||||||
|
except Exception as e:
|
||||||
|
logger.warning(f"Error notifying listener for {model_name}: {e}")
|
||||||
|
|
||||||
def update_progress(
|
def update_progress(
|
||||||
self,
|
self,
|
||||||
model_name: str,
|
model_name: str,
|
||||||
@@ -26,7 +65,9 @@ class ProgressManager:
|
|||||||
):
|
):
|
||||||
"""
|
"""
|
||||||
Update progress for a model download.
|
Update progress for a model download.
|
||||||
|
|
||||||
|
Thread-safe: can be called from background threads.
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
model_name: Name of the model (e.g., "qwen-tts-1.7B", "whisper-base")
|
model_name: Name of the model (e.g., "qwen-tts-1.7B", "whisper-base")
|
||||||
current: Current bytes downloaded
|
current: Current bytes downloaded
|
||||||
@@ -34,9 +75,12 @@ class ProgressManager:
|
|||||||
filename: Current file being downloaded
|
filename: Current file being downloaded
|
||||||
status: Status string (downloading, extracting, complete, error)
|
status: Status string (downloading, extracting, complete, error)
|
||||||
"""
|
"""
|
||||||
|
import logging
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
progress_pct = (current / total * 100) if total > 0 else 0
|
progress_pct = (current / total * 100) if total > 0 else 0
|
||||||
|
|
||||||
self._progress[model_name] = {
|
progress_data = {
|
||||||
"model_name": model_name,
|
"model_name": model_name,
|
||||||
"current": current,
|
"current": current,
|
||||||
"total": total,
|
"total": total,
|
||||||
@@ -45,26 +89,43 @@ class ProgressManager:
|
|||||||
"status": status,
|
"status": status,
|
||||||
"timestamp": datetime.now().isoformat(),
|
"timestamp": datetime.now().isoformat(),
|
||||||
}
|
}
|
||||||
|
|
||||||
# Notify all listeners
|
print(f"[DEBUG] update_progress called: {model_name}, {progress_pct:.1f}%")
|
||||||
if model_name in self._listeners:
|
|
||||||
for queue in self._listeners[model_name]:
|
# Thread-safe update of progress dict
|
||||||
try:
|
with self._lock:
|
||||||
queue.put_nowait(self._progress[model_name].copy())
|
self._progress[model_name] = progress_data
|
||||||
except asyncio.QueueFull:
|
|
||||||
pass
|
# Notify all listeners (thread-safe)
|
||||||
|
listener_count = len(self._listeners.get(model_name, []))
|
||||||
|
print(f"[DEBUG] Listener count for {model_name}: {listener_count}")
|
||||||
|
print(f"[DEBUG] All listeners: {list(self._listeners.keys())}")
|
||||||
|
print(f"[DEBUG] Main loop set: {self._main_loop is not None}")
|
||||||
|
if self._main_loop:
|
||||||
|
print(f"[DEBUG] Main loop running: {self._main_loop.is_running()}")
|
||||||
|
|
||||||
|
if listener_count > 0:
|
||||||
|
logger.debug(f"Notifying {listener_count} listeners for {model_name}: {progress_pct:.1f}% ({filename})")
|
||||||
|
print(f"[DEBUG] About to notify listeners...")
|
||||||
|
self._notify_listeners_threadsafe(model_name, progress_data)
|
||||||
|
print(f"[DEBUG] Notified listeners")
|
||||||
|
else:
|
||||||
|
logger.debug(f"No listeners for {model_name}, progress update stored: {progress_pct:.1f}%")
|
||||||
|
|
||||||
def get_progress(self, model_name: str) -> Optional[Dict]:
|
def get_progress(self, model_name: str) -> Optional[Dict]:
|
||||||
"""Get current progress for a model."""
|
"""Get current progress for a model. Thread-safe."""
|
||||||
return self._progress.get(model_name)
|
with self._lock:
|
||||||
|
progress = self._progress.get(model_name)
|
||||||
|
return progress.copy() if progress else None
|
||||||
|
|
||||||
def get_all_active(self) -> List[Dict]:
|
def get_all_active(self) -> List[Dict]:
|
||||||
"""Get all active downloads (status is 'downloading' or 'extracting')."""
|
"""Get all active downloads (status is 'downloading' or 'extracting'). Thread-safe."""
|
||||||
active = []
|
active = []
|
||||||
for model_name, progress in self._progress.items():
|
with self._lock:
|
||||||
status = progress.get("status", "")
|
for model_name, progress in self._progress.items():
|
||||||
if status in ("downloading", "extracting"):
|
status = progress.get("status", "")
|
||||||
active.append(progress.copy())
|
if status in ("downloading", "extracting"):
|
||||||
|
active.append(progress.copy())
|
||||||
return active
|
return active
|
||||||
|
|
||||||
def create_progress_callback(self, model_name: str, filename: Optional[str] = None):
|
def create_progress_callback(self, model_name: str, filename: Optional[str] = None):
|
||||||
@@ -98,30 +159,57 @@ class ProgressManager:
|
|||||||
async def subscribe(self, model_name: str):
|
async def subscribe(self, model_name: str):
|
||||||
"""
|
"""
|
||||||
Subscribe to progress updates for a model.
|
Subscribe to progress updates for a model.
|
||||||
|
|
||||||
Yields progress updates as Server-Sent Events.
|
Yields progress updates as Server-Sent Events.
|
||||||
"""
|
"""
|
||||||
queue = asyncio.Queue(maxsize=10)
|
import logging
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# Store the main event loop for thread-safe operations
|
||||||
|
try:
|
||||||
|
self._main_loop = asyncio.get_running_loop()
|
||||||
|
except RuntimeError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
queue = asyncio.Queue(maxsize=10)
|
||||||
|
|
||||||
# Add to listeners
|
# Add to listeners
|
||||||
if model_name not in self._listeners:
|
if model_name not in self._listeners:
|
||||||
self._listeners[model_name] = []
|
self._listeners[model_name] = []
|
||||||
self._listeners[model_name].append(queue)
|
self._listeners[model_name].append(queue)
|
||||||
|
|
||||||
|
logger.info(f"SSE client subscribed to {model_name}, total listeners: {len(self._listeners[model_name])}")
|
||||||
|
|
||||||
try:
|
try:
|
||||||
# Send initial progress if available
|
# Send initial progress if available and still in progress (thread-safe read)
|
||||||
if model_name in self._progress:
|
with self._lock:
|
||||||
yield f"data: {json.dumps(self._progress[model_name])}\n\n"
|
initial_progress = self._progress.get(model_name)
|
||||||
|
if initial_progress:
|
||||||
|
initial_progress = initial_progress.copy()
|
||||||
|
|
||||||
|
if initial_progress:
|
||||||
|
status = initial_progress.get('status')
|
||||||
|
# Only send initial progress if download is actually in progress
|
||||||
|
# Don't send old 'complete' or 'error' status from previous downloads
|
||||||
|
if status in ('downloading', 'extracting'):
|
||||||
|
logger.info(f"Sending initial progress for {model_name}: {status}")
|
||||||
|
yield f"data: {json.dumps(initial_progress)}\n\n"
|
||||||
|
else:
|
||||||
|
logger.info(f"Skipping initial progress for {model_name} (status: {status})")
|
||||||
|
else:
|
||||||
|
logger.info(f"No initial progress available for {model_name}")
|
||||||
|
|
||||||
# Stream updates
|
# Stream updates
|
||||||
while True:
|
while True:
|
||||||
try:
|
try:
|
||||||
# Wait for update with timeout
|
# Wait for update with timeout
|
||||||
progress = await asyncio.wait_for(queue.get(), timeout=1.0)
|
progress = await asyncio.wait_for(queue.get(), timeout=1.0)
|
||||||
|
logger.debug(f"Sending progress update for {model_name}: {progress.get('status')} - {progress.get('progress', 0):.1f}%")
|
||||||
yield f"data: {json.dumps(progress)}\n\n"
|
yield f"data: {json.dumps(progress)}\n\n"
|
||||||
|
|
||||||
# Stop if complete or error
|
# Stop if complete or error
|
||||||
if progress.get("status") in ("complete", "error"):
|
if progress.get("status") in ("complete", "error"):
|
||||||
|
logger.info(f"Download {progress.get('status')} for {model_name}, closing SSE connection")
|
||||||
break
|
break
|
||||||
except asyncio.TimeoutError:
|
except asyncio.TimeoutError:
|
||||||
# Send heartbeat
|
# Send heartbeat
|
||||||
@@ -133,32 +221,53 @@ class ProgressManager:
|
|||||||
self._listeners[model_name].remove(queue)
|
self._listeners[model_name].remove(queue)
|
||||||
if not self._listeners[model_name]:
|
if not self._listeners[model_name]:
|
||||||
del self._listeners[model_name]
|
del self._listeners[model_name]
|
||||||
|
logger.info(f"SSE client unsubscribed from {model_name}, remaining listeners: {len(self._listeners.get(model_name, []))}")
|
||||||
|
|
||||||
def mark_complete(self, model_name: str):
|
def mark_complete(self, model_name: str):
|
||||||
"""Mark a model download as complete."""
|
"""Mark a model download as complete. Thread-safe."""
|
||||||
if model_name in self._progress:
|
import logging
|
||||||
self._progress[model_name]["status"] = "complete"
|
logger = logging.getLogger(__name__)
|
||||||
self._progress[model_name]["progress"] = 100.0
|
|
||||||
# Notify listeners
|
with self._lock:
|
||||||
if model_name in self._listeners:
|
if model_name in self._progress:
|
||||||
for queue in self._listeners[model_name]:
|
self._progress[model_name]["status"] = "complete"
|
||||||
try:
|
self._progress[model_name]["progress"] = 100.0
|
||||||
queue.put_nowait(self._progress[model_name].copy())
|
progress_data = self._progress[model_name].copy()
|
||||||
except asyncio.QueueFull:
|
else:
|
||||||
pass
|
logger.warning(f"Cannot mark {model_name} as complete: not found in progress")
|
||||||
|
return
|
||||||
|
|
||||||
|
logger.info(f"Marked {model_name} as complete")
|
||||||
|
# Notify listeners (thread-safe)
|
||||||
|
self._notify_listeners_threadsafe(model_name, progress_data)
|
||||||
|
|
||||||
def mark_error(self, model_name: str, error: str):
|
def mark_error(self, model_name: str, error: str):
|
||||||
"""Mark a model download as failed."""
|
"""Mark a model download as failed. Thread-safe."""
|
||||||
if model_name in self._progress:
|
import logging
|
||||||
self._progress[model_name]["status"] = "error"
|
logger = logging.getLogger(__name__)
|
||||||
self._progress[model_name]["error"] = error
|
|
||||||
# Notify listeners
|
with self._lock:
|
||||||
if model_name in self._listeners:
|
if model_name in self._progress:
|
||||||
for queue in self._listeners[model_name]:
|
self._progress[model_name]["status"] = "error"
|
||||||
try:
|
self._progress[model_name]["error"] = error
|
||||||
queue.put_nowait(self._progress[model_name].copy())
|
progress_data = self._progress[model_name].copy()
|
||||||
except asyncio.QueueFull:
|
else:
|
||||||
pass
|
# Create new progress entry for error
|
||||||
|
progress_data = {
|
||||||
|
"model_name": model_name,
|
||||||
|
"current": 0,
|
||||||
|
"total": 0,
|
||||||
|
"progress": 0,
|
||||||
|
"filename": None,
|
||||||
|
"status": "error",
|
||||||
|
"error": error,
|
||||||
|
"timestamp": datetime.now().isoformat(),
|
||||||
|
}
|
||||||
|
self._progress[model_name] = progress_data
|
||||||
|
|
||||||
|
logger.error(f"Marked {model_name} as error: {error}")
|
||||||
|
# Notify listeners (thread-safe)
|
||||||
|
self._notify_listeners_threadsafe(model_name, progress_data)
|
||||||
|
|
||||||
|
|
||||||
# Global progress manager instance
|
# Global progress manager instance
|
||||||
|
|||||||
@@ -4,16 +4,20 @@ from PyInstaller.utils.hooks import collect_submodules
|
|||||||
from PyInstaller.utils.hooks import copy_metadata
|
from PyInstaller.utils.hooks import copy_metadata
|
||||||
|
|
||||||
datas = []
|
datas = []
|
||||||
hiddenimports = ['backend', 'backend.main', 'backend.config', 'backend.database', 'backend.models', 'backend.profiles', 'backend.history', 'backend.tts', 'backend.transcribe', 'backend.utils.audio', 'backend.utils.cache', 'backend.utils.progress', 'backend.utils.hf_progress', 'backend.utils.validation', 'torch', 'transformers', 'fastapi', 'uvicorn', 'sqlalchemy', 'librosa', 'soundfile', 'qwen_tts', 'qwen_tts.inference', 'qwen_tts.inference.qwen3_tts_model', 'qwen_tts.inference.qwen3_tts_tokenizer', 'qwen_tts.core', 'qwen_tts.cli', 'pkg_resources.extern']
|
hiddenimports = ['backend', 'backend.main', 'backend.config', 'backend.database', 'backend.models', 'backend.profiles', 'backend.history', 'backend.tts', 'backend.transcribe', 'backend.platform_detect', 'backend.backends', 'backend.backends.pytorch_backend', 'backend.utils.audio', 'backend.utils.cache', 'backend.utils.progress', 'backend.utils.hf_progress', 'backend.utils.validation', 'torch', 'transformers', 'fastapi', 'uvicorn', 'sqlalchemy', 'librosa', 'soundfile', 'qwen_tts', 'qwen_tts.inference', 'qwen_tts.inference.qwen3_tts_model', 'qwen_tts.inference.qwen3_tts_tokenizer', 'qwen_tts.core', 'qwen_tts.cli', 'pkg_resources.extern', 'backend.backends.mlx_backend', 'mlx', 'mlx.core', 'mlx.nn', 'mlx_audio', 'mlx_audio.tts', 'mlx_audio.stt']
|
||||||
datas += collect_data_files('qwen_tts')
|
datas += collect_data_files('qwen_tts')
|
||||||
|
datas += collect_data_files('mlx')
|
||||||
|
datas += collect_data_files('mlx_audio')
|
||||||
datas += copy_metadata('qwen-tts')
|
datas += copy_metadata('qwen-tts')
|
||||||
hiddenimports += collect_submodules('qwen_tts')
|
hiddenimports += collect_submodules('qwen_tts')
|
||||||
hiddenimports += collect_submodules('jaraco')
|
hiddenimports += collect_submodules('jaraco')
|
||||||
|
hiddenimports += collect_submodules('mlx')
|
||||||
|
hiddenimports += collect_submodules('mlx_audio')
|
||||||
|
|
||||||
|
|
||||||
a = Analysis(
|
a = Analysis(
|
||||||
['server.py'],
|
['server.py'],
|
||||||
pathex=['C:\\Users\\ijame\\Projects\\voice\\Qwen3-TTS'],
|
pathex=[],
|
||||||
binaries=[],
|
binaries=[],
|
||||||
datas=datas,
|
datas=datas,
|
||||||
hiddenimports=hiddenimports,
|
hiddenimports=hiddenimports,
|
||||||
|
|||||||
@@ -13,8 +13,11 @@
|
|||||||
},
|
},
|
||||||
"app": {
|
"app": {
|
||||||
"name": "@voicebox/app",
|
"name": "@voicebox/app",
|
||||||
"version": "0.1.0",
|
"version": "0.1.9",
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
|
"@dnd-kit/core": "^6.3.1",
|
||||||
|
"@dnd-kit/sortable": "^10.0.0",
|
||||||
|
"@dnd-kit/utilities": "^3.2.2",
|
||||||
"@hookform/resolvers": "^3.9.0",
|
"@hookform/resolvers": "^3.9.0",
|
||||||
"@radix-ui/react-alert-dialog": "^1.1.1",
|
"@radix-ui/react-alert-dialog": "^1.1.1",
|
||||||
"@radix-ui/react-avatar": "^1.1.0",
|
"@radix-ui/react-avatar": "^1.1.0",
|
||||||
@@ -32,6 +35,7 @@
|
|||||||
"@radix-ui/react-toast": "^1.2.1",
|
"@radix-ui/react-toast": "^1.2.1",
|
||||||
"@tanstack/react-query": "^5.0.0",
|
"@tanstack/react-query": "^5.0.0",
|
||||||
"@tanstack/react-query-devtools": "^5.0.0",
|
"@tanstack/react-query-devtools": "^5.0.0",
|
||||||
|
"@tanstack/react-router": "^1.157.16",
|
||||||
"@tauri-apps/api": "^2.0.0",
|
"@tauri-apps/api": "^2.0.0",
|
||||||
"@tauri-apps/plugin-dialog": "^2.0.0",
|
"@tauri-apps/plugin-dialog": "^2.0.0",
|
||||||
"@tauri-apps/plugin-fs": "^2.0.0",
|
"@tauri-apps/plugin-fs": "^2.0.0",
|
||||||
@@ -46,6 +50,7 @@
|
|||||||
"react": "^18.3.0",
|
"react": "^18.3.0",
|
||||||
"react-dom": "^18.3.0",
|
"react-dom": "^18.3.0",
|
||||||
"react-hook-form": "^7.53.0",
|
"react-hook-form": "^7.53.0",
|
||||||
|
"react-sound-visualizer": "^1.4.0",
|
||||||
"tailwind-merge": "^2.5.4",
|
"tailwind-merge": "^2.5.4",
|
||||||
"wavesurfer.js": "^7.0.0",
|
"wavesurfer.js": "^7.0.0",
|
||||||
"zod": "^3.23.8",
|
"zod": "^3.23.8",
|
||||||
@@ -63,7 +68,7 @@
|
|||||||
},
|
},
|
||||||
"landing": {
|
"landing": {
|
||||||
"name": "@voicebox/landing",
|
"name": "@voicebox/landing",
|
||||||
"version": "0.1.0",
|
"version": "0.1.9",
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@radix-ui/react-separator": "^1.1.8",
|
"@radix-ui/react-separator": "^1.1.8",
|
||||||
"@radix-ui/react-slot": "^1.2.4",
|
"@radix-ui/react-slot": "^1.2.4",
|
||||||
@@ -88,7 +93,7 @@
|
|||||||
},
|
},
|
||||||
"tauri": {
|
"tauri": {
|
||||||
"name": "@voicebox/tauri",
|
"name": "@voicebox/tauri",
|
||||||
"version": "0.1.0",
|
"version": "0.1.9",
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@tauri-apps/api": "^2.0.0",
|
"@tauri-apps/api": "^2.0.0",
|
||||||
"@tauri-apps/plugin-shell": "^2.0.0",
|
"@tauri-apps/plugin-shell": "^2.0.0",
|
||||||
@@ -107,7 +112,7 @@
|
|||||||
},
|
},
|
||||||
"web": {
|
"web": {
|
||||||
"name": "@voicebox/web",
|
"name": "@voicebox/web",
|
||||||
"version": "0.1.0",
|
"version": "0.1.9",
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@tanstack/react-query": "^5.0.0",
|
"@tanstack/react-query": "^5.0.0",
|
||||||
"react": "^18.3.0",
|
"react": "^18.3.0",
|
||||||
@@ -188,6 +193,14 @@
|
|||||||
|
|
||||||
"@biomejs/cli-win32-x64": ["@biomejs/[email protected]", "", { "os": "win32", "cpu": "x64" }, "sha512-qqGVWqNNek0KikwPZlOIoxtXgsNGsX+rgdEzgw82Re8nF02W+E2WokaQhpF5TdBh/D/RQ3TLppH+otp6ztN0lw=="],
|
"@biomejs/cli-win32-x64": ["@biomejs/[email protected]", "", { "os": "win32", "cpu": "x64" }, "sha512-qqGVWqNNek0KikwPZlOIoxtXgsNGsX+rgdEzgw82Re8nF02W+E2WokaQhpF5TdBh/D/RQ3TLppH+otp6ztN0lw=="],
|
||||||
|
|
||||||
|
"@dnd-kit/accessibility": ["@dnd-kit/[email protected]", "", { "dependencies": { "tslib": "^2.0.0" }, "peerDependencies": { "react": ">=16.8.0" } }, "sha512-2P+YgaXF+gRsIihwwY1gCsQSYnu9Zyj2py8kY5fFvUM1qm2WA2u639R6YNVfU4GWr+ZM5mqEsfHZZLoRONbemw=="],
|
||||||
|
|
||||||
|
"@dnd-kit/core": ["@dnd-kit/[email protected]", "", { "dependencies": { "@dnd-kit/accessibility": "^3.1.1", "@dnd-kit/utilities": "^3.2.2", "tslib": "^2.0.0" }, "peerDependencies": { "react": ">=16.8.0", "react-dom": ">=16.8.0" } }, "sha512-xkGBRQQab4RLwgXxoqETICr6S5JlogafbhNsidmrkVv2YRs5MLwpjoF2qpiGjQt8S9AoxtIV603s0GIUpY5eYQ=="],
|
||||||
|
|
||||||
|
"@dnd-kit/sortable": ["@dnd-kit/[email protected]", "", { "dependencies": { "@dnd-kit/utilities": "^3.2.2", "tslib": "^2.0.0" }, "peerDependencies": { "@dnd-kit/core": "^6.3.0", "react": ">=16.8.0" } }, "sha512-+xqhmIIzvAYMGfBYYnbKuNicfSsk4RksY2XdmJhT+HAC01nix6fHCztU68jooFiMUB01Ky3F0FyOvhG/BZrWkg=="],
|
||||||
|
|
||||||
|
"@dnd-kit/utilities": ["@dnd-kit/[email protected]", "", { "dependencies": { "tslib": "^2.0.0" }, "peerDependencies": { "react": ">=16.8.0" } }, "sha512-+MKAJEOfaBe5SmV6t34p80MMKhjvUz0vRrvVJbPT0WElzaOJ/1xs+D+KDv+tD/NE5ujfrChEcshd4fLn0wpiqg=="],
|
||||||
|
|
||||||
"@emnapi/runtime": ["@emnapi/[email protected]", "", { "dependencies": { "tslib": "^2.4.0" } }, "sha512-mehfKSMWjjNol8659Z8KxEMrdSJDDot5SXMq00dM8BN4o+CLNXQ0xH2V7EchNHV4RmbZLmmPdEaXZc5H2FXmDg=="],
|
"@emnapi/runtime": ["@emnapi/[email protected]", "", { "dependencies": { "tslib": "^2.4.0" } }, "sha512-mehfKSMWjjNol8659Z8KxEMrdSJDDot5SXMq00dM8BN4o+CLNXQ0xH2V7EchNHV4RmbZLmmPdEaXZc5H2FXmDg=="],
|
||||||
|
|
||||||
"@esbuild/aix-ppc64": ["@esbuild/[email protected]", "", { "os": "aix", "cpu": "ppc64" }, "sha512-1SDgH6ZSPTlggy1yI6+Dbkiz8xzpHJEVAlF/AM1tHPLsf5STom9rwtjE4hKAF20FfXXNTFqEYXyJNWh1GiZedQ=="],
|
"@esbuild/aix-ppc64": ["@esbuild/[email protected]", "", { "os": "aix", "cpu": "ppc64" }, "sha512-1SDgH6ZSPTlggy1yI6+Dbkiz8xzpHJEVAlF/AM1tHPLsf5STom9rwtjE4hKAF20FfXXNTFqEYXyJNWh1GiZedQ=="],
|
||||||
@@ -512,6 +525,8 @@
|
|||||||
|
|
||||||
"@tailwindcss/vite": ["@tailwindcss/[email protected]", "", { "dependencies": { "@tailwindcss/node": "4.1.18", "@tailwindcss/oxide": "4.1.18", "tailwindcss": "4.1.18" }, "peerDependencies": { "vite": "^5.2.0 || ^6 || ^7" } }, "sha512-jVA+/UpKL1vRLg6Hkao5jldawNmRo7mQYrZtNHMIVpLfLhDml5nMRUo/8MwoX2vNXvnaXNNMedrMfMugAVX1nA=="],
|
"@tailwindcss/vite": ["@tailwindcss/[email protected]", "", { "dependencies": { "@tailwindcss/node": "4.1.18", "@tailwindcss/oxide": "4.1.18", "tailwindcss": "4.1.18" }, "peerDependencies": { "vite": "^5.2.0 || ^6 || ^7" } }, "sha512-jVA+/UpKL1vRLg6Hkao5jldawNmRo7mQYrZtNHMIVpLfLhDml5nMRUo/8MwoX2vNXvnaXNNMedrMfMugAVX1nA=="],
|
||||||
|
|
||||||
|
"@tanstack/history": ["@tanstack/[email protected]", "", {}, "sha512-xyIfof8eHBuub1CkBnbKNKQXeRZC4dClhmzePHVOEel4G7lk/dW+TQ16da7CFdeNLv6u6Owf5VoBQxoo6DFTSA=="],
|
||||||
|
|
||||||
"@tanstack/query-core": ["@tanstack/[email protected]", "", {}, "sha512-OMD2HLpNouXEfZJWcKeVKUgQ5n+n3A2JFmBaScpNDUqSrQSjiveC7dKMe53uJUg1nDG16ttFPz2xfilz6i2uVg=="],
|
"@tanstack/query-core": ["@tanstack/[email protected]", "", {}, "sha512-OMD2HLpNouXEfZJWcKeVKUgQ5n+n3A2JFmBaScpNDUqSrQSjiveC7dKMe53uJUg1nDG16ttFPz2xfilz6i2uVg=="],
|
||||||
|
|
||||||
"@tanstack/query-devtools": ["@tanstack/[email protected]", "", {}, "sha512-N8D27KH1vEpVacvZgJL27xC6yPFUy0Zkezn5gnB3L3gRCxlDeSuiya7fKge8Y91uMTnC8aSxBQhcK6ocY7alpQ=="],
|
"@tanstack/query-devtools": ["@tanstack/[email protected]", "", {}, "sha512-N8D27KH1vEpVacvZgJL27xC6yPFUy0Zkezn5gnB3L3gRCxlDeSuiya7fKge8Y91uMTnC8aSxBQhcK6ocY7alpQ=="],
|
||||||
@@ -520,6 +535,14 @@
|
|||||||
|
|
||||||
"@tanstack/react-query-devtools": ["@tanstack/[email protected]", "", { "dependencies": { "@tanstack/query-devtools": "5.92.0" }, "peerDependencies": { "@tanstack/react-query": "^5.90.14", "react": "^18 || ^19" } }, "sha512-ZJ1503ay5fFeEYFUdo7LMNFzZryi6B0Cacrgr2h1JRkvikK1khgIq6Nq2EcblqEdIlgB/r7XDW8f8DQ89RuUgg=="],
|
"@tanstack/react-query-devtools": ["@tanstack/[email protected]", "", { "dependencies": { "@tanstack/query-devtools": "5.92.0" }, "peerDependencies": { "@tanstack/react-query": "^5.90.14", "react": "^18 || ^19" } }, "sha512-ZJ1503ay5fFeEYFUdo7LMNFzZryi6B0Cacrgr2h1JRkvikK1khgIq6Nq2EcblqEdIlgB/r7XDW8f8DQ89RuUgg=="],
|
||||||
|
|
||||||
|
"@tanstack/react-router": ["@tanstack/[email protected]", "", { "dependencies": { "@tanstack/history": "1.154.14", "@tanstack/react-store": "^0.8.0", "@tanstack/router-core": "1.157.16", "isbot": "^5.1.22", "tiny-invariant": "^1.3.3", "tiny-warning": "^1.0.3" }, "peerDependencies": { "react": ">=18.0.0 || >=19.0.0", "react-dom": ">=18.0.0 || >=19.0.0" } }, "sha512-xwFQa7S7dhBhm3aJYwU79cITEYgAKSrcL6wokaROIvl2JyIeazn8jueWqUPJzFjv+QF6Q8euKRlKUEyb5q2ymg=="],
|
||||||
|
|
||||||
|
"@tanstack/react-store": ["@tanstack/[email protected]", "", { "dependencies": { "@tanstack/store": "0.8.0", "use-sync-external-store": "^1.6.0" }, "peerDependencies": { "react": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0", "react-dom": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0" } }, "sha512-1vG9beLIuB7q69skxK9r5xiLN3ztzIPfSQSs0GfeqWGO2tGIyInZx0x1COhpx97RKaONSoAb8C3dxacWksm1ow=="],
|
||||||
|
|
||||||
|
"@tanstack/router-core": ["@tanstack/[email protected]", "", { "dependencies": { "@tanstack/history": "1.154.14", "@tanstack/store": "^0.8.0", "cookie-es": "^2.0.0", "seroval": "^1.4.2", "seroval-plugins": "^1.4.2", "tiny-invariant": "^1.3.3", "tiny-warning": "^1.0.3" } }, "sha512-eJuVgM7KZYTTr4uPorbUzUflmljMVcaX2g6VvhITLnHmg9SBx9RAgtQ1HmT+72mzyIbRSlQ1q0fY/m+of/fosA=="],
|
||||||
|
|
||||||
|
"@tanstack/store": ["@tanstack/[email protected]", "", {}, "sha512-Om+BO0YfMZe//X2z0uLF2j+75nQga6TpTJgLJQBiq85aOyZNIhkCgleNcud2KQg4k4v9Y9l+Uhru3qWMPGTOzQ=="],
|
||||||
|
|
||||||
"@tauri-apps/api": ["@tauri-apps/[email protected]", "", {}, "sha512-IGlhP6EivjXHepbBic618GOmiWe4URJiIeZFlB7x3czM0yDHHYviH1Xvoiv4FefdkQtn6v7TuwWCRfOGdnVUGw=="],
|
"@tauri-apps/api": ["@tauri-apps/[email protected]", "", {}, "sha512-IGlhP6EivjXHepbBic618GOmiWe4URJiIeZFlB7x3czM0yDHHYviH1Xvoiv4FefdkQtn6v7TuwWCRfOGdnVUGw=="],
|
||||||
|
|
||||||
"@tauri-apps/cli": ["@tauri-apps/[email protected]", "", { "optionalDependencies": { "@tauri-apps/cli-darwin-arm64": "2.9.6", "@tauri-apps/cli-darwin-x64": "2.9.6", "@tauri-apps/cli-linux-arm-gnueabihf": "2.9.6", "@tauri-apps/cli-linux-arm64-gnu": "2.9.6", "@tauri-apps/cli-linux-arm64-musl": "2.9.6", "@tauri-apps/cli-linux-riscv64-gnu": "2.9.6", "@tauri-apps/cli-linux-x64-gnu": "2.9.6", "@tauri-apps/cli-linux-x64-musl": "2.9.6", "@tauri-apps/cli-win32-arm64-msvc": "2.9.6", "@tauri-apps/cli-win32-ia32-msvc": "2.9.6", "@tauri-apps/cli-win32-x64-msvc": "2.9.6" }, "bin": { "tauri": "tauri.js" } }, "sha512-3xDdXL5omQ3sPfBfdC8fCtDKcnyV7OqyzQgfyT5P3+zY6lcPqIYKQBvUasNvppi21RSdfhy44ttvJmftb0PCDw=="],
|
"@tauri-apps/cli": ["@tauri-apps/[email protected]", "", { "optionalDependencies": { "@tauri-apps/cli-darwin-arm64": "2.9.6", "@tauri-apps/cli-darwin-x64": "2.9.6", "@tauri-apps/cli-linux-arm-gnueabihf": "2.9.6", "@tauri-apps/cli-linux-arm64-gnu": "2.9.6", "@tauri-apps/cli-linux-arm64-musl": "2.9.6", "@tauri-apps/cli-linux-riscv64-gnu": "2.9.6", "@tauri-apps/cli-linux-x64-gnu": "2.9.6", "@tauri-apps/cli-linux-x64-musl": "2.9.6", "@tauri-apps/cli-win32-arm64-msvc": "2.9.6", "@tauri-apps/cli-win32-ia32-msvc": "2.9.6", "@tauri-apps/cli-win32-x64-msvc": "2.9.6" }, "bin": { "tauri": "tauri.js" } }, "sha512-3xDdXL5omQ3sPfBfdC8fCtDKcnyV7OqyzQgfyT5P3+zY6lcPqIYKQBvUasNvppi21RSdfhy44ttvJmftb0PCDw=="],
|
||||||
@@ -664,6 +687,8 @@
|
|||||||
|
|
||||||
"convert-source-map": ["[email protected]", "", {}, "sha512-Kvp459HrV2FEJ1CAsi1Ku+MY3kasH19TFykTz2xWmMeq6bk2NU3XXvfJ+Q61m0xktWwt+1HSYf3JZsTms3aRJg=="],
|
"convert-source-map": ["[email protected]", "", {}, "sha512-Kvp459HrV2FEJ1CAsi1Ku+MY3kasH19TFykTz2xWmMeq6bk2NU3XXvfJ+Q61m0xktWwt+1HSYf3JZsTms3aRJg=="],
|
||||||
|
|
||||||
|
"cookie-es": ["[email protected]", "", {}, "sha512-RAj4E421UYRgqokKUmotqAwuplYw15qtdXfY+hGzgCJ/MBjCVZcSoHK/kH9kocfjRjcDME7IiDWR/1WX1TM2Pg=="],
|
||||||
|
|
||||||
"cross-spawn": ["[email protected]", "", { "dependencies": { "path-key": "^3.1.0", "shebang-command": "^2.0.0", "which": "^2.0.1" } }, "sha512-uV2QOWP2nWzsy2aMp8aRibhi9dlzF5Hgh5SHaB9OiTGEyDTiJJyx0uy51QXdyWbtAHNua4XJzUKca3OzKUd3vA=="],
|
"cross-spawn": ["[email protected]", "", { "dependencies": { "path-key": "^3.1.0", "shebang-command": "^2.0.0", "which": "^2.0.1" } }, "sha512-uV2QOWP2nWzsy2aMp8aRibhi9dlzF5Hgh5SHaB9OiTGEyDTiJJyx0uy51QXdyWbtAHNua4XJzUKca3OzKUd3vA=="],
|
||||||
|
|
||||||
"cssesc": ["[email protected]", "", { "bin": { "cssesc": "bin/cssesc" } }, "sha512-/Tb/JcjK111nNScGob5MNtsntNM1aCNUDipB/TkwZFhyDrrE47SOx/18wF2bbjgc3ZzCSKW1T5nt5EbFoAz/Vg=="],
|
"cssesc": ["[email protected]", "", { "bin": { "cssesc": "bin/cssesc" } }, "sha512-/Tb/JcjK111nNScGob5MNtsntNM1aCNUDipB/TkwZFhyDrrE47SOx/18wF2bbjgc3ZzCSKW1T5nt5EbFoAz/Vg=="],
|
||||||
@@ -792,6 +817,8 @@
|
|||||||
|
|
||||||
"is-path-inside": ["[email protected]", "", {}, "sha512-Fd4gABb+ycGAmKou8eMftCupSir5lRxqf4aD/vd0cD2qc4HL07OjCeuHMr8Ro4CoMaeCKDB0/ECBOVWjTwUvPQ=="],
|
"is-path-inside": ["[email protected]", "", {}, "sha512-Fd4gABb+ycGAmKou8eMftCupSir5lRxqf4aD/vd0cD2qc4HL07OjCeuHMr8Ro4CoMaeCKDB0/ECBOVWjTwUvPQ=="],
|
||||||
|
|
||||||
|
"isbot": ["[email protected]", "", {}, "sha512-aCMIBSKd/XPRYdiCQTLC8QHH4YT8B3JUADu+7COgYIZPvkeoMcUHMRjZLM9/7V8fCj+l7FSREc1lOPNjzogo/A=="],
|
||||||
|
|
||||||
"isexe": ["[email protected]", "", {}, "sha512-RHxMLp9lnKHGHRng9QFhRCMbYAcVpn69smSGcq3f36xjgVVWThj4qqLbTLlq7Ssj8B+fIQ1EuCEGI2lKsyQeIw=="],
|
"isexe": ["[email protected]", "", {}, "sha512-RHxMLp9lnKHGHRng9QFhRCMbYAcVpn69smSGcq3f36xjgVVWThj4qqLbTLlq7Ssj8B+fIQ1EuCEGI2lKsyQeIw=="],
|
||||||
|
|
||||||
"jiti": ["[email protected]", "", { "bin": { "jiti": "bin/jiti.js" } }, "sha512-/imKNG4EbWNrVjoNC/1H5/9GFy+tqjGBHCaSsN+P2RnPqjsLmv6UD3Ej+Kj8nBWaRAwyk7kK5ZUc+OEatnTR3A=="],
|
"jiti": ["[email protected]", "", { "bin": { "jiti": "bin/jiti.js" } }, "sha512-/imKNG4EbWNrVjoNC/1H5/9GFy+tqjGBHCaSsN+P2RnPqjsLmv6UD3Ej+Kj8nBWaRAwyk7kK5ZUc+OEatnTR3A=="],
|
||||||
@@ -944,6 +971,8 @@
|
|||||||
|
|
||||||
"react-remove-scroll-bar": ["[email protected]", "", { "dependencies": { "react-style-singleton": "^2.2.2", "tslib": "^2.0.0" }, "peerDependencies": { "@types/react": "*", "react": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0" }, "optionalPeers": ["@types/react"] }, "sha512-9r+yi9+mgU33AKcj6IbT9oRCO78WriSj6t/cF8DWBZJ9aOGPOTEDvdUDz1FwKim7QXWwmHqtdHnRJfhAxEG46Q=="],
|
"react-remove-scroll-bar": ["[email protected]", "", { "dependencies": { "react-style-singleton": "^2.2.2", "tslib": "^2.0.0" }, "peerDependencies": { "@types/react": "*", "react": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0" }, "optionalPeers": ["@types/react"] }, "sha512-9r+yi9+mgU33AKcj6IbT9oRCO78WriSj6t/cF8DWBZJ9aOGPOTEDvdUDz1FwKim7QXWwmHqtdHnRJfhAxEG46Q=="],
|
||||||
|
|
||||||
|
"react-sound-visualizer": ["[email protected]", "", { "dependencies": { "sound-visualizer": "^1.2.0" }, "peerDependencies": { "react": ">= 16" } }, "sha512-Qe7tFTd1owtQ8nYrUYXg7QLt8mw7iUy86mqj/+IwmXzSw+NlhnMnAGPuisb1Lk3ncliFnM+AQbZb3C4RQN9uMQ=="],
|
||||||
|
|
||||||
"react-style-singleton": ["[email protected]", "", { "dependencies": { "get-nonce": "^1.0.0", "tslib": "^2.0.0" }, "peerDependencies": { "@types/react": "*", "react": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0 || ^19.0.0-rc" }, "optionalPeers": ["@types/react"] }, "sha512-b6jSvxvVnyptAiLjbkWLE/lOnR4lfTtDAl+eUC7RZy+QQWc6wRzIV2CE6xBuMmDxc2qIihtDCZD5NPOFl7fRBQ=="],
|
"react-style-singleton": ["[email protected]", "", { "dependencies": { "get-nonce": "^1.0.0", "tslib": "^2.0.0" }, "peerDependencies": { "@types/react": "*", "react": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0 || ^19.0.0-rc" }, "optionalPeers": ["@types/react"] }, "sha512-b6jSvxvVnyptAiLjbkWLE/lOnR4lfTtDAl+eUC7RZy+QQWc6wRzIV2CE6xBuMmDxc2qIihtDCZD5NPOFl7fRBQ=="],
|
||||||
|
|
||||||
"read-cache": ["[email protected]", "", { "dependencies": { "pify": "^2.3.0" } }, "sha512-Owdv/Ft7IjOgm/i0xvNDZ1LrRANRfew4b2prF3OWMQLxLfu3bS8FVhCsrSCMK4lR56Y9ya+AThoTpDCTxCmpRA=="],
|
"read-cache": ["[email protected]", "", { "dependencies": { "pify": "^2.3.0" } }, "sha512-Owdv/Ft7IjOgm/i0xvNDZ1LrRANRfew4b2prF3OWMQLxLfu3bS8FVhCsrSCMK4lR56Y9ya+AThoTpDCTxCmpRA=="],
|
||||||
@@ -966,6 +995,10 @@
|
|||||||
|
|
||||||
"semver": ["[email protected]", "", { "bin": { "semver": "bin/semver.js" } }, "sha512-BR7VvDCVHO+q2xBEWskxS6DJE1qRnb7DxzUrogb71CWoSficBxYsiAGd+Kl0mmq/MprG9yArRkyrQxTO6XjMzA=="],
|
"semver": ["[email protected]", "", { "bin": { "semver": "bin/semver.js" } }, "sha512-BR7VvDCVHO+q2xBEWskxS6DJE1qRnb7DxzUrogb71CWoSficBxYsiAGd+Kl0mmq/MprG9yArRkyrQxTO6XjMzA=="],
|
||||||
|
|
||||||
|
"seroval": ["[email protected]", "", {}, "sha512-OE4cvmJ1uSPrKorFIH9/w/Qwuvi/IMcGbv5RKgcJ/zjA/IohDLU6SVaxFN9FwajbP7nsX0dQqMDes1whk3y+yw=="],
|
||||||
|
|
||||||
|
"seroval-plugins": ["[email protected]", "", { "peerDependencies": { "seroval": "^1.0" } }, "sha512-EAHqADIQondwRZIdeW2I636zgsODzoBDwb3PT/+7TLDWyw1Dy/Xv7iGUIEXXav7usHDE9HVhOU61irI3EnyyHA=="],
|
||||||
|
|
||||||
"sharp": ["[email protected]", "", { "dependencies": { "@img/colour": "^1.0.0", "detect-libc": "^2.1.2", "semver": "^7.7.3" }, "optionalDependencies": { "@img/sharp-darwin-arm64": "0.34.5", "@img/sharp-darwin-x64": "0.34.5", "@img/sharp-libvips-darwin-arm64": "1.2.4", "@img/sharp-libvips-darwin-x64": "1.2.4", "@img/sharp-libvips-linux-arm": "1.2.4", "@img/sharp-libvips-linux-arm64": "1.2.4", "@img/sharp-libvips-linux-ppc64": "1.2.4", "@img/sharp-libvips-linux-riscv64": "1.2.4", "@img/sharp-libvips-linux-s390x": "1.2.4", "@img/sharp-libvips-linux-x64": "1.2.4", "@img/sharp-libvips-linuxmusl-arm64": "1.2.4", "@img/sharp-libvips-linuxmusl-x64": "1.2.4", "@img/sharp-linux-arm": "0.34.5", "@img/sharp-linux-arm64": "0.34.5", "@img/sharp-linux-ppc64": "0.34.5", "@img/sharp-linux-riscv64": "0.34.5", "@img/sharp-linux-s390x": "0.34.5", "@img/sharp-linux-x64": "0.34.5", "@img/sharp-linuxmusl-arm64": "0.34.5", "@img/sharp-linuxmusl-x64": "0.34.5", "@img/sharp-wasm32": "0.34.5", "@img/sharp-win32-arm64": "0.34.5", "@img/sharp-win32-ia32": "0.34.5", "@img/sharp-win32-x64": "0.34.5" } }, "sha512-Ou9I5Ft9WNcCbXrU9cMgPBcCK8LiwLqcbywW3t4oDV37n1pzpuNLsYiAV8eODnjbtQlSDwZ2cUEeQz4E54Hltg=="],
|
"sharp": ["[email protected]", "", { "dependencies": { "@img/colour": "^1.0.0", "detect-libc": "^2.1.2", "semver": "^7.7.3" }, "optionalDependencies": { "@img/sharp-darwin-arm64": "0.34.5", "@img/sharp-darwin-x64": "0.34.5", "@img/sharp-libvips-darwin-arm64": "1.2.4", "@img/sharp-libvips-darwin-x64": "1.2.4", "@img/sharp-libvips-linux-arm": "1.2.4", "@img/sharp-libvips-linux-arm64": "1.2.4", "@img/sharp-libvips-linux-ppc64": "1.2.4", "@img/sharp-libvips-linux-riscv64": "1.2.4", "@img/sharp-libvips-linux-s390x": "1.2.4", "@img/sharp-libvips-linux-x64": "1.2.4", "@img/sharp-libvips-linuxmusl-arm64": "1.2.4", "@img/sharp-libvips-linuxmusl-x64": "1.2.4", "@img/sharp-linux-arm": "0.34.5", "@img/sharp-linux-arm64": "0.34.5", "@img/sharp-linux-ppc64": "0.34.5", "@img/sharp-linux-riscv64": "0.34.5", "@img/sharp-linux-s390x": "0.34.5", "@img/sharp-linux-x64": "0.34.5", "@img/sharp-linuxmusl-arm64": "0.34.5", "@img/sharp-linuxmusl-x64": "0.34.5", "@img/sharp-wasm32": "0.34.5", "@img/sharp-win32-arm64": "0.34.5", "@img/sharp-win32-ia32": "0.34.5", "@img/sharp-win32-x64": "0.34.5" } }, "sha512-Ou9I5Ft9WNcCbXrU9cMgPBcCK8LiwLqcbywW3t4oDV37n1pzpuNLsYiAV8eODnjbtQlSDwZ2cUEeQz4E54Hltg=="],
|
||||||
|
|
||||||
"shebang-command": ["[email protected]", "", { "dependencies": { "shebang-regex": "^3.0.0" } }, "sha512-kHxr2zZpYtdmrN1qDjrrX/Z1rR1kG8Dx+gkpK1G4eXmvXswmcE1hTWBWYUzlraYw1/yZp6YuDY77YtvbN0dmDA=="],
|
"shebang-command": ["[email protected]", "", { "dependencies": { "shebang-regex": "^3.0.0" } }, "sha512-kHxr2zZpYtdmrN1qDjrrX/Z1rR1kG8Dx+gkpK1G4eXmvXswmcE1hTWBWYUzlraYw1/yZp6YuDY77YtvbN0dmDA=="],
|
||||||
@@ -974,6 +1007,8 @@
|
|||||||
|
|
||||||
"slash": ["[email protected]", "", {}, "sha512-g9Q1haeby36OSStwb4ntCGGGaKsaVSjQ68fBxoQcutl5fS1vuY18H3wSt3jFyFtrkx+Kz0V1G85A4MyAdDMi2Q=="],
|
"slash": ["[email protected]", "", {}, "sha512-g9Q1haeby36OSStwb4ntCGGGaKsaVSjQ68fBxoQcutl5fS1vuY18H3wSt3jFyFtrkx+Kz0V1G85A4MyAdDMi2Q=="],
|
||||||
|
|
||||||
|
"sound-visualizer": ["[email protected]", "", {}, "sha512-2+Un0PrrBgXylnCjrVYUoRW7KEDH29h7O8/MGzeDOgFGBPb9oX/2n/RGBxJXvVv2U3KFwX5olUWeJKf0Rr5TLQ=="],
|
||||||
|
|
||||||
"source-map-js": ["[email protected]", "", {}, "sha512-UXWMKhLOwVKb728IUtQPXxfYU+usdybtUrK/8uGE8CQMvrhOpwvzDBwj0QhSL7MQc7vIsISBG8VQ8+IDQxpfQA=="],
|
"source-map-js": ["[email protected]", "", {}, "sha512-UXWMKhLOwVKb728IUtQPXxfYU+usdybtUrK/8uGE8CQMvrhOpwvzDBwj0QhSL7MQc7vIsISBG8VQ8+IDQxpfQA=="],
|
||||||
|
|
||||||
"strip-ansi": ["[email protected]", "", { "dependencies": { "ansi-regex": "^5.0.1" } }, "sha512-Y38VPSHcqkFrCpFnQ9vuSXmquuv5oXOKpGeT6aGrr3o3Gc9AlVa6JBfUSOCnbxGGZF+/0ooI7KrPuUSztUdU5A=="],
|
"strip-ansi": ["[email protected]", "", { "dependencies": { "ansi-regex": "^5.0.1" } }, "sha512-Y38VPSHcqkFrCpFnQ9vuSXmquuv5oXOKpGeT6aGrr3o3Gc9AlVa6JBfUSOCnbxGGZF+/0ooI7KrPuUSztUdU5A=="],
|
||||||
@@ -1002,6 +1037,10 @@
|
|||||||
|
|
||||||
"thenify-all": ["[email protected]", "", { "dependencies": { "thenify": ">= 3.1.0 < 4" } }, "sha512-RNxQH/qI8/t3thXJDwcstUO4zeqo64+Uy/+sNVRBx4Xn2OX+OZ9oP+iJnNFqplFra2ZUVeKCSa2oVWi3T4uVmA=="],
|
"thenify-all": ["[email protected]", "", { "dependencies": { "thenify": ">= 3.1.0 < 4" } }, "sha512-RNxQH/qI8/t3thXJDwcstUO4zeqo64+Uy/+sNVRBx4Xn2OX+OZ9oP+iJnNFqplFra2ZUVeKCSa2oVWi3T4uVmA=="],
|
||||||
|
|
||||||
|
"tiny-invariant": ["[email protected]", "", {}, "sha512-+FbBPE1o9QAYvviau/qC5SE3caw21q3xkvWKBtja5vgqOWIHHJ3ioaq1VPfn/Szqctz2bU/oYeKd9/z5BL+PVg=="],
|
||||||
|
|
||||||
|
"tiny-warning": ["[email protected]", "", {}, "sha512-lBN9zLN/oAf68o3zNXYrdCt1kP8WsiGW8Oo2ka41b2IM5JL/S1CTyX1rW0mb/zSuJun0ZUrDxx4sqvYS2FWzPA=="],
|
||||||
|
|
||||||
"tinyglobby": ["[email protected]", "", { "dependencies": { "fdir": "^6.5.0", "picomatch": "^4.0.3" } }, "sha512-j2Zq4NyQYG5XMST4cbs02Ak8iJUdxRM0XI5QyxXuZOzKOINmWurp3smXu3y5wDcJrptwpSjgXHzIQxR0omXljQ=="],
|
"tinyglobby": ["[email protected]", "", { "dependencies": { "fdir": "^6.5.0", "picomatch": "^4.0.3" } }, "sha512-j2Zq4NyQYG5XMST4cbs02Ak8iJUdxRM0XI5QyxXuZOzKOINmWurp3smXu3y5wDcJrptwpSjgXHzIQxR0omXljQ=="],
|
||||||
|
|
||||||
"to-regex-range": ["[email protected]", "", { "dependencies": { "is-number": "^7.0.0" } }, "sha512-65P7iz6X5yEr1cwcgvQxbbIw7Uk3gOy5dIdtZ4rDveLqhrdJP+Li/Hx6tyK0NEb+2GCyneCMJiGqrADCSNk8sQ=="],
|
"to-regex-range": ["[email protected]", "", { "dependencies": { "is-number": "^7.0.0" } }, "sha512-65P7iz6X5yEr1cwcgvQxbbIw7Uk3gOy5dIdtZ4rDveLqhrdJP+Li/Hx6tyK0NEb+2GCyneCMJiGqrADCSNk8sQ=="],
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
node_modules
|
||||||
|
.mintlify
|
||||||
|
.DS_Store
|
||||||
@@ -0,0 +1,64 @@
|
|||||||
|
# Voicebox Documentation
|
||||||
|
|
||||||
|
This directory contains the documentation for Voicebox, built with [Mintlify](https://mintlify.com).
|
||||||
|
|
||||||
|
## Development
|
||||||
|
|
||||||
|
### Prerequisites
|
||||||
|
|
||||||
|
Install Mintlify globally using bun:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bun add -g mintlify
|
||||||
|
```
|
||||||
|
|
||||||
|
Or use the helper script:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bun run install:mintlify
|
||||||
|
```
|
||||||
|
|
||||||
|
### Running Locally
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bun run dev
|
||||||
|
```
|
||||||
|
|
||||||
|
This will start the Mintlify dev server.
|
||||||
|
|
||||||
|
The docs will be available at `http://localhost:3000`
|
||||||
|
|
||||||
|
### Structure
|
||||||
|
|
||||||
|
```
|
||||||
|
docs/
|
||||||
|
├── mint.json # Mintlify configuration
|
||||||
|
├── custom.css # Custom styles
|
||||||
|
├── overview/ # Getting started & feature docs
|
||||||
|
├── guides/ # User guides
|
||||||
|
├── api/ # API reference
|
||||||
|
├── development/ # Developer documentation
|
||||||
|
├── logo/ # Logo assets
|
||||||
|
└── public/ # Static assets
|
||||||
|
```
|
||||||
|
|
||||||
|
### Writing Docs
|
||||||
|
|
||||||
|
- Use `.mdx` files for all documentation pages
|
||||||
|
- Follow the existing structure in `mint.json` for navigation
|
||||||
|
- Use Mintlify components for enhanced formatting (Card, CardGroup, Accordion, etc.)
|
||||||
|
- Reference the [Mintlify documentation](https://mintlify.com/docs) for available components
|
||||||
|
|
||||||
|
## Deployment
|
||||||
|
|
||||||
|
Docs are automatically deployed when changes are pushed to the main branch.
|
||||||
|
|
||||||
|
To manually deploy:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mintlify deploy
|
||||||
|
```
|
||||||
|
|
||||||
|
## Contributing
|
||||||
|
|
||||||
|
See [CONTRIBUTING.md](../CONTRIBUTING.md) for contribution guidelines.
|
||||||
+32
-4
@@ -90,6 +90,26 @@ chmod +x voicebox-*.AppImage
|
|||||||
- Slower but works without GPU
|
- Slower but works without GPU
|
||||||
- Backend automatically falls back to CPU
|
- Backend automatically falls back to CPU
|
||||||
|
|
||||||
|
### MLX "Failed to load the default metallib" error (Apple Silicon)
|
||||||
|
|
||||||
|
**Symptoms:** Generation fails with "library not found" or "metallib" errors
|
||||||
|
|
||||||
|
**Solutions:**
|
||||||
|
1. **Rebuild server binary**
|
||||||
|
```bash
|
||||||
|
bun run build:server
|
||||||
|
```
|
||||||
|
The build script should automatically include MLX Metal shader libraries.
|
||||||
|
|
||||||
|
2. **Check MLX installation**
|
||||||
|
```bash
|
||||||
|
pip install -r backend/requirements-mlx.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
3. **Verify backend detection**
|
||||||
|
- Check server logs for "Backend: MLX"
|
||||||
|
- If showing "Backend: PYTORCH", MLX may not be installed correctly
|
||||||
|
|
||||||
### Audio playback issues
|
### Audio playback issues
|
||||||
|
|
||||||
**Symptoms:** Generated audio won't play
|
**Symptoms:** Generated audio won't play
|
||||||
@@ -111,19 +131,27 @@ chmod +x voicebox-*.AppImage
|
|||||||
**Symptoms:** Generation takes >30 seconds
|
**Symptoms:** Generation takes >30 seconds
|
||||||
|
|
||||||
**Solutions:**
|
**Solutions:**
|
||||||
1. **Use GPU** (if available)
|
1. **Check backend type** (Apple Silicon)
|
||||||
|
- Check Settings → Server Status
|
||||||
|
- Should show "Backend: MLX" on Apple Silicon
|
||||||
|
- If showing "Backend: PYTORCH", install MLX: `pip install -r backend/requirements-mlx.txt`
|
||||||
|
- MLX provides 4-5x faster inference on Apple Silicon
|
||||||
|
|
||||||
|
2. **Use GPU** (if available)
|
||||||
- Check Settings → Server Status
|
- Check Settings → Server Status
|
||||||
- Should show "GPU available: true"
|
- Should show "GPU available: true"
|
||||||
|
- Apple Silicon: Should show "Metal (Apple Silicon via MLX)"
|
||||||
|
- Windows/Linux: Should show "CUDA" if GPU available
|
||||||
|
|
||||||
2. **Enable caching**
|
3. **Enable caching**
|
||||||
- Voice prompts are cached automatically
|
- Voice prompts are cached automatically
|
||||||
- Second generation with same voice should be faster
|
- Second generation with same voice should be faster
|
||||||
|
|
||||||
3. **Use smaller model**
|
4. **Use smaller model**
|
||||||
- 0.6B model is faster than 1.7B
|
- 0.6B model is faster than 1.7B
|
||||||
- Quality difference is minimal for most voices
|
- Quality difference is minimal for most voices
|
||||||
|
|
||||||
4. **Check system resources**
|
5. **Check system resources**
|
||||||
- Close other CPU/GPU intensive apps
|
- Close other CPU/GPU intensive apps
|
||||||
- Ensure adequate RAM (8GB+ recommended)
|
- Ensure adequate RAM (8GB+ recommended)
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,55 @@
|
|||||||
|
---
|
||||||
|
title: "Authentication"
|
||||||
|
description: "API authentication and security"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Current Status
|
||||||
|
|
||||||
|
<Warning>
|
||||||
|
Authentication is not currently implemented in Voicebox. The API is intended for local use only.
|
||||||
|
</Warning>
|
||||||
|
|
||||||
|
## Local Usage
|
||||||
|
|
||||||
|
For local development and usage:
|
||||||
|
- API runs on `localhost:17493`
|
||||||
|
- No authentication required
|
||||||
|
- Access restricted to local machine
|
||||||
|
|
||||||
|
## Future Implementation
|
||||||
|
|
||||||
|
Authentication will be added in a future release for:
|
||||||
|
- Remote deployments
|
||||||
|
- Multi-user access
|
||||||
|
- Production environments
|
||||||
|
|
||||||
|
Planned authentication methods:
|
||||||
|
- API keys
|
||||||
|
- OAuth 2.0
|
||||||
|
- JWT tokens
|
||||||
|
|
||||||
|
## Security Best Practices
|
||||||
|
|
||||||
|
Until authentication is implemented:
|
||||||
|
|
||||||
|
<CardGroup cols={2}>
|
||||||
|
<Card title="Use VPN" icon="shield">
|
||||||
|
Use WireGuard or Tailscale for remote access
|
||||||
|
</Card>
|
||||||
|
<Card title="Reverse Proxy" icon="server">
|
||||||
|
Run behind nginx with basic auth
|
||||||
|
</Card>
|
||||||
|
<Card title="Firewall" icon="fire">
|
||||||
|
Restrict access to trusted IPs only
|
||||||
|
</Card>
|
||||||
|
<Card title="Local Only" icon="laptop">
|
||||||
|
Don't expose to public internet
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
|
|
||||||
|
## Coming Soon
|
||||||
|
|
||||||
|
- API key management
|
||||||
|
- User accounts
|
||||||
|
- Rate limiting
|
||||||
|
- Access control
|
||||||
@@ -0,0 +1,119 @@
|
|||||||
|
---
|
||||||
|
title: "Generation API"
|
||||||
|
description: "Generate speech from text"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Generate Speech
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /generate
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"text": "Hello world",
|
||||||
|
"profile_id": "abc123",
|
||||||
|
"language": "en"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "gen123",
|
||||||
|
"text": "Hello world",
|
||||||
|
"profile_id": "abc123",
|
||||||
|
"language": "en",
|
||||||
|
"audio_url": "/audio/gen123.wav",
|
||||||
|
"duration": 2.3,
|
||||||
|
"created_at": "2024-01-29T12:00:00Z"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## List History
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /history
|
||||||
|
```
|
||||||
|
|
||||||
|
**Query Parameters:**
|
||||||
|
- `profile_id` (optional) - Filter by voice profile
|
||||||
|
- `limit` (optional) - Number of results (default: 50)
|
||||||
|
- `offset` (optional) - Pagination offset
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"generations": [
|
||||||
|
{
|
||||||
|
"id": "gen123",
|
||||||
|
"text": "Hello world",
|
||||||
|
"profile_id": "abc123",
|
||||||
|
"duration": 2.3,
|
||||||
|
"created_at": "2024-01-29T12:00:00Z"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"total": 100
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Get Generation
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /history/{id}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "gen123",
|
||||||
|
"text": "Hello world",
|
||||||
|
"profile_id": "abc123",
|
||||||
|
"language": "en",
|
||||||
|
"audio_url": "/audio/gen123.wav",
|
||||||
|
"duration": 2.3,
|
||||||
|
"created_at": "2024-01-29T12:00:00Z"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Delete Generation
|
||||||
|
|
||||||
|
```http
|
||||||
|
DELETE /history/{id}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## TypeScript Example
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { VoiceboxClient } from '@/lib/api'
|
||||||
|
|
||||||
|
const client = new VoiceboxClient({
|
||||||
|
baseUrl: 'http://localhost:17493'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Generate speech
|
||||||
|
const generation = await client.generate({
|
||||||
|
text: 'Hello world',
|
||||||
|
profile_id: 'abc123',
|
||||||
|
language: 'en'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Get audio URL
|
||||||
|
const audioUrl = generation.audio_url
|
||||||
|
|
||||||
|
// List history
|
||||||
|
const history = await client.listHistory({
|
||||||
|
profile_id: 'abc123',
|
||||||
|
limit: 20
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
For full API documentation, visit `http://localhost:17493/docs` when the server is running.
|
||||||
@@ -0,0 +1,219 @@
|
|||||||
|
---
|
||||||
|
title: "API Overview"
|
||||||
|
description: "Integrate voice synthesis into your applications with the Voicebox REST API"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Introduction
|
||||||
|
|
||||||
|
Voicebox exposes a full REST API that allows you to integrate voice synthesis into your own applications. The API runs on `http://localhost:17493` by default.
|
||||||
|
|
||||||
|
<Card title="Interactive API Docs" icon="book" href="http://localhost:17493/docs">
|
||||||
|
When Voicebox is running, visit the auto-generated API documentation at `http://localhost:17493/docs`
|
||||||
|
</Card>
|
||||||
|
|
||||||
|
## Base URL
|
||||||
|
|
||||||
|
```
|
||||||
|
http://localhost:17493
|
||||||
|
```
|
||||||
|
|
||||||
|
For remote deployments, replace `localhost` with your server's IP or hostname.
|
||||||
|
|
||||||
|
## Authentication
|
||||||
|
|
||||||
|
<Note>
|
||||||
|
Currently, the API does not require authentication for local development. Authentication will be added in a future release for production deployments.
|
||||||
|
</Note>
|
||||||
|
|
||||||
|
## Quick Example
|
||||||
|
|
||||||
|
Here's a simple example of generating speech:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Generate speech
|
||||||
|
curl -X POST http://localhost:17493/generate \
|
||||||
|
-H "Content-Type: application/json" \
|
||||||
|
-d '{
|
||||||
|
"text": "Hello world",
|
||||||
|
"profile_id": "abc123",
|
||||||
|
"language": "en"
|
||||||
|
}'
|
||||||
|
```
|
||||||
|
|
||||||
|
## API Endpoints
|
||||||
|
|
||||||
|
The Voicebox API is organized into several categories:
|
||||||
|
|
||||||
|
<CardGroup cols={2}>
|
||||||
|
<Card title="Voice Profiles" icon="user" href="/api/voice-profiles">
|
||||||
|
Create, list, update, and delete voice profiles
|
||||||
|
</Card>
|
||||||
|
<Card title="Generation" icon="waveform" href="/api/generation">
|
||||||
|
Generate speech from text using voice profiles
|
||||||
|
</Card>
|
||||||
|
<Card title="Recordings" icon="microphone" href="/api/recordings">
|
||||||
|
Record and transcribe audio
|
||||||
|
</Card>
|
||||||
|
<Card title="Stories" icon="film">
|
||||||
|
Create and manage multi-voice stories (coming soon)
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
|
|
||||||
|
## Core Endpoints
|
||||||
|
|
||||||
|
### Voice Profiles
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /profiles # List all profiles
|
||||||
|
POST /profiles # Create a new profile
|
||||||
|
GET /profiles/{id} # Get profile details
|
||||||
|
PUT /profiles/{id} # Update a profile
|
||||||
|
DELETE /profiles/{id} # Delete a profile
|
||||||
|
POST /profiles/{id}/samples # Add voice sample
|
||||||
|
```
|
||||||
|
|
||||||
|
### Generation
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /generate # Generate speech
|
||||||
|
GET /history # List generation history
|
||||||
|
GET /history/{id} # Get generation details
|
||||||
|
DELETE /history/{id} # Delete from history
|
||||||
|
```
|
||||||
|
|
||||||
|
### Recordings
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /recordings # Start recording
|
||||||
|
POST /recordings/stop # Stop recording
|
||||||
|
POST /transcribe # Transcribe audio
|
||||||
|
```
|
||||||
|
|
||||||
|
## Response Format
|
||||||
|
|
||||||
|
All API responses follow a consistent JSON format:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"success": true,
|
||||||
|
"data": {
|
||||||
|
// Response data
|
||||||
|
},
|
||||||
|
"error": null
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Error responses:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"success": false,
|
||||||
|
"data": null,
|
||||||
|
"error": {
|
||||||
|
"message": "Error description",
|
||||||
|
"code": "ERROR_CODE"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Data Models
|
||||||
|
|
||||||
|
### Voice Profile
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "abc123",
|
||||||
|
"name": "John Smith",
|
||||||
|
"language": "en",
|
||||||
|
"description": "Professional narrator voice",
|
||||||
|
"created_at": "2024-01-29T12:00:00Z",
|
||||||
|
"samples": [
|
||||||
|
{
|
||||||
|
"id": "sample123",
|
||||||
|
"audio_path": "/path/to/sample.wav",
|
||||||
|
"duration": 15.5
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Generation
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "gen123",
|
||||||
|
"text": "Hello world",
|
||||||
|
"profile_id": "abc123",
|
||||||
|
"language": "en",
|
||||||
|
"audio_path": "/path/to/output.wav",
|
||||||
|
"duration": 2.3,
|
||||||
|
"created_at": "2024-01-29T12:00:00Z"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## TypeScript Client
|
||||||
|
|
||||||
|
Voicebox provides an auto-generated TypeScript client with full type safety:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { VoiceboxClient } from '@/lib/api'
|
||||||
|
|
||||||
|
const client = new VoiceboxClient({
|
||||||
|
baseUrl: 'http://localhost:17493'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Create a profile
|
||||||
|
const profile = await client.createProfile({
|
||||||
|
name: 'John Smith',
|
||||||
|
language: 'en'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Generate speech
|
||||||
|
const generation = await client.generate({
|
||||||
|
text: 'Hello world',
|
||||||
|
profile_id: profile.id,
|
||||||
|
language: 'en'
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
The client is automatically generated from the OpenAPI schema. See [Development Setup](/development/setup#generate-openapi-client) for details.
|
||||||
|
|
||||||
|
## Rate Limiting
|
||||||
|
|
||||||
|
<Info>
|
||||||
|
Currently, there are no rate limits for local usage. Rate limiting will be added in a future release for production deployments.
|
||||||
|
</Info>
|
||||||
|
|
||||||
|
## WebSocket Support
|
||||||
|
|
||||||
|
<Note>
|
||||||
|
Real-time streaming generation via WebSockets is planned for a future release.
|
||||||
|
</Note>
|
||||||
|
|
||||||
|
## Use Cases
|
||||||
|
|
||||||
|
<CardGroup cols={2}>
|
||||||
|
<Card title="Game Development" icon="gamepad">
|
||||||
|
Generate dynamic dialogue for NPCs and characters
|
||||||
|
</Card>
|
||||||
|
<Card title="Content Creation" icon="video">
|
||||||
|
Automate voiceovers for videos and podcasts
|
||||||
|
</Card>
|
||||||
|
<Card title="Accessibility" icon="universal-access">
|
||||||
|
Build text-to-speech tools for visually impaired users
|
||||||
|
</Card>
|
||||||
|
<Card title="Voice Assistants" icon="robot">
|
||||||
|
Create custom voice interfaces
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
<CardGroup cols={2}>
|
||||||
|
<Card title="Voice Profiles API" icon="user" href="/api/voice-profiles">
|
||||||
|
Learn how to manage voice profiles
|
||||||
|
</Card>
|
||||||
|
<Card title="Generation API" icon="waveform" href="/api/generation">
|
||||||
|
Generate speech from text
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
@@ -0,0 +1,95 @@
|
|||||||
|
---
|
||||||
|
title: "Recordings API"
|
||||||
|
description: "Record and transcribe audio"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Start Recording
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /recordings/start
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"source": "microphone"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"recording_id": "rec123",
|
||||||
|
"status": "recording"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Stop Recording
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /recordings/stop
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"recording_id": "rec123"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"recording_id": "rec123",
|
||||||
|
"audio_url": "/audio/rec123.wav",
|
||||||
|
"duration": 15.5
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Transcribe Audio
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /transcribe
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request:** (multipart/form-data)
|
||||||
|
```
|
||||||
|
audio: <file>
|
||||||
|
language: "en" (optional)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"text": "Transcribed speech text here",
|
||||||
|
"language": "en",
|
||||||
|
"duration": 15.5,
|
||||||
|
"confidence": 0.95
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## TypeScript Example
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { VoiceboxClient } from '@/lib/api'
|
||||||
|
|
||||||
|
const client = new VoiceboxClient({
|
||||||
|
baseUrl: 'http://localhost:17493'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Start recording
|
||||||
|
const recording = await client.startRecording({
|
||||||
|
source: 'microphone'
|
||||||
|
})
|
||||||
|
|
||||||
|
// ... record audio ...
|
||||||
|
|
||||||
|
// Stop recording
|
||||||
|
const result = await client.stopRecording(recording.id)
|
||||||
|
|
||||||
|
// Transcribe
|
||||||
|
const transcription = await client.transcribe(audioFile, 'en')
|
||||||
|
console.log(transcription.text)
|
||||||
|
```
|
||||||
|
|
||||||
|
For full API documentation, visit `http://localhost:17493/docs` when the server is running.
|
||||||
@@ -0,0 +1,149 @@
|
|||||||
|
---
|
||||||
|
title: "Voice Profiles API"
|
||||||
|
description: "Manage voice profiles programmatically"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Endpoints
|
||||||
|
|
||||||
|
### List Profiles
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /profiles
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"profiles": [
|
||||||
|
{
|
||||||
|
"id": "abc123",
|
||||||
|
"name": "John Smith",
|
||||||
|
"language": "en",
|
||||||
|
"description": "Professional narrator",
|
||||||
|
"created_at": "2024-01-29T12:00:00Z",
|
||||||
|
"sample_count": 2
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Get Profile
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /profiles/{id}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "abc123",
|
||||||
|
"name": "John Smith",
|
||||||
|
"language": "en",
|
||||||
|
"description": "Professional narrator",
|
||||||
|
"created_at": "2024-01-29T12:00:00Z",
|
||||||
|
"samples": [
|
||||||
|
{
|
||||||
|
"id": "sample123",
|
||||||
|
"duration": 15.5,
|
||||||
|
"created_at": "2024-01-29T12:00:00Z"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Create Profile
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /profiles
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"name": "John Smith",
|
||||||
|
"language": "en",
|
||||||
|
"description": "Professional narrator"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "abc123",
|
||||||
|
"name": "John Smith",
|
||||||
|
"language": "en",
|
||||||
|
"description": "Professional narrator",
|
||||||
|
"created_at": "2024-01-29T12:00:00Z"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Update Profile
|
||||||
|
|
||||||
|
```http
|
||||||
|
PUT /profiles/{id}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"name": "Updated Name",
|
||||||
|
"description": "Updated description"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Delete Profile
|
||||||
|
|
||||||
|
```http
|
||||||
|
DELETE /profiles/{id}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Add Voice Sample
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /profiles/{id}/samples
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request:** (multipart/form-data)
|
||||||
|
```
|
||||||
|
audio: <file>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response:**
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"sample_id": "sample123",
|
||||||
|
"duration": 15.5
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## TypeScript Example
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { VoiceboxClient } from '@/lib/api'
|
||||||
|
|
||||||
|
const client = new VoiceboxClient({
|
||||||
|
baseUrl: 'http://localhost:17493'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Create profile
|
||||||
|
const profile = await client.createProfile({
|
||||||
|
name: 'John Smith',
|
||||||
|
language: 'en',
|
||||||
|
description: 'Professional narrator'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Add sample
|
||||||
|
await client.addSample(profile.id, audioFile)
|
||||||
|
|
||||||
|
// List all profiles
|
||||||
|
const profiles = await client.listProfiles()
|
||||||
|
```
|
||||||
|
|
||||||
|
For full API documentation, visit `http://localhost:17493/docs` when the server is running.
|
||||||
+1831
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,15 @@
|
|||||||
|
/* Anchor hover styles */
|
||||||
|
.nav-anchor:hover {
|
||||||
|
@apply text-[#BF9E40];
|
||||||
|
}
|
||||||
|
|
||||||
|
/* Icon wrapper on hover */
|
||||||
|
.nav-anchor:hover div {
|
||||||
|
background: #BF9E40 !important;
|
||||||
|
filter: brightness(1) !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* Icon SVG on hover */
|
||||||
|
.nav-anchor:hover svg {
|
||||||
|
@apply bg-white !important;
|
||||||
|
}
|
||||||
@@ -0,0 +1,206 @@
|
|||||||
|
---
|
||||||
|
title: "Architecture"
|
||||||
|
description: "Understanding Voicebox's technical architecture"
|
||||||
|
---
|
||||||
|
|
||||||
|
## System Overview
|
||||||
|
|
||||||
|
Voicebox uses a client-server architecture with a React frontend and Python backend. The desktop app is built with Tauri and contains two main layers:
|
||||||
|
|
||||||
|
**Frontend Layer:** A React application that handles the UI components, state management with Zustand, and data fetching with React Query (TanStack Query).
|
||||||
|
|
||||||
|
**Backend Layer:** A Python FastAPI server that provides the REST API, runs the TTS engine (Qwen3-TTS), manages the SQLite database, and handles audio processing.
|
||||||
|
|
||||||
|
These two layers communicate via HTTP, with the frontend making API requests to the backend.
|
||||||
|
|
||||||
|
## Frontend Architecture
|
||||||
|
|
||||||
|
### Tech Stack
|
||||||
|
|
||||||
|
- **Framework**: React 18 with TypeScript
|
||||||
|
- **State Management**: Zustand stores
|
||||||
|
- **Data Fetching**: React Query (TanStack Query)
|
||||||
|
- **Styling**: Tailwind CSS
|
||||||
|
- **Audio**: WaveSurfer.js
|
||||||
|
- **Desktop**: Tauri (Rust)
|
||||||
|
|
||||||
|
### Component Structure
|
||||||
|
|
||||||
|
```
|
||||||
|
app/src/
|
||||||
|
├── components/ # React components
|
||||||
|
│ ├── profiles/ # Voice profile UI
|
||||||
|
│ ├── generation/ # Speech generation UI
|
||||||
|
│ ├── stories/ # Timeline editor
|
||||||
|
│ └── shared/ # Reusable components
|
||||||
|
├── lib/ # Utilities
|
||||||
|
│ ├── api/ # Generated API client
|
||||||
|
│ └── utils/ # Helper functions
|
||||||
|
├── hooks/ # React hooks
|
||||||
|
└── stores/ # Zustand state stores
|
||||||
|
```
|
||||||
|
|
||||||
|
### State Management
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Example: Profile store
|
||||||
|
const useProfileStore = create((set) => ({
|
||||||
|
profiles: [],
|
||||||
|
selectedProfile: null,
|
||||||
|
setProfiles: (profiles) => set({ profiles }),
|
||||||
|
selectProfile: (id) => set({ selectedProfile: id })
|
||||||
|
}))
|
||||||
|
```
|
||||||
|
|
||||||
|
## Backend Architecture
|
||||||
|
|
||||||
|
### Tech Stack
|
||||||
|
|
||||||
|
- **Framework**: FastAPI (Python 3.11+)
|
||||||
|
- **TTS Model**: Qwen3-TTS
|
||||||
|
- **Transcription**: Whisper
|
||||||
|
- **Database**: SQLite
|
||||||
|
- **Audio**: librosa, soundfile
|
||||||
|
|
||||||
|
### API Structure
|
||||||
|
|
||||||
|
```python
|
||||||
|
# main.py - API routes
|
||||||
|
@app.post("/generate")
|
||||||
|
async def generate_speech(request: GenerateRequest):
|
||||||
|
# 1. Validate request
|
||||||
|
# 2. Load voice profile
|
||||||
|
# 3. Generate audio with TTS
|
||||||
|
# 4. Save to database
|
||||||
|
# 5. Return response
|
||||||
|
```
|
||||||
|
|
||||||
|
### Data Model
|
||||||
|
|
||||||
|
The database uses three main tables:
|
||||||
|
|
||||||
|
**Profile Table:** Stores voice profiles with fields for id, name, and language.
|
||||||
|
|
||||||
|
**Sample Table:** Stores audio samples linked to profiles via profile_id, with fields for audio_path and duration.
|
||||||
|
|
||||||
|
**Generation Table:** Stores generated audio with fields for id, profile_id, text, and audio_path.
|
||||||
|
|
||||||
|
## Desktop App (Tauri)
|
||||||
|
|
||||||
|
### Rust Backend
|
||||||
|
|
||||||
|
```rust
|
||||||
|
// Sidecar process management
|
||||||
|
// File system access
|
||||||
|
// Native integrations
|
||||||
|
```
|
||||||
|
|
||||||
|
### Responsibilities
|
||||||
|
|
||||||
|
- Launch Python backend as sidecar process
|
||||||
|
- Native file dialogs
|
||||||
|
- System tray integration
|
||||||
|
- Auto-updates
|
||||||
|
- OS-specific features
|
||||||
|
|
||||||
|
## Build Process
|
||||||
|
|
||||||
|
### Development
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Frontend (Vite dev server)
|
||||||
|
cd app && bun run dev
|
||||||
|
|
||||||
|
# Backend (manual start)
|
||||||
|
cd backend && uvicorn main:app --reload
|
||||||
|
|
||||||
|
# Desktop app (connects to manual backend)
|
||||||
|
bun run dev
|
||||||
|
```
|
||||||
|
|
||||||
|
### Production
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Build everything (server binary + Tauri app)
|
||||||
|
bun run build
|
||||||
|
|
||||||
|
# Or build separately:
|
||||||
|
# 1. Build server binary (PyInstaller)
|
||||||
|
bun run build:server
|
||||||
|
|
||||||
|
# 2. Build Tauri app (includes server)
|
||||||
|
cd tauri && bun run tauri build
|
||||||
|
```
|
||||||
|
|
||||||
|
## Data Flow
|
||||||
|
|
||||||
|
### Generation Flow
|
||||||
|
|
||||||
|
When a user generates speech, the data flows through the following stages:
|
||||||
|
|
||||||
|
1. **User Input** - User enters text in a React component
|
||||||
|
2. **State Update** - Text is stored in Zustand state
|
||||||
|
3. **API Request** - React Query mutation triggers an API call via fetch
|
||||||
|
4. **Backend Processing** - FastAPI endpoint receives the request
|
||||||
|
5. **TTS Generation** - Qwen3-TTS model generates the audio
|
||||||
|
6. **Storage** - Audio file is saved to disk and a database record is created
|
||||||
|
7. **Response** - Backend returns the audio URL
|
||||||
|
8. **Cache Update** - React Query updates its cache with the response
|
||||||
|
9. **UI Update** - Component re-renders with new data
|
||||||
|
10. **Playback** - User can play the generated audio
|
||||||
|
|
||||||
|
## Performance Considerations
|
||||||
|
|
||||||
|
### Frontend
|
||||||
|
|
||||||
|
- **Code splitting** - Lazy load routes
|
||||||
|
- **Memoization** - React.memo for heavy components
|
||||||
|
- **Virtual scrolling** - For large lists
|
||||||
|
- **Debouncing** - Search and input handling
|
||||||
|
|
||||||
|
### Backend
|
||||||
|
|
||||||
|
- **Async operations** - All I/O is async
|
||||||
|
- **Model caching** - Keep TTS model in memory
|
||||||
|
- **Voice prompt caching** - Reuse embeddings
|
||||||
|
- **Connection pooling** - Database connections
|
||||||
|
|
||||||
|
## Security
|
||||||
|
|
||||||
|
### Current
|
||||||
|
|
||||||
|
- Local-only by default
|
||||||
|
- No authentication (localhost trust)
|
||||||
|
- File system sandboxing via Tauri
|
||||||
|
|
||||||
|
### Planned
|
||||||
|
|
||||||
|
- API key authentication
|
||||||
|
- User accounts
|
||||||
|
- Rate limiting
|
||||||
|
- HTTPS support
|
||||||
|
|
||||||
|
## Deployment Modes
|
||||||
|
|
||||||
|
### Local Mode
|
||||||
|
|
||||||
|
- Backend runs as sidecar
|
||||||
|
- All data stays on device
|
||||||
|
- No network required
|
||||||
|
|
||||||
|
### Remote Mode
|
||||||
|
|
||||||
|
- Backend on separate machine
|
||||||
|
- Frontend connects via HTTP
|
||||||
|
- Shared infrastructure possible
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
<CardGroup cols={2}>
|
||||||
|
<Card title="Development Setup" icon="code" href="/development/setup">
|
||||||
|
Set up your dev environment
|
||||||
|
</Card>
|
||||||
|
<Card title="Contributing" icon="code-pull-request" href="/development/contributing">
|
||||||
|
Contribute to Voicebox
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
@@ -0,0 +1,310 @@
|
|||||||
|
---
|
||||||
|
title: "Audio Channels"
|
||||||
|
description: "How audio output routing works in Voicebox"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Audio channels allow routing voice output to different audio devices. This is useful for multi-output setups where different voices should play through different speakers or applications.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
**Channel:** A named audio bus that can be assigned to output devices.
|
||||||
|
|
||||||
|
**Device Mapping:** Links channels to OS audio device identifiers.
|
||||||
|
|
||||||
|
**Profile Mapping:** Links voice profiles to channels (many-to-many).
|
||||||
|
|
||||||
|
## Data Model
|
||||||
|
|
||||||
|
### AudioChannel Table
|
||||||
|
|
||||||
|
```python
|
||||||
|
class AudioChannel(Base):
|
||||||
|
__tablename__ = "audio_channels"
|
||||||
|
|
||||||
|
id = Column(String, primary_key=True)
|
||||||
|
name = Column(String, nullable=False)
|
||||||
|
is_default = Column(Boolean, default=False)
|
||||||
|
created_at = Column(DateTime)
|
||||||
|
```
|
||||||
|
|
||||||
|
### ChannelDeviceMapping Table
|
||||||
|
|
||||||
|
```python
|
||||||
|
class ChannelDeviceMapping(Base):
|
||||||
|
__tablename__ = "channel_device_mappings"
|
||||||
|
|
||||||
|
id = Column(String, primary_key=True)
|
||||||
|
channel_id = Column(String, ForeignKey("audio_channels.id"))
|
||||||
|
device_id = Column(String) # OS device identifier
|
||||||
|
```
|
||||||
|
|
||||||
|
### ProfileChannelMapping Table
|
||||||
|
|
||||||
|
```python
|
||||||
|
class ProfileChannelMapping(Base):
|
||||||
|
__tablename__ = "profile_channel_mappings"
|
||||||
|
|
||||||
|
profile_id = Column(String, ForeignKey("profiles.id"), primary_key=True)
|
||||||
|
channel_id = Column(String, ForeignKey("audio_channels.id"), primary_key=True)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Default Channel
|
||||||
|
|
||||||
|
A default channel is created on database initialization:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def init_db():
|
||||||
|
# Create default channel if it doesn't exist
|
||||||
|
default_channel = db.query(AudioChannel).filter(
|
||||||
|
AudioChannel.is_default == True
|
||||||
|
).first()
|
||||||
|
|
||||||
|
if not default_channel:
|
||||||
|
default_channel = AudioChannel(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
name="Default",
|
||||||
|
is_default=True
|
||||||
|
)
|
||||||
|
db.add(default_channel)
|
||||||
|
|
||||||
|
# Assign all existing profiles to default channel
|
||||||
|
profiles = db.query(VoiceProfile).all()
|
||||||
|
for profile in profiles:
|
||||||
|
mapping = ProfileChannelMapping(
|
||||||
|
profile_id=profile.id,
|
||||||
|
channel_id=default_channel.id
|
||||||
|
)
|
||||||
|
db.add(mapping)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Core Operations
|
||||||
|
|
||||||
|
### Creating a Channel
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def create_channel(
|
||||||
|
data: AudioChannelCreate,
|
||||||
|
db: Session,
|
||||||
|
) -> AudioChannelResponse:
|
||||||
|
# Check name uniqueness
|
||||||
|
existing = db.query(DBAudioChannel).filter_by(name=data.name).first()
|
||||||
|
if existing:
|
||||||
|
raise ValueError(f"Channel with name '{data.name}' already exists")
|
||||||
|
|
||||||
|
# Create channel
|
||||||
|
channel = DBAudioChannel(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
name=data.name,
|
||||||
|
is_default=False,
|
||||||
|
)
|
||||||
|
db.add(channel)
|
||||||
|
|
||||||
|
# Add device mappings
|
||||||
|
for device_id in data.device_ids:
|
||||||
|
mapping = DBChannelDeviceMapping(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
channel_id=channel.id,
|
||||||
|
device_id=device_id,
|
||||||
|
)
|
||||||
|
db.add(mapping)
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Updating a Channel
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def update_channel(
|
||||||
|
channel_id: str,
|
||||||
|
data: AudioChannelUpdate,
|
||||||
|
db: Session,
|
||||||
|
) -> AudioChannelResponse:
|
||||||
|
channel = db.query(DBAudioChannel).filter_by(id=channel_id).first()
|
||||||
|
|
||||||
|
# Cannot modify default channel
|
||||||
|
if channel.is_default:
|
||||||
|
raise ValueError("Cannot modify the default channel")
|
||||||
|
|
||||||
|
# Update name
|
||||||
|
if data.name is not None:
|
||||||
|
channel.name = data.name
|
||||||
|
|
||||||
|
# Update device mappings
|
||||||
|
if data.device_ids is not None:
|
||||||
|
# Delete existing
|
||||||
|
db.query(DBChannelDeviceMapping).filter_by(channel_id=channel_id).delete()
|
||||||
|
|
||||||
|
# Add new
|
||||||
|
for device_id in data.device_ids:
|
||||||
|
mapping = DBChannelDeviceMapping(
|
||||||
|
channel_id=channel.id,
|
||||||
|
device_id=device_id,
|
||||||
|
)
|
||||||
|
db.add(mapping)
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Deleting a Channel
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def delete_channel(channel_id: str, db: Session) -> bool:
|
||||||
|
channel = db.query(DBAudioChannel).filter_by(id=channel_id).first()
|
||||||
|
|
||||||
|
# Cannot delete default channel
|
||||||
|
if channel.is_default:
|
||||||
|
raise ValueError("Cannot delete the default channel")
|
||||||
|
|
||||||
|
# Delete device mappings
|
||||||
|
db.query(DBChannelDeviceMapping).filter_by(channel_id=channel_id).delete()
|
||||||
|
|
||||||
|
# Delete profile-channel mappings
|
||||||
|
db.query(DBProfileChannelMapping).filter_by(channel_id=channel_id).delete()
|
||||||
|
|
||||||
|
# Delete channel
|
||||||
|
db.delete(channel)
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
## Voice Assignment
|
||||||
|
|
||||||
|
### Assigning Voices to Channel
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def set_channel_voices(
|
||||||
|
channel_id: str,
|
||||||
|
data: ChannelVoiceAssignment,
|
||||||
|
db: Session,
|
||||||
|
) -> None:
|
||||||
|
# Verify channel exists
|
||||||
|
channel = db.query(DBAudioChannel).filter_by(id=channel_id).first()
|
||||||
|
if not channel:
|
||||||
|
raise ValueError(f"Channel {channel_id} not found")
|
||||||
|
|
||||||
|
# Verify all profiles exist
|
||||||
|
for profile_id in data.profile_ids:
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=profile_id).first()
|
||||||
|
if not profile:
|
||||||
|
raise ValueError(f"Profile {profile_id} not found")
|
||||||
|
|
||||||
|
# Delete existing mappings
|
||||||
|
db.query(DBProfileChannelMapping).filter_by(channel_id=channel_id).delete()
|
||||||
|
|
||||||
|
# Add new mappings
|
||||||
|
for profile_id in data.profile_ids:
|
||||||
|
mapping = DBProfileChannelMapping(
|
||||||
|
profile_id=profile_id,
|
||||||
|
channel_id=channel_id,
|
||||||
|
)
|
||||||
|
db.add(mapping)
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Assigning Channels to Voice
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def set_profile_channels(
|
||||||
|
profile_id: str,
|
||||||
|
data: ProfileChannelAssignment,
|
||||||
|
db: Session,
|
||||||
|
) -> None:
|
||||||
|
# Verify profile exists
|
||||||
|
profile = db.query(DBVoiceProfile).filter_by(id=profile_id).first()
|
||||||
|
if not profile:
|
||||||
|
raise ValueError(f"Profile {profile_id} not found")
|
||||||
|
|
||||||
|
# Delete existing mappings
|
||||||
|
db.query(DBProfileChannelMapping).filter_by(profile_id=profile_id).delete()
|
||||||
|
|
||||||
|
# Add new mappings
|
||||||
|
for channel_id in data.channel_ids:
|
||||||
|
mapping = DBProfileChannelMapping(
|
||||||
|
profile_id=profile_id,
|
||||||
|
channel_id=channel_id,
|
||||||
|
)
|
||||||
|
db.add(mapping)
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
## API Endpoints
|
||||||
|
|
||||||
|
| Method | Endpoint | Description |
|
||||||
|
|--------|----------|-------------|
|
||||||
|
| GET | `/channels` | List all channels |
|
||||||
|
| POST | `/channels` | Create a channel |
|
||||||
|
| GET | `/channels/{id}` | Get channel by ID |
|
||||||
|
| PUT | `/channels/{id}` | Update channel |
|
||||||
|
| DELETE | `/channels/{id}` | Delete channel |
|
||||||
|
| GET | `/channels/{id}/voices` | Get assigned voices |
|
||||||
|
| PUT | `/channels/{id}/voices` | Set assigned voices |
|
||||||
|
| GET | `/profiles/{id}/channels` | Get profile's channels |
|
||||||
|
| PUT | `/profiles/{id}/channels` | Set profile's channels |
|
||||||
|
|
||||||
|
## Request/Response Schemas
|
||||||
|
|
||||||
|
### AudioChannelCreate
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"name": "Speakers",
|
||||||
|
"device_ids": ["device_uuid_1", "device_uuid_2"]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### AudioChannelResponse
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "channel_uuid",
|
||||||
|
"name": "Speakers",
|
||||||
|
"is_default": false,
|
||||||
|
"device_ids": ["device_uuid_1", "device_uuid_2"],
|
||||||
|
"created_at": "2024-01-15T10:30:00Z"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### ChannelVoiceAssignment
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"profile_ids": ["profile_1", "profile_2"]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Use Cases
|
||||||
|
|
||||||
|
### Multi-Output Setup
|
||||||
|
|
||||||
|
**Scenario:** Stream with different voice characters
|
||||||
|
|
||||||
|
1. Create "Stream" channel → OBS virtual audio
|
||||||
|
2. Create "Monitor" channel → Headphones
|
||||||
|
3. Assign "Narrator" profile → Both channels
|
||||||
|
4. Assign "Character 1" profile → Stream only
|
||||||
|
|
||||||
|
### Virtual Audio Cables
|
||||||
|
|
||||||
|
Common device IDs for virtual audio:
|
||||||
|
- VB-Audio Virtual Cable
|
||||||
|
- BlackHole (macOS)
|
||||||
|
- Soundflower (macOS)
|
||||||
|
|
||||||
|
## Frontend Integration
|
||||||
|
|
||||||
|
The frontend needs to:
|
||||||
|
|
||||||
|
1. **Enumerate devices** using Web Audio API or Tauri
|
||||||
|
2. **Display channel list** with device assignments
|
||||||
|
3. **Allow profile assignment** via drag/drop or dropdown
|
||||||
|
4. **Route playback** to correct device based on profile's channel
|
||||||
|
|
||||||
|
## Limitations
|
||||||
|
|
||||||
|
- Device IDs are OS-specific
|
||||||
|
- Hot-plugging may invalidate device IDs
|
||||||
|
- Default channel cannot be modified/deleted
|
||||||
|
- Frontend handles actual audio routing (backend just stores config)
|
||||||
@@ -0,0 +1,84 @@
|
|||||||
|
---
|
||||||
|
title: "Auto-Updater"
|
||||||
|
description: "Configure and use the Tauri auto-updater"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Voicebox uses Tauri's built-in auto-updater to deliver updates to users automatically.
|
||||||
|
|
||||||
|
## Quick Reference
|
||||||
|
|
||||||
|
For detailed setup instructions, see the existing documentation:
|
||||||
|
|
||||||
|
- [AUTOUPDATER_QUICKSTART.md](https://github.com/jamiepine/voicebox/blob/main/docs/AUTOUPDATER_QUICKSTART.md)
|
||||||
|
- [AUTOUPDATER.md](https://github.com/jamiepine/voicebox/blob/main/docs/AUTOUPDATER.md)
|
||||||
|
|
||||||
|
## How It Works
|
||||||
|
|
||||||
|
The auto-updater follows a secure update process:
|
||||||
|
|
||||||
|
1. **Check for Updates** - The Voicebox app periodically checks GitHub Releases for new versions
|
||||||
|
2. **Download Update** - If a new version is found, the update package is downloaded
|
||||||
|
3. **Verify Signature** - The downloaded package is cryptographically verified using the public key
|
||||||
|
4. **Install** - After verification, the update is installed
|
||||||
|
5. **Restart** - The app restarts with the new version
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
Updates are configured in `tauri/src-tauri/tauri.conf.json`:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"updater": {
|
||||||
|
"active": true,
|
||||||
|
"endpoints": [
|
||||||
|
"https://github.com/jamiepine/voicebox/releases/latest/download/latest.json"
|
||||||
|
],
|
||||||
|
"dialog": true,
|
||||||
|
"pubkey": "YOUR_PUBLIC_KEY"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Generating Keys
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Generate signing keys
|
||||||
|
bun run generate:keys
|
||||||
|
|
||||||
|
# Keys saved to ~/.tauri/voicebox.key
|
||||||
|
```
|
||||||
|
|
||||||
|
<Warning>
|
||||||
|
Keep your private key secure! Never commit it to the repository.
|
||||||
|
</Warning>
|
||||||
|
|
||||||
|
## Release Process
|
||||||
|
|
||||||
|
1. **Bump version** using bumpversion
|
||||||
|
2. **Push tag** to trigger CI/CD
|
||||||
|
3. **GitHub Actions** builds and signs releases
|
||||||
|
4. **Users** receive update notification
|
||||||
|
|
||||||
|
## User Experience
|
||||||
|
|
||||||
|
When an update is available:
|
||||||
|
|
||||||
|
1. User sees a notification dialog
|
||||||
|
2. User clicks "Update"
|
||||||
|
3. Update downloads in background
|
||||||
|
4. App restarts with new version
|
||||||
|
|
||||||
|
## For Developers
|
||||||
|
|
||||||
|
See the full documentation files for:
|
||||||
|
|
||||||
|
- Setting up signing keys
|
||||||
|
- Configuring GitHub releases
|
||||||
|
- Testing updates locally
|
||||||
|
- Troubleshooting update failures
|
||||||
|
|
||||||
|
<Card title="View Full Docs" href="https://github.com/jamiepine/voicebox/tree/main/docs">
|
||||||
|
Access AUTOUPDATER.md and AUTOUPDATER_QUICKSTART.md in the repository
|
||||||
|
</Card>
|
||||||
@@ -0,0 +1,270 @@
|
|||||||
|
---
|
||||||
|
title: "Building"
|
||||||
|
description: "Build Voicebox for production"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Voicebox uses a multi-step build process to create platform-specific installers.
|
||||||
|
|
||||||
|
## Quick Build
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Build for your current platform (automatically builds server binary first)
|
||||||
|
make build
|
||||||
|
|
||||||
|
# Or manually
|
||||||
|
bun run build
|
||||||
|
```
|
||||||
|
|
||||||
|
This automatically:
|
||||||
|
1. Builds the Python server binary (`bun run build:server`)
|
||||||
|
2. Builds the Tauri app (`cd tauri && bun run tauri build`)
|
||||||
|
|
||||||
|
## Build Process
|
||||||
|
|
||||||
|
The build process consists of two steps, but `bun run build` handles both automatically:
|
||||||
|
|
||||||
|
### 1. Server Binary Build (Automatic)
|
||||||
|
|
||||||
|
The Python backend is compiled into a standalone executable using PyInstaller. This happens automatically when you run `bun run build`.
|
||||||
|
|
||||||
|
**Platform-specific binaries:**
|
||||||
|
- macOS (Apple Silicon): `voicebox-server-aarch64-apple-darwin` (includes MLX backend)
|
||||||
|
- macOS (Intel): `voicebox-server-x86_64-apple-darwin` (PyTorch backend)
|
||||||
|
- Windows: `voicebox-server-x86_64-pc-windows-msvc.exe` (PyTorch backend)
|
||||||
|
- Linux: `voicebox-server-x86_64-unknown-linux-gnu` (PyTorch backend)
|
||||||
|
|
||||||
|
<Note>
|
||||||
|
The build script automatically detects your platform and includes the appropriate backend (MLX for Apple Silicon, PyTorch for others).
|
||||||
|
</Note>
|
||||||
|
|
||||||
|
**Manual build (if needed):**
|
||||||
|
```bash
|
||||||
|
bun run build:server
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Tauri App Build (Automatic)
|
||||||
|
|
||||||
|
The Tauri app build is also handled automatically, which:
|
||||||
|
1. Builds the React frontend (Vite)
|
||||||
|
2. Compiles the Rust backend
|
||||||
|
3. Bundles the server binary as a sidecar
|
||||||
|
4. Creates platform-specific installers
|
||||||
|
|
||||||
|
**Manual build (if needed):**
|
||||||
|
```bash
|
||||||
|
cd tauri && bun run tauri build
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Output
|
||||||
|
|
||||||
|
Installers are created in `tauri/src-tauri/target/release/bundle/`:
|
||||||
|
|
||||||
|
**macOS:**
|
||||||
|
- `dmg/` - Disk image installer
|
||||||
|
- `macos/` - App bundle
|
||||||
|
|
||||||
|
**Windows:**
|
||||||
|
- `msi/` - MSI installer
|
||||||
|
- `nsis/` - NSIS installer
|
||||||
|
|
||||||
|
**Linux:**
|
||||||
|
- `deb/` - Debian package
|
||||||
|
- `appimage/` - AppImage
|
||||||
|
|
||||||
|
## Advanced Options
|
||||||
|
|
||||||
|
### Building for Specific Platform
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Build for macOS (Apple Silicon)
|
||||||
|
bun run tauri build -- --target aarch64-apple-darwin
|
||||||
|
|
||||||
|
# Build for macOS (Intel)
|
||||||
|
bun run tauri build -- --target x86_64-apple-darwin
|
||||||
|
|
||||||
|
# Build for Windows
|
||||||
|
bun run tauri build -- --target x86_64-pc-windows-msvc
|
||||||
|
|
||||||
|
# Build for Linux
|
||||||
|
bun run tauri build -- --target x86_64-unknown-linux-gnu
|
||||||
|
```
|
||||||
|
|
||||||
|
### Using Local Qwen3-TTS
|
||||||
|
|
||||||
|
If you're developing Qwen3-TTS locally:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export QWEN_TTS_PATH=~/path/to/Qwen3-TTS
|
||||||
|
bun run build:server # Build server binary only
|
||||||
|
# or
|
||||||
|
bun run build # Build everything
|
||||||
|
```
|
||||||
|
|
||||||
|
This makes PyInstaller use your local version instead of the pip package.
|
||||||
|
|
||||||
|
### Debug Build
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd tauri
|
||||||
|
bun run tauri build --debug
|
||||||
|
```
|
||||||
|
|
||||||
|
Creates a debug build with symbols and logging.
|
||||||
|
|
||||||
|
## Build Configuration
|
||||||
|
|
||||||
|
### Tauri Config
|
||||||
|
|
||||||
|
Edit `tauri/src-tauri/tauri.conf.json`:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"bundle": {
|
||||||
|
"identifier": "com.voicebox.app",
|
||||||
|
"icon": [
|
||||||
|
"icons/32x32.png",
|
||||||
|
"icons/128x128.png",
|
||||||
|
"icons/icon.icns",
|
||||||
|
"icons/icon.ico"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Sidecar Configuration
|
||||||
|
|
||||||
|
The Python server is bundled as a sidecar:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"tauri": {
|
||||||
|
"bundle": {
|
||||||
|
"externalBin": [
|
||||||
|
"binaries/voicebox-server"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Code Signing
|
||||||
|
|
||||||
|
### macOS
|
||||||
|
|
||||||
|
To sign the app for distribution:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Set signing identity
|
||||||
|
export APPLE_SIGNING_IDENTITY="Developer ID Application: Your Name"
|
||||||
|
|
||||||
|
# Build with signing
|
||||||
|
bun run tauri build
|
||||||
|
```
|
||||||
|
|
||||||
|
For notarization:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Set credentials
|
||||||
|
export APPLE_ID="[email protected]"
|
||||||
|
export APPLE_PASSWORD="app-specific-password"
|
||||||
|
|
||||||
|
# Build and notarize
|
||||||
|
bun run tauri build
|
||||||
|
```
|
||||||
|
|
||||||
|
### Windows
|
||||||
|
|
||||||
|
For Windows code signing:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Set certificate
|
||||||
|
export WINDOWS_CERTIFICATE_PATH="/path/to/cert.pfx"
|
||||||
|
export WINDOWS_CERTIFICATE_PASSWORD="password"
|
||||||
|
|
||||||
|
# Build with signing
|
||||||
|
bun run tauri build
|
||||||
|
```
|
||||||
|
|
||||||
|
## Release Process
|
||||||
|
|
||||||
|
The full release process is automated:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. Bump version
|
||||||
|
bumpversion patch # or minor/major
|
||||||
|
|
||||||
|
# 2. Build all platforms (CI/CD handles this)
|
||||||
|
git push --tags
|
||||||
|
|
||||||
|
# 3. GitHub Actions creates releases
|
||||||
|
```
|
||||||
|
|
||||||
|
See [CONTRIBUTING.md](/development/contributing) for the full release workflow.
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
<AccordionGroup>
|
||||||
|
<Accordion title="Server Binary Build Fails">
|
||||||
|
**Common issues:**
|
||||||
|
- Missing Python dependencies: `pip install -r requirements.txt`
|
||||||
|
- PyInstaller not found: `pip install pyinstaller`
|
||||||
|
- Qwen3-TTS not installed: `pip install git+https://github.com/QwenLM/Qwen3-TTS.git`
|
||||||
|
|
||||||
|
**Solution:**
|
||||||
|
```bash
|
||||||
|
cd backend
|
||||||
|
source venv/bin/activate
|
||||||
|
pip install -r requirements.txt
|
||||||
|
pip install pyinstaller
|
||||||
|
```
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="Tauri Build Fails">
|
||||||
|
**Common issues:**
|
||||||
|
- Rust not installed: `curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh`
|
||||||
|
- Server binary missing: Usually auto-built, but can run manually: `./scripts/build-server.sh`
|
||||||
|
- Node modules outdated: `bun install`
|
||||||
|
|
||||||
|
**Solution:**
|
||||||
|
```bash
|
||||||
|
# Clean and rebuild
|
||||||
|
cd tauri/src-tauri
|
||||||
|
cargo clean
|
||||||
|
cd ../..
|
||||||
|
bun run build # Automatically builds server binary first
|
||||||
|
```
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="App Won't Launch After Build">
|
||||||
|
**Check:**
|
||||||
|
- Server binary has execute permissions
|
||||||
|
- All dependencies are bundled
|
||||||
|
- Check logs in the app's data directory
|
||||||
|
|
||||||
|
**macOS:**
|
||||||
|
```bash
|
||||||
|
tail -f ~/Library/Application\ Support/com.voicebox.app/logs/server.log
|
||||||
|
```
|
||||||
|
|
||||||
|
**Windows:**
|
||||||
|
```bash
|
||||||
|
type %APPDATA%\com.voicebox.app\logs\server.log
|
||||||
|
```
|
||||||
|
</Accordion>
|
||||||
|
</AccordionGroup>
|
||||||
|
|
||||||
|
## CI/CD
|
||||||
|
|
||||||
|
GitHub Actions automatically builds releases when tags are pushed:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
# .github/workflows/release.yml
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
tags:
|
||||||
|
- 'v*'
|
||||||
|
```
|
||||||
|
|
||||||
|
See the [repository](https://github.com/jamiepine/voicebox) for the full CI/CD configuration.
|
||||||
@@ -0,0 +1,326 @@
|
|||||||
|
---
|
||||||
|
title: "Contributing"
|
||||||
|
description: "How to contribute to Voicebox"
|
||||||
|
---
|
||||||
|
|
||||||
|
Thank you for your interest in contributing to Voicebox! This guide will help you get started.
|
||||||
|
|
||||||
|
## Code of Conduct
|
||||||
|
|
||||||
|
- Be respectful and inclusive
|
||||||
|
- Welcome newcomers and help them learn
|
||||||
|
- Focus on constructive feedback
|
||||||
|
- Respect different viewpoints and experiences
|
||||||
|
|
||||||
|
## Getting Started
|
||||||
|
|
||||||
|
Before you start contributing, make sure you have:
|
||||||
|
|
||||||
|
1. **Read the documentation** to understand how Voicebox works
|
||||||
|
2. **Set up your development environment** - see [Development Setup](/development/setup)
|
||||||
|
3. **Explored the codebase** to understand the project structure
|
||||||
|
4. **Checked existing issues** to see if someone else is working on something similar
|
||||||
|
|
||||||
|
## Ways to Contribute
|
||||||
|
|
||||||
|
<CardGroup cols={2}>
|
||||||
|
<Card title="Report Bugs" icon="bug">
|
||||||
|
Found a bug? Open an issue with reproduction steps
|
||||||
|
</Card>
|
||||||
|
<Card title="Request Features" icon="lightbulb">
|
||||||
|
Have an idea? Start a discussion or open an issue
|
||||||
|
</Card>
|
||||||
|
<Card title="Improve Docs" icon="book">
|
||||||
|
Fix typos, add examples, or clarify instructions
|
||||||
|
</Card>
|
||||||
|
<Card title="Write Code" icon="code">
|
||||||
|
Fix bugs, add features, or optimize performance
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
|
|
||||||
|
## Development Workflow
|
||||||
|
|
||||||
|
### 1. Fork & Clone
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Fork the repository on GitHub
|
||||||
|
# Then clone your fork
|
||||||
|
git clone https://github.com/YOUR_USERNAME/voicebox.git
|
||||||
|
cd voicebox
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Create a Branch
|
||||||
|
|
||||||
|
Use descriptive branch names:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# For features
|
||||||
|
git checkout -b feature/voice-effects
|
||||||
|
|
||||||
|
# For bug fixes
|
||||||
|
git checkout -b fix/audio-playback-issue
|
||||||
|
|
||||||
|
# For documentation
|
||||||
|
git checkout -b docs/api-examples
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Make Your Changes
|
||||||
|
|
||||||
|
Follow these guidelines:
|
||||||
|
|
||||||
|
<AccordionGroup>
|
||||||
|
<Accordion title="Code Style">
|
||||||
|
**TypeScript/React:**
|
||||||
|
- Use TypeScript strict mode
|
||||||
|
- Prefer functional components with hooks
|
||||||
|
- Use named exports
|
||||||
|
- Format with Biome (runs automatically)
|
||||||
|
|
||||||
|
**Python:**
|
||||||
|
- Follow PEP 8
|
||||||
|
- Use type hints
|
||||||
|
- Use async/await for I/O
|
||||||
|
- Document functions with docstrings
|
||||||
|
|
||||||
|
**Rust:**
|
||||||
|
- Follow Rust conventions
|
||||||
|
- Use meaningful names
|
||||||
|
- Handle errors explicitly
|
||||||
|
- Run `rustfmt`
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="Commit Messages">
|
||||||
|
Write clear, descriptive commit messages:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Good
|
||||||
|
git commit -m "Add voice profile export feature"
|
||||||
|
git commit -m "Fix audio playback stopping after 30 seconds"
|
||||||
|
|
||||||
|
# Avoid
|
||||||
|
git commit -m "Update code"
|
||||||
|
git commit -m "Fix bug"
|
||||||
|
```
|
||||||
|
|
||||||
|
Format:
|
||||||
|
- Use imperative mood ("Add feature" not "Added feature")
|
||||||
|
- Keep first line under 50 characters
|
||||||
|
- Add detailed description if needed
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="Testing">
|
||||||
|
- Test your changes manually in the app
|
||||||
|
- Ensure backend API endpoints work
|
||||||
|
- Check for TypeScript/Python errors
|
||||||
|
- Verify UI components render correctly
|
||||||
|
- Add automated tests when possible
|
||||||
|
</Accordion>
|
||||||
|
</AccordionGroup>
|
||||||
|
|
||||||
|
### 4. Push & Create PR
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Push your branch
|
||||||
|
git push origin feature/your-feature-name
|
||||||
|
|
||||||
|
# Then create a pull request on GitHub
|
||||||
|
```
|
||||||
|
|
||||||
|
## Pull Request Guidelines
|
||||||
|
|
||||||
|
When creating a pull request:
|
||||||
|
|
||||||
|
<Steps>
|
||||||
|
<Step title="Use a Clear Title">
|
||||||
|
Examples:
|
||||||
|
- "Add voice profile export functionality"
|
||||||
|
- "Fix audio playback stopping after 30 seconds"
|
||||||
|
- "Improve generation speed with caching"
|
||||||
|
</Step>
|
||||||
|
|
||||||
|
<Step title="Provide Description">
|
||||||
|
Include:
|
||||||
|
- What changes you made
|
||||||
|
- Why you made them
|
||||||
|
- How to test them
|
||||||
|
- Screenshots (for UI changes)
|
||||||
|
- Reference related issues
|
||||||
|
</Step>
|
||||||
|
|
||||||
|
<Step title="Update Documentation">
|
||||||
|
- Update relevant docs if behavior changes
|
||||||
|
- Add API documentation for new endpoints
|
||||||
|
- Update README if needed
|
||||||
|
</Step>
|
||||||
|
|
||||||
|
<Step title="Check the Checklist">
|
||||||
|
- [ ] Code follows style guidelines
|
||||||
|
- [ ] Documentation updated
|
||||||
|
- [ ] Changes tested
|
||||||
|
- [ ] No breaking changes (or documented)
|
||||||
|
- [ ] CHANGELOG.md updated
|
||||||
|
</Step>
|
||||||
|
</Steps>
|
||||||
|
|
||||||
|
## Project Structure
|
||||||
|
|
||||||
|
Understanding the codebase:
|
||||||
|
|
||||||
|
```
|
||||||
|
voicebox/
|
||||||
|
├── app/ # Shared React frontend
|
||||||
|
│ ├── src/
|
||||||
|
│ │ ├── components/ # UI components
|
||||||
|
│ │ ├── lib/ # Utilities and API client
|
||||||
|
│ │ ├── hooks/ # React hooks
|
||||||
|
│ │ └── stores/ # Zustand state stores
|
||||||
|
├── backend/ # Python FastAPI server
|
||||||
|
│ ├── main.py # API routes
|
||||||
|
│ ├── tts.py # Voice synthesis logic
|
||||||
|
│ ├── database.py # SQLite operations
|
||||||
|
│ └── models.py # Pydantic models
|
||||||
|
├── tauri/ # Desktop app wrapper
|
||||||
|
│ └── src-tauri/ # Rust backend
|
||||||
|
├── web/ # Web deployment
|
||||||
|
├── landing/ # Marketing website
|
||||||
|
└── scripts/ # Build & release scripts
|
||||||
|
```
|
||||||
|
|
||||||
|
## Areas for Contribution
|
||||||
|
|
||||||
|
### Bug Fixes
|
||||||
|
|
||||||
|
- Check [existing issues](https://github.com/jamiepine/voicebox/issues) for bugs
|
||||||
|
- Test your fix thoroughly
|
||||||
|
- Add regression tests if possible
|
||||||
|
|
||||||
|
### New Features
|
||||||
|
|
||||||
|
- Check the [roadmap](https://github.com/jamiepine/voicebox#roadmap) for planned features
|
||||||
|
- Discuss major features in an issue first
|
||||||
|
- Keep features focused and well-scoped
|
||||||
|
|
||||||
|
### Documentation
|
||||||
|
|
||||||
|
- Improve clarity and fix typos
|
||||||
|
- Add code examples
|
||||||
|
- Create tutorials or guides
|
||||||
|
- Document API endpoints
|
||||||
|
|
||||||
|
### UI/UX Improvements
|
||||||
|
|
||||||
|
- Improve accessibility
|
||||||
|
- Enhance visual design
|
||||||
|
- Optimize performance
|
||||||
|
- Add animations/transitions
|
||||||
|
|
||||||
|
### Infrastructure
|
||||||
|
|
||||||
|
- Improve build process
|
||||||
|
- Add CI/CD improvements
|
||||||
|
- Optimize bundle size
|
||||||
|
- Add testing infrastructure
|
||||||
|
|
||||||
|
## API Development
|
||||||
|
|
||||||
|
When adding new API endpoints:
|
||||||
|
|
||||||
|
<Steps>
|
||||||
|
<Step title="Add Route">
|
||||||
|
In `backend/main.py`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
@app.post("/api/new-endpoint")
|
||||||
|
async def new_endpoint(data: RequestModel) -> ResponseModel:
|
||||||
|
"""Endpoint description."""
|
||||||
|
# Implementation
|
||||||
|
return response
|
||||||
|
```
|
||||||
|
</Step>
|
||||||
|
|
||||||
|
<Step title="Create Models">
|
||||||
|
In `backend/models.py`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class RequestModel(BaseModel):
|
||||||
|
field: str
|
||||||
|
|
||||||
|
class ResponseModel(BaseModel):
|
||||||
|
result: str
|
||||||
|
```
|
||||||
|
</Step>
|
||||||
|
|
||||||
|
<Step title="Regenerate Client">
|
||||||
|
```bash
|
||||||
|
bun run generate:api
|
||||||
|
```
|
||||||
|
|
||||||
|
This updates the TypeScript client with type-safe bindings.
|
||||||
|
</Step>
|
||||||
|
|
||||||
|
<Step title="Update Docs">
|
||||||
|
Add documentation in `/docs/api/`
|
||||||
|
</Step>
|
||||||
|
</Steps>
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
Currently testing is primarily manual. When adding tests:
|
||||||
|
|
||||||
|
**Backend:**
|
||||||
|
```bash
|
||||||
|
cd backend
|
||||||
|
pytest
|
||||||
|
```
|
||||||
|
|
||||||
|
**Frontend:**
|
||||||
|
```bash
|
||||||
|
bun run test
|
||||||
|
```
|
||||||
|
|
||||||
|
**E2E (future):**
|
||||||
|
```bash
|
||||||
|
bun run test:e2e
|
||||||
|
```
|
||||||
|
|
||||||
|
## Release Process
|
||||||
|
|
||||||
|
Releases are managed by maintainers using `bumpversion`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Bump version (patch, minor, or major)
|
||||||
|
bumpversion patch
|
||||||
|
|
||||||
|
# Push with tags
|
||||||
|
git push && git push --tags
|
||||||
|
```
|
||||||
|
|
||||||
|
GitHub Actions automatically builds and publishes releases when tags are pushed.
|
||||||
|
|
||||||
|
## Community
|
||||||
|
|
||||||
|
- **GitHub Issues:** Bug reports and feature requests
|
||||||
|
- **GitHub Discussions:** General questions and ideas
|
||||||
|
- **Discord:** Real-time chat (coming soon)
|
||||||
|
|
||||||
|
## Recognition
|
||||||
|
|
||||||
|
Contributors are recognized in:
|
||||||
|
- [CHANGELOG.md](https://github.com/jamiepine/voicebox/blob/main/CHANGELOG.md)
|
||||||
|
- GitHub contributor list
|
||||||
|
- Release notes
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
By contributing, you agree that your contributions will be licensed under the MIT License.
|
||||||
|
|
||||||
|
## Questions?
|
||||||
|
|
||||||
|
If you have questions:
|
||||||
|
|
||||||
|
1. Check the [documentation](/overview/introduction)
|
||||||
|
2. Search [existing issues](https://github.com/jamiepine/voicebox/issues)
|
||||||
|
3. Open a new issue or discussion
|
||||||
|
4. See [CONTRIBUTING.md](https://github.com/jamiepine/voicebox/blob/main/CONTRIBUTING.md) in the repo
|
||||||
|
|
||||||
|
Thank you for contributing to Voicebox! 🎉
|
||||||
@@ -0,0 +1,260 @@
|
|||||||
|
---
|
||||||
|
title: "Generation History"
|
||||||
|
description: "How generation history tracking works in Voicebox"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
The history module tracks all generated audio, providing a searchable record of past generations. Each generation stores the text, settings, and a reference to the audio file.
|
||||||
|
|
||||||
|
## Data Model
|
||||||
|
|
||||||
|
### Generation Table
|
||||||
|
|
||||||
|
```python
|
||||||
|
class Generation(Base):
|
||||||
|
__tablename__ = "generations"
|
||||||
|
|
||||||
|
id = Column(String, primary_key=True)
|
||||||
|
profile_id = Column(String, ForeignKey("profiles.id"))
|
||||||
|
text = Column(Text, nullable=False)
|
||||||
|
language = Column(String, default="en")
|
||||||
|
audio_path = Column(String, nullable=False)
|
||||||
|
duration = Column(Float, nullable=False)
|
||||||
|
seed = Column(Integer)
|
||||||
|
instruct = Column(Text)
|
||||||
|
created_at = Column(DateTime)
|
||||||
|
```
|
||||||
|
|
||||||
|
## File Storage
|
||||||
|
|
||||||
|
Generated audio is stored in:
|
||||||
|
|
||||||
|
```
|
||||||
|
data/
|
||||||
|
└── generations/
|
||||||
|
└── {generation_id}.wav
|
||||||
|
```
|
||||||
|
|
||||||
|
## Core Functions
|
||||||
|
|
||||||
|
### Creating a Generation Record
|
||||||
|
|
||||||
|
After TTS generates audio, a history entry is created:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def create_generation(
|
||||||
|
profile_id: str,
|
||||||
|
text: str,
|
||||||
|
language: str,
|
||||||
|
audio_path: str,
|
||||||
|
duration: float,
|
||||||
|
seed: Optional[int],
|
||||||
|
db: Session,
|
||||||
|
instruct: Optional[str] = None,
|
||||||
|
) -> GenerationResponse:
|
||||||
|
db_generation = DBGeneration(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
profile_id=profile_id,
|
||||||
|
text=text,
|
||||||
|
language=language,
|
||||||
|
audio_path=audio_path,
|
||||||
|
duration=duration,
|
||||||
|
seed=seed,
|
||||||
|
instruct=instruct,
|
||||||
|
created_at=datetime.utcnow(),
|
||||||
|
)
|
||||||
|
|
||||||
|
db.add(db_generation)
|
||||||
|
db.commit()
|
||||||
|
|
||||||
|
return GenerationResponse.model_validate(db_generation)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Listing Generations
|
||||||
|
|
||||||
|
Supports filtering and pagination:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def list_generations(
|
||||||
|
query: HistoryQuery,
|
||||||
|
db: Session,
|
||||||
|
) -> HistoryListResponse:
|
||||||
|
# Build query with profile name join
|
||||||
|
q = db.query(
|
||||||
|
DBGeneration,
|
||||||
|
DBVoiceProfile.name.label('profile_name')
|
||||||
|
).join(
|
||||||
|
DBVoiceProfile,
|
||||||
|
DBGeneration.profile_id == DBVoiceProfile.id
|
||||||
|
)
|
||||||
|
|
||||||
|
# Apply filters
|
||||||
|
if query.profile_id:
|
||||||
|
q = q.filter(DBGeneration.profile_id == query.profile_id)
|
||||||
|
|
||||||
|
if query.search:
|
||||||
|
q = q.filter(DBGeneration.text.like(f"%{query.search}%"))
|
||||||
|
|
||||||
|
# Order and paginate
|
||||||
|
total = q.count()
|
||||||
|
q = q.order_by(DBGeneration.created_at.desc())
|
||||||
|
q = q.offset(query.offset).limit(query.limit)
|
||||||
|
|
||||||
|
return HistoryListResponse(items=results, total=total)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Getting Statistics
|
||||||
|
|
||||||
|
Aggregate statistics for the dashboard:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def get_generation_stats(db: Session) -> dict:
|
||||||
|
total = db.query(func.count(DBGeneration.id)).scalar()
|
||||||
|
total_duration = db.query(func.sum(DBGeneration.duration)).scalar()
|
||||||
|
|
||||||
|
by_profile = db.query(
|
||||||
|
DBGeneration.profile_id,
|
||||||
|
func.count(DBGeneration.id).label('count')
|
||||||
|
).group_by(DBGeneration.profile_id).all()
|
||||||
|
|
||||||
|
return {
|
||||||
|
"total_generations": total,
|
||||||
|
"total_duration_seconds": total_duration,
|
||||||
|
"generations_by_profile": {
|
||||||
|
profile_id: count for profile_id, count in by_profile
|
||||||
|
},
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Deletion
|
||||||
|
|
||||||
|
Deleting a generation removes both the database record and audio file:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def delete_generation(generation_id: str, db: Session) -> bool:
|
||||||
|
generation = db.query(DBGeneration).filter_by(id=generation_id).first()
|
||||||
|
if not generation:
|
||||||
|
return False
|
||||||
|
|
||||||
|
# Delete audio file
|
||||||
|
audio_path = Path(generation.audio_path)
|
||||||
|
if audio_path.exists():
|
||||||
|
audio_path.unlink()
|
||||||
|
|
||||||
|
# Delete database record
|
||||||
|
db.delete(generation)
|
||||||
|
db.commit()
|
||||||
|
|
||||||
|
return True
|
||||||
|
```
|
||||||
|
|
||||||
|
### Cascade Delete
|
||||||
|
|
||||||
|
When deleting a profile, all its generations are also deleted:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def delete_generations_by_profile(profile_id: str, db: Session) -> int:
|
||||||
|
generations = db.query(DBGeneration).filter_by(profile_id=profile_id).all()
|
||||||
|
|
||||||
|
for generation in generations:
|
||||||
|
Path(generation.audio_path).unlink(missing_ok=True)
|
||||||
|
db.delete(generation)
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
return len(generations)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Export/Import
|
||||||
|
|
||||||
|
### Exporting a Generation
|
||||||
|
|
||||||
|
Generations can be exported as ZIP archives:
|
||||||
|
|
||||||
|
```
|
||||||
|
generation_export.zip
|
||||||
|
├── generation.json # Metadata
|
||||||
|
└── audio.wav # Audio file
|
||||||
|
```
|
||||||
|
|
||||||
|
### Importing a Generation
|
||||||
|
|
||||||
|
The import process:
|
||||||
|
|
||||||
|
1. Extract ZIP archive
|
||||||
|
2. Validate metadata and audio
|
||||||
|
3. Create new generation ID
|
||||||
|
4. Copy audio to generations directory
|
||||||
|
5. Create database record
|
||||||
|
|
||||||
|
## API Endpoints
|
||||||
|
|
||||||
|
| Method | Endpoint | Description |
|
||||||
|
|--------|----------|-------------|
|
||||||
|
| GET | `/history` | List generations with filters |
|
||||||
|
| GET | `/history/stats` | Get aggregate statistics |
|
||||||
|
| GET | `/history/{id}` | Get generation by ID |
|
||||||
|
| DELETE | `/history/{id}` | Delete generation |
|
||||||
|
| GET | `/history/{id}/export` | Export as ZIP |
|
||||||
|
| GET | `/history/{id}/export-audio` | Export audio only |
|
||||||
|
| POST | `/history/import` | Import from ZIP |
|
||||||
|
|
||||||
|
### Query Parameters
|
||||||
|
|
||||||
|
```
|
||||||
|
GET /history?profile_id=uuid&search=hello&limit=50&offset=0
|
||||||
|
```
|
||||||
|
|
||||||
|
| Parameter | Type | Default | Description |
|
||||||
|
|-----------|------|---------|-------------|
|
||||||
|
| `profile_id` | string | null | Filter by profile |
|
||||||
|
| `search` | string | null | Search in text |
|
||||||
|
| `limit` | int | 50 | Results per page |
|
||||||
|
| `offset` | int | 0 | Pagination offset |
|
||||||
|
|
||||||
|
### Response Schema
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"items": [
|
||||||
|
{
|
||||||
|
"id": "uuid",
|
||||||
|
"profile_id": "uuid",
|
||||||
|
"profile_name": "My Voice",
|
||||||
|
"text": "Hello world",
|
||||||
|
"language": "en",
|
||||||
|
"audio_path": "/path/to/audio.wav",
|
||||||
|
"duration": 1.5,
|
||||||
|
"seed": 42,
|
||||||
|
"instruct": null,
|
||||||
|
"created_at": "2024-01-15T10:30:00Z"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"total": 150
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Usage in Stories
|
||||||
|
|
||||||
|
Generations can be added to stories for multi-voice narratives. The story system references generations by ID:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class StoryItem(Base):
|
||||||
|
generation_id = Column(String, ForeignKey("generations.id"))
|
||||||
|
```
|
||||||
|
|
||||||
|
This allows the same generation to be reused across multiple stories without duplicating audio files.
|
||||||
|
|
||||||
|
## Storage Considerations
|
||||||
|
|
||||||
|
### Disk Usage
|
||||||
|
|
||||||
|
Each generation creates a WAV file. For a 10-second clip at 24kHz:
|
||||||
|
- ~480KB per file (mono, 16-bit)
|
||||||
|
|
||||||
|
### Cleanup Strategy
|
||||||
|
|
||||||
|
Consider implementing:
|
||||||
|
- Automatic cleanup of old generations
|
||||||
|
- Storage quota per profile
|
||||||
|
- Compression for archival
|
||||||
@@ -0,0 +1,341 @@
|
|||||||
|
---
|
||||||
|
title: "Model Management"
|
||||||
|
description: "How model downloading, loading, and status tracking works in Voicebox"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Voicebox manages two types of models:
|
||||||
|
|
||||||
|
**TTS Models:** Qwen3-TTS for voice cloning (0.6B and 1.7B variants).
|
||||||
|
|
||||||
|
**ASR Models:** Whisper for transcription (tiny through large).
|
||||||
|
|
||||||
|
Models are downloaded from HuggingFace Hub on first use and cached locally.
|
||||||
|
|
||||||
|
## Available Models
|
||||||
|
|
||||||
|
### TTS Models
|
||||||
|
|
||||||
|
| Model | HuggingFace ID | Size | VRAM |
|
||||||
|
|-------|----------------|------|------|
|
||||||
|
| 0.6B | `Qwen/Qwen3-TTS-12Hz-0.6B-Base` | ~1.2GB | ~2GB |
|
||||||
|
| 1.7B | `Qwen/Qwen3-TTS-12Hz-1.7B-Base` | ~3.4GB | ~6GB |
|
||||||
|
|
||||||
|
### Whisper Models
|
||||||
|
|
||||||
|
| Model | HuggingFace ID | Size | VRAM |
|
||||||
|
|-------|----------------|------|------|
|
||||||
|
| tiny | `openai/whisper-tiny` | ~150MB | ~1GB |
|
||||||
|
| base | `openai/whisper-base` | ~300MB | ~1GB |
|
||||||
|
| small | `openai/whisper-small` | ~500MB | ~2GB |
|
||||||
|
| medium | `openai/whisper-medium` | ~1.5GB | ~5GB |
|
||||||
|
| large | `openai/whisper-large` | ~3GB | ~10GB |
|
||||||
|
|
||||||
|
## Model Storage
|
||||||
|
|
||||||
|
Models are cached in the HuggingFace cache directory:
|
||||||
|
|
||||||
|
```
|
||||||
|
~/.cache/huggingface/hub/
|
||||||
|
├── models--Qwen--Qwen3-TTS-12Hz-1.7B-Base/
|
||||||
|
├── models--Qwen--Qwen3-TTS-12Hz-0.6B-Base/
|
||||||
|
├── models--openai--whisper-base/
|
||||||
|
└── ...
|
||||||
|
```
|
||||||
|
|
||||||
|
## Progress Tracking
|
||||||
|
|
||||||
|
### Progress Manager
|
||||||
|
|
||||||
|
Tracks download progress across all models:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class ProgressManager:
|
||||||
|
def __init__(self):
|
||||||
|
self._progress = {} # model_name -> progress_info
|
||||||
|
|
||||||
|
def update_progress(
|
||||||
|
self,
|
||||||
|
model_name: str,
|
||||||
|
current: int,
|
||||||
|
total: int,
|
||||||
|
filename: str,
|
||||||
|
status: str,
|
||||||
|
):
|
||||||
|
self._progress[model_name] = {
|
||||||
|
"current": current,
|
||||||
|
"total": total,
|
||||||
|
"filename": filename,
|
||||||
|
"status": status, # downloading, complete, error
|
||||||
|
"updated_at": datetime.utcnow(),
|
||||||
|
}
|
||||||
|
|
||||||
|
def get_progress(self, model_name: str) -> Optional[dict]:
|
||||||
|
return self._progress.get(model_name)
|
||||||
|
```
|
||||||
|
|
||||||
|
### HuggingFace Progress Callback
|
||||||
|
|
||||||
|
Hooks into HuggingFace's download system:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class HFProgressTracker:
|
||||||
|
def __init__(self, callback):
|
||||||
|
self.callback = callback
|
||||||
|
|
||||||
|
@contextmanager
|
||||||
|
def patch_download(self):
|
||||||
|
"""Context manager to intercept HF downloads."""
|
||||||
|
original_download = hf_hub_download
|
||||||
|
|
||||||
|
def patched_download(*args, **kwargs):
|
||||||
|
# Intercept progress
|
||||||
|
result = original_download(*args, **kwargs)
|
||||||
|
self.callback(progress_info)
|
||||||
|
return result
|
||||||
|
|
||||||
|
# Apply patch
|
||||||
|
with patch('huggingface_hub.hf_hub_download', patched_download):
|
||||||
|
yield
|
||||||
|
```
|
||||||
|
|
||||||
|
### Server-Sent Events (SSE)
|
||||||
|
|
||||||
|
Progress is streamed to the frontend:
|
||||||
|
|
||||||
|
```python
|
||||||
|
@app.get("/models/progress/{model_name}")
|
||||||
|
async def get_model_progress(model_name: str):
|
||||||
|
async def event_generator():
|
||||||
|
while True:
|
||||||
|
progress = progress_manager.get_progress(model_name)
|
||||||
|
if progress:
|
||||||
|
yield f"data: {json.dumps(progress)}\n\n"
|
||||||
|
|
||||||
|
if progress and progress["status"] in ["complete", "error"]:
|
||||||
|
break
|
||||||
|
|
||||||
|
await asyncio.sleep(0.5)
|
||||||
|
|
||||||
|
return StreamingResponse(
|
||||||
|
event_generator(),
|
||||||
|
media_type="text/event-stream"
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Task Manager
|
||||||
|
|
||||||
|
Tracks active downloads and generations:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class TaskManager:
|
||||||
|
def __init__(self):
|
||||||
|
self._active_downloads = {}
|
||||||
|
self._active_generations = {}
|
||||||
|
|
||||||
|
def start_download(self, model_name: str):
|
||||||
|
self._active_downloads[model_name] = {
|
||||||
|
"status": "downloading",
|
||||||
|
"started_at": datetime.utcnow(),
|
||||||
|
}
|
||||||
|
|
||||||
|
def complete_download(self, model_name: str):
|
||||||
|
if model_name in self._active_downloads:
|
||||||
|
del self._active_downloads[model_name]
|
||||||
|
|
||||||
|
def get_active_tasks(self) -> dict:
|
||||||
|
return {
|
||||||
|
"downloads": list(self._active_downloads.values()),
|
||||||
|
"generations": list(self._active_generations.values()),
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Model Status
|
||||||
|
|
||||||
|
Check which models are downloaded and loaded:
|
||||||
|
|
||||||
|
```python
|
||||||
|
@app.get("/models/status")
|
||||||
|
async def get_model_status() -> ModelStatusListResponse:
|
||||||
|
models = []
|
||||||
|
|
||||||
|
# Check TTS models
|
||||||
|
for size, hf_id in [("1.7B", "Qwen/Qwen3-TTS-12Hz-1.7B-Base"), ...]:
|
||||||
|
downloaded = is_model_downloaded(hf_id)
|
||||||
|
loaded = tts_model._current_model_size == size
|
||||||
|
|
||||||
|
models.append(ModelStatus(
|
||||||
|
model_name=f"qwen-tts-{size}",
|
||||||
|
display_name=f"Qwen3-TTS {size}",
|
||||||
|
downloaded=downloaded,
|
||||||
|
size_mb=get_model_size_mb(hf_id),
|
||||||
|
loaded=loaded,
|
||||||
|
))
|
||||||
|
|
||||||
|
# Check Whisper models
|
||||||
|
for size in ["tiny", "base", "small", "medium", "large"]:
|
||||||
|
hf_id = f"openai/whisper-{size}"
|
||||||
|
downloaded = is_model_downloaded(hf_id)
|
||||||
|
|
||||||
|
models.append(ModelStatus(
|
||||||
|
model_name=f"whisper-{size}",
|
||||||
|
display_name=f"Whisper {size}",
|
||||||
|
downloaded=downloaded,
|
||||||
|
size_mb=get_model_size_mb(hf_id),
|
||||||
|
loaded=False, # Whisper is loaded on-demand
|
||||||
|
))
|
||||||
|
|
||||||
|
return ModelStatusListResponse(models=models)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Manual Model Operations
|
||||||
|
|
||||||
|
### Load Model
|
||||||
|
|
||||||
|
```python
|
||||||
|
@app.post("/models/load")
|
||||||
|
async def load_model(model_size: str = "1.7B"):
|
||||||
|
tts_model = get_tts_model()
|
||||||
|
await tts_model.load_model_async(model_size)
|
||||||
|
return {"status": "loaded", "model_size": model_size}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Unload Model
|
||||||
|
|
||||||
|
```python
|
||||||
|
@app.post("/models/unload")
|
||||||
|
async def unload_model():
|
||||||
|
tts_model = get_tts_model()
|
||||||
|
tts_model.unload_model()
|
||||||
|
return {"status": "unloaded"}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Trigger Download
|
||||||
|
|
||||||
|
```python
|
||||||
|
@app.post("/models/download")
|
||||||
|
async def trigger_model_download(request: ModelDownloadRequest):
|
||||||
|
# This triggers the download in background
|
||||||
|
# Progress is tracked via /models/progress/{model_name}
|
||||||
|
|
||||||
|
if request.model_name.startswith("qwen-tts"):
|
||||||
|
size = request.model_name.split("-")[-1]
|
||||||
|
asyncio.create_task(download_tts_model(size))
|
||||||
|
elif request.model_name.startswith("whisper"):
|
||||||
|
size = request.model_name.split("-")[-1]
|
||||||
|
asyncio.create_task(download_whisper_model(size))
|
||||||
|
|
||||||
|
return {"status": "downloading"}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Delete Model
|
||||||
|
|
||||||
|
```python
|
||||||
|
@app.delete("/models/{model_name}")
|
||||||
|
async def delete_model(model_name: str):
|
||||||
|
# Find and delete from HuggingFace cache
|
||||||
|
cache_dir = Path.home() / ".cache" / "huggingface" / "hub"
|
||||||
|
|
||||||
|
model_dirs = list(cache_dir.glob(f"models--*--{model_name}*"))
|
||||||
|
for model_dir in model_dirs:
|
||||||
|
shutil.rmtree(model_dir)
|
||||||
|
|
||||||
|
return {"status": "deleted"}
|
||||||
|
```
|
||||||
|
|
||||||
|
## API Endpoints
|
||||||
|
|
||||||
|
| Method | Endpoint | Description |
|
||||||
|
|--------|----------|-------------|
|
||||||
|
| GET | `/models/status` | Get status of all models |
|
||||||
|
| POST | `/models/load` | Load TTS model |
|
||||||
|
| POST | `/models/unload` | Unload TTS model |
|
||||||
|
| POST | `/models/download` | Trigger model download |
|
||||||
|
| GET | `/models/progress/{name}` | Stream download progress (SSE) |
|
||||||
|
| DELETE | `/models/{name}` | Delete downloaded model |
|
||||||
|
| GET | `/tasks/active` | Get active downloads/generations |
|
||||||
|
|
||||||
|
## Response Schemas
|
||||||
|
|
||||||
|
### ModelStatus
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"model_name": "qwen-tts-1.7B",
|
||||||
|
"display_name": "Qwen3-TTS 1.7B",
|
||||||
|
"downloaded": true,
|
||||||
|
"size_mb": 3400,
|
||||||
|
"loaded": true
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### ActiveTasksResponse
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"downloads": [
|
||||||
|
{
|
||||||
|
"model_name": "whisper-medium",
|
||||||
|
"status": "downloading",
|
||||||
|
"started_at": "2024-01-15T10:30:00Z"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"generations": [
|
||||||
|
{
|
||||||
|
"task_id": "uuid",
|
||||||
|
"profile_id": "uuid",
|
||||||
|
"text_preview": "Hello world...",
|
||||||
|
"started_at": "2024-01-15T10:30:00Z"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Frontend Integration
|
||||||
|
|
||||||
|
### Progress Display
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Subscribe to download progress via SSE
|
||||||
|
const eventSource = new EventSource(`/models/progress/${modelName}`);
|
||||||
|
|
||||||
|
eventSource.onmessage = (event) => {
|
||||||
|
const progress = JSON.parse(event.data);
|
||||||
|
updateProgressBar(progress.current / progress.total);
|
||||||
|
|
||||||
|
if (progress.status === 'complete') {
|
||||||
|
eventSource.close();
|
||||||
|
}
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
### Model Status UI
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Fetch model status
|
||||||
|
const { data: models } = useQuery({
|
||||||
|
queryKey: ['models', 'status'],
|
||||||
|
queryFn: () => api.getModelStatus(),
|
||||||
|
});
|
||||||
|
|
||||||
|
// Display download/load buttons based on status
|
||||||
|
models.map(model => (
|
||||||
|
<ModelCard
|
||||||
|
name={model.display_name}
|
||||||
|
downloaded={model.downloaded}
|
||||||
|
loaded={model.loaded}
|
||||||
|
onDownload={() => triggerDownload(model.model_name)}
|
||||||
|
onLoad={() => loadModel(model.model_name)}
|
||||||
|
/>
|
||||||
|
));
|
||||||
|
```
|
||||||
|
|
||||||
|
## Error Handling
|
||||||
|
|
||||||
|
| Error | Cause | Solution |
|
||||||
|
|-------|-------|----------|
|
||||||
|
| Download failed | Network issue | Retry download |
|
||||||
|
| OOM on load | Model too large | Use smaller model |
|
||||||
|
| Model not found | Cache corrupted | Re-download |
|
||||||
|
| Slow download | HF rate limit | Wait and retry |
|
||||||
@@ -0,0 +1,239 @@
|
|||||||
|
---
|
||||||
|
title: "Development Setup"
|
||||||
|
description: "Set up your local development environment for Voicebox"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Prerequisites
|
||||||
|
|
||||||
|
Before you begin, ensure you have the following installed:
|
||||||
|
|
||||||
|
<CardGroup cols={3}>
|
||||||
|
<Card title="Bun" icon="package">
|
||||||
|
[Download Bun](https://bun.sh)
|
||||||
|
```bash
|
||||||
|
curl -fsSL https://bun.sh/install | bash
|
||||||
|
```
|
||||||
|
</Card>
|
||||||
|
<Card title="Python 3.11+" icon="python">
|
||||||
|
[Download Python](https://python.org)
|
||||||
|
```bash
|
||||||
|
python --version
|
||||||
|
```
|
||||||
|
</Card>
|
||||||
|
<Card title="Rust" icon="rust">
|
||||||
|
[Install Rust](https://rustup.rs)
|
||||||
|
```bash
|
||||||
|
rustc --version
|
||||||
|
```
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
|
|
||||||
|
## Clone the Repository
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git clone https://github.com/jamiepine/voicebox.git
|
||||||
|
cd voicebox
|
||||||
|
```
|
||||||
|
|
||||||
|
## Quick Setup (Recommended)
|
||||||
|
|
||||||
|
The easiest way to get started is using the Makefile:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Setup everything
|
||||||
|
make setup
|
||||||
|
|
||||||
|
# Start development
|
||||||
|
make dev
|
||||||
|
```
|
||||||
|
|
||||||
|
<Note>
|
||||||
|
The Makefile is available on macOS and Linux. Windows users should follow the manual setup below.
|
||||||
|
</Note>
|
||||||
|
|
||||||
|
## Manual Setup
|
||||||
|
|
||||||
|
### 1. Install JavaScript Dependencies
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bun install
|
||||||
|
```
|
||||||
|
|
||||||
|
This installs dependencies for:
|
||||||
|
- `app/` - Shared React frontend
|
||||||
|
- `tauri/` - Tauri desktop wrapper
|
||||||
|
- `web/` - Web deployment wrapper
|
||||||
|
|
||||||
|
### 2. Set Up Python Backend
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd backend
|
||||||
|
|
||||||
|
# Create virtual environment
|
||||||
|
python -m venv venv
|
||||||
|
|
||||||
|
# Activate virtual environment
|
||||||
|
source venv/bin/activate # macOS/Linux
|
||||||
|
# or
|
||||||
|
venv\Scripts\activate # Windows
|
||||||
|
|
||||||
|
# Install Python dependencies
|
||||||
|
pip install -r requirements.txt
|
||||||
|
|
||||||
|
# Install MLX dependencies (Apple Silicon only - for faster inference)
|
||||||
|
# On Apple Silicon, this enables native Metal acceleration
|
||||||
|
if [[ $(uname -m) == "arm64" ]]; then
|
||||||
|
pip install -r requirements-mlx.txt
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Install Qwen3-TTS
|
||||||
|
pip install git+https://github.com/QwenLM/Qwen3-TTS.git
|
||||||
|
```
|
||||||
|
|
||||||
|
## Running in Development
|
||||||
|
|
||||||
|
Development requires **two terminals**: one for the Python backend, one for the Tauri app.
|
||||||
|
|
||||||
|
<Tabs>
|
||||||
|
<Tab title="Terminal 1: Backend">
|
||||||
|
Start the Python server first:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd backend
|
||||||
|
source venv/bin/activate # Activate venv
|
||||||
|
bun run dev:server
|
||||||
|
```
|
||||||
|
|
||||||
|
Or manually:
|
||||||
|
```bash
|
||||||
|
uvicorn main:app --reload --port 17493
|
||||||
|
```
|
||||||
|
|
||||||
|
Backend will be available at `http://localhost:17493`
|
||||||
|
</Tab>
|
||||||
|
|
||||||
|
<Tab title="Terminal 2: Desktop App">
|
||||||
|
Then start the Tauri app:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bun run dev
|
||||||
|
```
|
||||||
|
|
||||||
|
This will:
|
||||||
|
- Create a placeholder sidecar binary
|
||||||
|
- Start Vite dev server on port 5173
|
||||||
|
- Launch Tauri window
|
||||||
|
- Enable hot reload
|
||||||
|
</Tab>
|
||||||
|
</Tabs>
|
||||||
|
|
||||||
|
<Info>
|
||||||
|
In dev mode, the app connects to your manually-started Python server. The bundled server binary is only used in production builds.
|
||||||
|
</Info>
|
||||||
|
|
||||||
|
### Optional: Web App
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bun run dev:web
|
||||||
|
```
|
||||||
|
|
||||||
|
Web app will be available at `http://localhost:5174`
|
||||||
|
|
||||||
|
## Model Downloads
|
||||||
|
|
||||||
|
Models are automatically downloaded from HuggingFace Hub on first use:
|
||||||
|
|
||||||
|
- **Whisper** (transcription): Auto-downloads on first transcription
|
||||||
|
- **Qwen3-TTS** (voice cloning): Auto-downloads on first generation (~2-4GB)
|
||||||
|
|
||||||
|
<Warning>
|
||||||
|
First-time usage will be slower due to model downloads, but subsequent runs will use cached models.
|
||||||
|
</Warning>
|
||||||
|
|
||||||
|
## Project Structure
|
||||||
|
|
||||||
|
```
|
||||||
|
voicebox/
|
||||||
|
├── app/ # Shared React frontend
|
||||||
|
│ └── src/
|
||||||
|
│ ├── components/ # UI components
|
||||||
|
│ ├── lib/ # Utilities and API client
|
||||||
|
│ └── hooks/ # React hooks
|
||||||
|
├── backend/ # Python FastAPI server
|
||||||
|
│ ├── main.py # API routes
|
||||||
|
│ ├── tts.py # Voice synthesis
|
||||||
|
│ └── database.py # SQLite operations
|
||||||
|
├── tauri/ # Desktop app wrapper
|
||||||
|
│ └── src-tauri/ # Rust backend
|
||||||
|
├── web/ # Web deployment
|
||||||
|
├── landing/ # Marketing website
|
||||||
|
└── scripts/ # Build & release scripts
|
||||||
|
```
|
||||||
|
|
||||||
|
## Available Make Commands
|
||||||
|
|
||||||
|
Run `make help` to see all available commands:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make setup # Install all dependencies
|
||||||
|
make dev # Start development servers
|
||||||
|
make dev-web # Start web development server
|
||||||
|
make build # Build desktop app
|
||||||
|
make build-web # Build web app
|
||||||
|
make clean # Clean build artifacts
|
||||||
|
make test # Run tests
|
||||||
|
```
|
||||||
|
|
||||||
|
## Generate OpenAPI Client
|
||||||
|
|
||||||
|
After starting the backend server, generate the TypeScript API client:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/generate-api.sh
|
||||||
|
# or
|
||||||
|
bun run generate:api
|
||||||
|
```
|
||||||
|
|
||||||
|
This downloads the OpenAPI schema and generates the TypeScript client in `app/src/lib/api/`
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
<CardGroup cols={2}>
|
||||||
|
<Card title="Architecture" icon="diagram-project" href="/development/architecture">
|
||||||
|
Understand the system architecture
|
||||||
|
</Card>
|
||||||
|
<Card title="Contributing" icon="code-pull-request" href="/development/contributing">
|
||||||
|
Read the contribution guidelines
|
||||||
|
</Card>
|
||||||
|
<Card title="Building" icon="hammer" href="/development/building">
|
||||||
|
Learn how to build production releases
|
||||||
|
</Card>
|
||||||
|
<Card title="API Reference" icon="code" href="/api/overview">
|
||||||
|
Explore the REST API
|
||||||
|
</Card>
|
||||||
|
</CardGroup>
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
<AccordionGroup>
|
||||||
|
<Accordion title="Backend won't start">
|
||||||
|
- Check Python version (must be 3.11+)
|
||||||
|
- Ensure virtual environment is activated
|
||||||
|
- Verify all dependencies are installed: `pip install -r requirements.txt`
|
||||||
|
- Check if port 17493 is available
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="Tauri build fails">
|
||||||
|
- Ensure Rust is installed: `rustc --version`
|
||||||
|
- Clean the build: `cd tauri/src-tauri && cargo clean`
|
||||||
|
- Try rebuilding: `bun run dev`
|
||||||
|
</Accordion>
|
||||||
|
|
||||||
|
<Accordion title="OpenAPI client generation fails">
|
||||||
|
- Ensure backend is running: `curl http://localhost:17493/openapi.json`
|
||||||
|
- Check network connectivity
|
||||||
|
- Verify the backend is accessible at localhost:17493
|
||||||
|
</Accordion>
|
||||||
|
</AccordionGroup>
|
||||||
|
|
||||||
|
See the full [Troubleshooting Guide](/guides/troubleshooting) for more issues and solutions.
|
||||||
@@ -0,0 +1,320 @@
|
|||||||
|
---
|
||||||
|
title: "Stories & Timeline"
|
||||||
|
description: "How the multi-voice timeline editor works in Voicebox"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Stories allow users to arrange multiple voice generations on a timeline to create multi-voice narratives. The system supports tracks, trimming, splitting, and audio mixing.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
**Story:** A container that holds story items with metadata.
|
||||||
|
|
||||||
|
**Story Item:** Links a generation to a story with timeline position, track, and trim data.
|
||||||
|
|
||||||
|
**Export:** Combines all items into a single mixed audio file.
|
||||||
|
|
||||||
|
## Data Model
|
||||||
|
|
||||||
|
### Story Table
|
||||||
|
|
||||||
|
```python
|
||||||
|
class Story(Base):
|
||||||
|
__tablename__ = "stories"
|
||||||
|
|
||||||
|
id = Column(String, primary_key=True)
|
||||||
|
name = Column(String, nullable=False)
|
||||||
|
description = Column(Text)
|
||||||
|
created_at = Column(DateTime)
|
||||||
|
updated_at = Column(DateTime)
|
||||||
|
```
|
||||||
|
|
||||||
|
### StoryItem Table
|
||||||
|
|
||||||
|
```python
|
||||||
|
class StoryItem(Base):
|
||||||
|
__tablename__ = "story_items"
|
||||||
|
|
||||||
|
id = Column(String, primary_key=True)
|
||||||
|
story_id = Column(String, ForeignKey("stories.id"))
|
||||||
|
generation_id = Column(String, ForeignKey("generations.id"))
|
||||||
|
start_time_ms = Column(Integer, default=0) # Timeline position
|
||||||
|
track = Column(Integer, default=0) # Track number
|
||||||
|
trim_start_ms = Column(Integer, default=0) # Trim from start
|
||||||
|
trim_end_ms = Column(Integer, default=0) # Trim from end
|
||||||
|
created_at = Column(DateTime)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Timeline Concepts
|
||||||
|
|
||||||
|
### Start Time
|
||||||
|
|
||||||
|
`start_time_ms` defines when an item begins on the timeline:
|
||||||
|
|
||||||
|
```
|
||||||
|
Timeline (ms): 0----1000----2000----3000----4000
|
||||||
|
Item 1: [======]
|
||||||
|
Item 2: [==========]
|
||||||
|
Item 3: [====]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Tracks
|
||||||
|
|
||||||
|
Multiple tracks allow overlapping audio:
|
||||||
|
|
||||||
|
```
|
||||||
|
Track 0: [Item 1] [Item 3]
|
||||||
|
Track 1: [Item 2]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Trimming
|
||||||
|
|
||||||
|
Trim values cut audio from the start or end without destroying the original:
|
||||||
|
|
||||||
|
```
|
||||||
|
Original: [=========AUDIO=========]
|
||||||
|
trim_start: ^^
|
||||||
|
trim_end: ^^
|
||||||
|
Result: [=====AUDIO=====]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Core Operations
|
||||||
|
|
||||||
|
### Adding Items
|
||||||
|
|
||||||
|
When adding a generation to a story:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def add_item_to_story(
|
||||||
|
story_id: str,
|
||||||
|
data: StoryItemCreate,
|
||||||
|
db: Session,
|
||||||
|
) -> StoryItemDetail:
|
||||||
|
# Calculate start time if not provided
|
||||||
|
if data.start_time_ms is None:
|
||||||
|
# Find the end of all existing items
|
||||||
|
existing_items = get_items_with_durations(story_id, db)
|
||||||
|
max_end_time_ms = max(
|
||||||
|
item.start_time_ms + int(gen.duration * 1000)
|
||||||
|
for item, gen in existing_items
|
||||||
|
)
|
||||||
|
start_time_ms = max_end_time_ms + 200 # 200ms gap
|
||||||
|
|
||||||
|
# Create the item
|
||||||
|
item = DBStoryItem(
|
||||||
|
id=str(uuid.uuid4()),
|
||||||
|
story_id=story_id,
|
||||||
|
generation_id=data.generation_id,
|
||||||
|
start_time_ms=start_time_ms,
|
||||||
|
track=data.track or 0,
|
||||||
|
)
|
||||||
|
db.add(item)
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Moving Items
|
||||||
|
|
||||||
|
Update position and/or track:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def move_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: StoryItemMove,
|
||||||
|
db: Session,
|
||||||
|
) -> StoryItemDetail:
|
||||||
|
item = get_item(story_id, item_id, db)
|
||||||
|
|
||||||
|
item.start_time_ms = data.start_time_ms
|
||||||
|
item.track = data.track
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Trimming Items
|
||||||
|
|
||||||
|
Non-destructive trimming:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def trim_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: StoryItemTrim,
|
||||||
|
db: Session,
|
||||||
|
) -> StoryItemDetail:
|
||||||
|
item = get_item(story_id, item_id, db)
|
||||||
|
generation = get_generation(item.generation_id, db)
|
||||||
|
|
||||||
|
# Validate trim doesn't exceed duration
|
||||||
|
max_duration_ms = int(generation.duration * 1000)
|
||||||
|
if data.trim_start_ms + data.trim_end_ms >= max_duration_ms:
|
||||||
|
return None # Invalid trim
|
||||||
|
|
||||||
|
item.trim_start_ms = data.trim_start_ms
|
||||||
|
item.trim_end_ms = data.trim_end_ms
|
||||||
|
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Splitting Items
|
||||||
|
|
||||||
|
Split one item into two at a specific time:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def split_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
data: StoryItemSplit,
|
||||||
|
db: Session,
|
||||||
|
) -> List[StoryItemDetail]:
|
||||||
|
item = get_item(story_id, item_id, db)
|
||||||
|
generation = get_generation(item.generation_id, db)
|
||||||
|
|
||||||
|
# Calculate split point
|
||||||
|
current_trim_start = item.trim_start_ms
|
||||||
|
current_trim_end = item.trim_end_ms
|
||||||
|
original_duration_ms = int(generation.duration * 1000)
|
||||||
|
absolute_split_ms = current_trim_start + data.split_time_ms
|
||||||
|
|
||||||
|
# Update original: trim from end
|
||||||
|
item.trim_end_ms = original_duration_ms - absolute_split_ms
|
||||||
|
|
||||||
|
# Create new item: trim from start
|
||||||
|
new_item = DBStoryItem(
|
||||||
|
generation_id=item.generation_id, # Same generation
|
||||||
|
start_time_ms=item.start_time_ms + data.split_time_ms,
|
||||||
|
track=item.track,
|
||||||
|
trim_start_ms=absolute_split_ms,
|
||||||
|
trim_end_ms=current_trim_end,
|
||||||
|
)
|
||||||
|
|
||||||
|
db.add(new_item)
|
||||||
|
db.commit()
|
||||||
|
|
||||||
|
return [item, new_item]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Duplicating Items
|
||||||
|
|
||||||
|
Create a copy with all properties:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def duplicate_story_item(
|
||||||
|
story_id: str,
|
||||||
|
item_id: str,
|
||||||
|
db: Session,
|
||||||
|
) -> StoryItemDetail:
|
||||||
|
original = get_item(story_id, item_id, db)
|
||||||
|
generation = get_generation(original.generation_id, db)
|
||||||
|
|
||||||
|
# Calculate effective duration for positioning
|
||||||
|
effective_duration_ms = (
|
||||||
|
int(generation.duration * 1000)
|
||||||
|
- original.trim_start_ms
|
||||||
|
- original.trim_end_ms
|
||||||
|
)
|
||||||
|
|
||||||
|
# Place copy after original with 200ms gap
|
||||||
|
new_item = DBStoryItem(
|
||||||
|
generation_id=original.generation_id,
|
||||||
|
start_time_ms=original.start_time_ms + effective_duration_ms + 200,
|
||||||
|
track=original.track,
|
||||||
|
trim_start_ms=original.trim_start_ms,
|
||||||
|
trim_end_ms=original.trim_end_ms,
|
||||||
|
)
|
||||||
|
|
||||||
|
db.add(new_item)
|
||||||
|
db.commit()
|
||||||
|
```
|
||||||
|
|
||||||
|
## Audio Export
|
||||||
|
|
||||||
|
### Mixing Algorithm
|
||||||
|
|
||||||
|
The export function mixes all items into a single audio file:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def export_story_audio(story_id: str, db: Session) -> bytes:
|
||||||
|
items = get_all_items_with_generations(story_id, db)
|
||||||
|
|
||||||
|
# Calculate total duration
|
||||||
|
max_end_time_ms = max(
|
||||||
|
data['start_time_ms'] + data['duration_ms']
|
||||||
|
for data in audio_data
|
||||||
|
)
|
||||||
|
|
||||||
|
# Create output buffer
|
||||||
|
total_samples = int((max_end_time_ms / 1000.0) * sample_rate)
|
||||||
|
final_audio = np.zeros(total_samples, dtype=np.float32)
|
||||||
|
|
||||||
|
# Mix each item at its position
|
||||||
|
for data in audio_data:
|
||||||
|
audio = data['audio']
|
||||||
|
start_sample = int((data['start_time_ms'] / 1000.0) * sample_rate)
|
||||||
|
|
||||||
|
# Apply trim
|
||||||
|
trimmed_audio = audio[trim_start_sample:len(audio) - trim_end_sample]
|
||||||
|
|
||||||
|
# Add to buffer (overlapping items sum together)
|
||||||
|
final_audio[start_sample:start_sample + len(trimmed_audio)] += trimmed_audio
|
||||||
|
|
||||||
|
# Normalize to prevent clipping
|
||||||
|
max_val = np.abs(final_audio).max()
|
||||||
|
if max_val > 1.0:
|
||||||
|
final_audio = final_audio / max_val
|
||||||
|
|
||||||
|
return audio_to_bytes(final_audio, sample_rate)
|
||||||
|
```
|
||||||
|
|
||||||
|
## API Endpoints
|
||||||
|
|
||||||
|
| Method | Endpoint | Description |
|
||||||
|
|--------|----------|-------------|
|
||||||
|
| GET | `/stories` | List all stories |
|
||||||
|
| POST | `/stories` | Create a story |
|
||||||
|
| GET | `/stories/{id}` | Get story with items |
|
||||||
|
| PUT | `/stories/{id}` | Update story metadata |
|
||||||
|
| DELETE | `/stories/{id}` | Delete story |
|
||||||
|
| POST | `/stories/{id}/items` | Add item to story |
|
||||||
|
| DELETE | `/stories/{id}/items/{item_id}` | Remove item |
|
||||||
|
| PUT | `/stories/{id}/items/{item_id}/move` | Move item |
|
||||||
|
| PUT | `/stories/{id}/items/{item_id}/trim` | Trim item |
|
||||||
|
| POST | `/stories/{id}/items/{item_id}/split` | Split item |
|
||||||
|
| POST | `/stories/{id}/items/{item_id}/duplicate` | Duplicate item |
|
||||||
|
| PUT | `/stories/{id}/items/times` | Batch update times |
|
||||||
|
| PUT | `/stories/{id}/items/reorder` | Reorder items |
|
||||||
|
| GET | `/stories/{id}/export-audio` | Export mixed audio |
|
||||||
|
|
||||||
|
## Response Schemas
|
||||||
|
|
||||||
|
### StoryItemDetail
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "item_uuid",
|
||||||
|
"story_id": "story_uuid",
|
||||||
|
"generation_id": "generation_uuid",
|
||||||
|
"start_time_ms": 1500,
|
||||||
|
"track": 0,
|
||||||
|
"trim_start_ms": 200,
|
||||||
|
"trim_end_ms": 100,
|
||||||
|
"profile_id": "profile_uuid",
|
||||||
|
"profile_name": "Narrator",
|
||||||
|
"text": "Hello world",
|
||||||
|
"audio_path": "/path/to/audio.wav",
|
||||||
|
"duration": 2.5,
|
||||||
|
"created_at": "2024-01-15T10:30:00Z"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Frontend Integration
|
||||||
|
|
||||||
|
The timeline UI needs to:
|
||||||
|
|
||||||
|
1. **Fetch story** with all items
|
||||||
|
2. **Render waveforms** for each item
|
||||||
|
3. **Handle drag/drop** to move items
|
||||||
|
4. **Handle edge drag** for trimming
|
||||||
|
5. **Sync playhead** across all tracks
|
||||||
|
6. **Export** when user clicks download
|
||||||
@@ -0,0 +1,299 @@
|
|||||||
|
---
|
||||||
|
title: "Transcription"
|
||||||
|
description: "How Whisper-based audio transcription works in Voicebox"
|
||||||
|
---
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Voicebox uses OpenAI's Whisper model for automatic speech recognition (ASR). This powers the transcription feature for creating reference text from audio recordings.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The transcription system is built around the `WhisperModel` class:
|
||||||
|
|
||||||
|
**Model Loading:** Lazy loading with HuggingFace Hub download.
|
||||||
|
|
||||||
|
**Audio Processing:** Resampling and preprocessing for Whisper.
|
||||||
|
|
||||||
|
**Inference:** Running transcription with optional language hints.
|
||||||
|
|
||||||
|
## WhisperModel Class
|
||||||
|
|
||||||
|
```python
|
||||||
|
class WhisperModel:
|
||||||
|
def __init__(self, model_size: str = "base"):
|
||||||
|
self.model = None
|
||||||
|
self.processor = None
|
||||||
|
self.model_size = model_size
|
||||||
|
self.device = self._get_device()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Model Sizes
|
||||||
|
|
||||||
|
| Size | Parameters | VRAM | Speed | Quality |
|
||||||
|
|------|------------|------|-------|---------|
|
||||||
|
| tiny | 39M | ~1GB | Fastest | Basic |
|
||||||
|
| base | 74M | ~1GB | Fast | Good |
|
||||||
|
| small | 244M | ~2GB | Medium | Better |
|
||||||
|
| medium | 769M | ~5GB | Slow | High |
|
||||||
|
| large | 1550M | ~10GB | Slowest | Best |
|
||||||
|
|
||||||
|
Default is `base` for balance of speed and quality.
|
||||||
|
|
||||||
|
## Model Loading
|
||||||
|
|
||||||
|
Models are downloaded from HuggingFace Hub:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def load_model(self, model_size: Optional[str] = None):
|
||||||
|
from transformers import WhisperProcessor, WhisperForConditionalGeneration
|
||||||
|
|
||||||
|
model_name = f"openai/whisper-{model_size}"
|
||||||
|
|
||||||
|
# Track download progress
|
||||||
|
progress_manager = get_progress_manager()
|
||||||
|
task_manager = get_task_manager()
|
||||||
|
task_manager.start_download(f"whisper-{model_size}")
|
||||||
|
|
||||||
|
# Load processor and model
|
||||||
|
with tracker.patch_download():
|
||||||
|
self.processor = WhisperProcessor.from_pretrained(model_name)
|
||||||
|
self.model = WhisperForConditionalGeneration.from_pretrained(model_name)
|
||||||
|
|
||||||
|
self.model.to(self.device)
|
||||||
|
|
||||||
|
# Mark complete
|
||||||
|
progress_manager.mark_complete(f"whisper-{model_size}")
|
||||||
|
task_manager.complete_download(f"whisper-{model_size}")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Async Loading
|
||||||
|
|
||||||
|
Like TTS, loading runs in a thread pool:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def load_model_async(self, model_size: Optional[str] = None):
|
||||||
|
if self.model is not None and self.model_size == model_size:
|
||||||
|
return
|
||||||
|
await asyncio.to_thread(self.load_model, model_size)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Transcription
|
||||||
|
|
||||||
|
### Basic Transcription
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def transcribe(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
language: Optional[str] = None,
|
||||||
|
) -> str:
|
||||||
|
await self.load_model_async()
|
||||||
|
|
||||||
|
def _transcribe_sync():
|
||||||
|
# Load and resample to 16kHz (Whisper requirement)
|
||||||
|
audio, sr = load_audio(audio_path, sample_rate=16000)
|
||||||
|
|
||||||
|
# Process audio
|
||||||
|
inputs = self.processor(
|
||||||
|
audio,
|
||||||
|
sampling_rate=16000,
|
||||||
|
return_tensors="pt",
|
||||||
|
)
|
||||||
|
inputs = inputs.to(self.device)
|
||||||
|
|
||||||
|
# Set language hint if provided
|
||||||
|
forced_decoder_ids = None
|
||||||
|
if language:
|
||||||
|
forced_decoder_ids = self.processor.get_decoder_prompt_ids(
|
||||||
|
language=language,
|
||||||
|
task="transcribe",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Generate
|
||||||
|
with torch.no_grad():
|
||||||
|
predicted_ids = self.model.generate(
|
||||||
|
inputs["input_features"],
|
||||||
|
forced_decoder_ids=forced_decoder_ids,
|
||||||
|
)
|
||||||
|
|
||||||
|
# Decode
|
||||||
|
transcription = self.processor.batch_decode(
|
||||||
|
predicted_ids,
|
||||||
|
skip_special_tokens=True,
|
||||||
|
)[0]
|
||||||
|
|
||||||
|
return transcription.strip()
|
||||||
|
|
||||||
|
return await asyncio.to_thread(_transcribe_sync)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Supported Languages
|
||||||
|
|
||||||
|
Whisper supports 99+ languages. Common ones in Voicebox:
|
||||||
|
|
||||||
|
| Code | Language |
|
||||||
|
|------|----------|
|
||||||
|
| en | English |
|
||||||
|
| zh | Chinese |
|
||||||
|
| ja | Japanese |
|
||||||
|
| ko | Korean |
|
||||||
|
| de | German |
|
||||||
|
| fr | French |
|
||||||
|
| ru | Russian |
|
||||||
|
| pt | Portuguese |
|
||||||
|
| es | Spanish |
|
||||||
|
| it | Italian |
|
||||||
|
|
||||||
|
### Language Detection
|
||||||
|
|
||||||
|
When no language is specified, Whisper auto-detects:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Without language hint - auto-detect
|
||||||
|
transcription = await whisper.transcribe(audio_path)
|
||||||
|
|
||||||
|
# With language hint - more accurate for short clips
|
||||||
|
transcription = await whisper.transcribe(audio_path, language="en")
|
||||||
|
```
|
||||||
|
|
||||||
|
## Transcription with Timestamps
|
||||||
|
|
||||||
|
For advanced use cases, word-level timestamps are available:
|
||||||
|
|
||||||
|
```python
|
||||||
|
async def transcribe_with_timestamps(
|
||||||
|
self,
|
||||||
|
audio_path: str,
|
||||||
|
language: Optional[str] = None,
|
||||||
|
) -> List[Dict[str, any]]:
|
||||||
|
await self.load_model_async()
|
||||||
|
|
||||||
|
def _transcribe_timestamps_sync():
|
||||||
|
audio, sr = load_audio(audio_path, sample_rate=16000)
|
||||||
|
inputs = self.processor(audio, sampling_rate=16000, return_tensors="pt")
|
||||||
|
|
||||||
|
with torch.no_grad():
|
||||||
|
predicted_ids = self.model.generate(
|
||||||
|
inputs["input_features"],
|
||||||
|
return_timestamps=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
# Parse timestamps
|
||||||
|
return [
|
||||||
|
{
|
||||||
|
"text": transcription,
|
||||||
|
"start": 0.0,
|
||||||
|
"end": len(audio) / sr,
|
||||||
|
}
|
||||||
|
]
|
||||||
|
|
||||||
|
return await asyncio.to_thread(_transcribe_timestamps_sync)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Memory Management
|
||||||
|
|
||||||
|
### Unloading
|
||||||
|
|
||||||
|
Free memory when not needed:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def unload_model(self):
|
||||||
|
if self.model is not None:
|
||||||
|
del self.model
|
||||||
|
del self.processor
|
||||||
|
self.model = None
|
||||||
|
self.processor = None
|
||||||
|
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Global Instance
|
||||||
|
|
||||||
|
A singleton pattern manages the model:
|
||||||
|
|
||||||
|
```python
|
||||||
|
_whisper_model: Optional[WhisperModel] = None
|
||||||
|
|
||||||
|
def get_whisper_model() -> WhisperModel:
|
||||||
|
global _whisper_model
|
||||||
|
if _whisper_model is None:
|
||||||
|
_whisper_model = WhisperModel()
|
||||||
|
return _whisper_model
|
||||||
|
```
|
||||||
|
|
||||||
|
## Audio Preprocessing
|
||||||
|
|
||||||
|
### Resampling
|
||||||
|
|
||||||
|
Whisper requires 16kHz audio:
|
||||||
|
|
||||||
|
```python
|
||||||
|
audio, sr = load_audio(audio_path, sample_rate=16000)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Format Support
|
||||||
|
|
||||||
|
The `load_audio` utility handles:
|
||||||
|
- WAV
|
||||||
|
- MP3
|
||||||
|
- FLAC
|
||||||
|
- OGG
|
||||||
|
- M4A
|
||||||
|
|
||||||
|
All formats are converted to mono 16kHz.
|
||||||
|
|
||||||
|
## API Endpoints
|
||||||
|
|
||||||
|
| Method | Endpoint | Description |
|
||||||
|
|--------|----------|-------------|
|
||||||
|
| POST | `/transcribe` | Transcribe audio file |
|
||||||
|
|
||||||
|
### Request
|
||||||
|
|
||||||
|
Multipart form data:
|
||||||
|
|
||||||
|
```
|
||||||
|
POST /transcribe
|
||||||
|
Content-Type: multipart/form-data
|
||||||
|
|
||||||
|
file: <audio_file>
|
||||||
|
language: en (optional)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Response
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"text": "Hello, this is a test transcription.",
|
||||||
|
"duration": 3.5
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Use Cases
|
||||||
|
|
||||||
|
### Reference Text for Voice Cloning
|
||||||
|
|
||||||
|
1. User records audio sample
|
||||||
|
2. Audio is sent to `/transcribe`
|
||||||
|
3. Transcription becomes `reference_text`
|
||||||
|
4. Both are added to voice profile
|
||||||
|
|
||||||
|
### Quality Tips
|
||||||
|
|
||||||
|
- Provide language hint for short audio
|
||||||
|
- Use clean audio with minimal noise
|
||||||
|
- Longer audio (>5s) improves accuracy
|
||||||
|
- Consider `small` or `medium` model for better quality
|
||||||
|
|
||||||
|
## Error Handling
|
||||||
|
|
||||||
|
Common issues:
|
||||||
|
|
||||||
|
| Error | Cause | Solution |
|
||||||
|
|-------|-------|----------|
|
||||||
|
| Model not found | First run, download failed | Retry with network |
|
||||||
|
| OOM | Model too large | Use smaller model |
|
||||||
|
| Empty result | No speech detected | Check audio has speech |
|
||||||
|
| Wrong language | Auto-detect failed | Provide language hint |
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user