feat: add CosyVoice2/3 TTS engine with instruct and voice cloning

Integrate Alibaba's CosyVoice2-0.5B and Fun-CosyVoice3-0.5B as a new
TTS engine supporting 9 languages, zero-shot voice cloning, and instruct
control (emotions, speed, volume, dialects).

The CosyVoice source is cloned at setup time into backend/vendors/ since
no PyPI package exists. A modelscope→HuggingFace shim redirects model
downloads to the public HF repos, and a lightweight pylogger shim avoids
pulling in pytorch-lightning as a transitive dependency.

Backend: cosyvoice_backend.py, __init__.py registry, models.py regex
Frontend: engine selector, language map, Zod schema, model descriptions
Infra: requirements.txt, justfile, release.yml, Dockerfile, PyInstaller
This commit is contained in:
James Pine
2026-03-17 12:31:20 -07:00
parent c9f38dd496
commit f77dd621e2
15 changed files with 494 additions and 14 deletions
+9
View File
@@ -40,6 +40,15 @@ pyloudnorm
# provides the only class TADA uses: Snake1d.)
torchaudio
# CosyVoice2/3 sub-dependencies (the cosyvoice source is cloned at
# setup time into backend/vendors/CosyVoice — no PyPI package exists)
hyperpyyaml>=1.2.0
onnxruntime>=1.18.0
openai-whisper>=20231117
tiktoken
einops
inflect
# Audio processing
librosa>=0.10.0
soundfile>=0.12.0