add HumeAI TADA TTS engine (1B English + 3B Multilingual)

Integrates HumeAI's TADA (Text-Acoustic Dual Alignment) speech-language
model as a new TTS engine. TADA uses a novel 1:1 token-audio alignment
that produces coherent speech over long sequences (700s+).

Two model variants:
- tada-1b: English-only, ~4GB, built on Llama 3.2 1B
- tada-3b-ml: 10 languages, ~8GB, built on Llama 3.2 3B

Backend uses the Encoder for voice prompt encoding with caching, and
TadaForCausalLM with flow-matching diffusion for generation. Supports
bf16 inference on CUDA, forces CPU on macOS (MPS compatibility).

Installed with --no-deps due to torch>=2.7 pin conflict; descript-audio-codec
and torchaudio added as explicit sub-dependencies.
This commit is contained in:
James Pine
2026-03-17 01:55:15 -07:00
parent 51fb320b8c
commit 4e7772a21d
13 changed files with 437 additions and 16 deletions
+44
View File
@@ -186,6 +186,50 @@ def build_server(cuda=False):
# needed by LuxTTS for text-to-phoneme conversion
"--collect-all",
"piper_phonemize",
# HumeAI TADA — speech-language model using Llama + flow matching
"--hidden-import",
"backend.backends.hume_backend",
"--hidden-import",
"tada",
"--hidden-import",
"tada.modules",
"--hidden-import",
"tada.modules.tada",
"--hidden-import",
"tada.modules.encoder",
"--hidden-import",
"tada.modules.decoder",
"--hidden-import",
"tada.modules.aligner",
"--hidden-import",
"tada.modules.acoustic_spkr_verf",
"--hidden-import",
"tada.nn",
"--hidden-import",
"tada.nn.vibevoice",
"--hidden-import",
"tada.utils",
"--hidden-import",
"tada.utils.gray_code",
"--hidden-import",
"tada.utils.text",
# descript-audio-codec (DAC) — used by TADA for Snake1d layers
"--hidden-import",
"dac",
"--hidden-import",
"dac.nn",
"--hidden-import",
"dac.nn.layers",
"--hidden-import",
"dac.model",
"--hidden-import",
"dac.model.dac",
"--collect-all",
"dac",
"--hidden-import",
"torchaudio",
"--collect-submodules",
"tada",
]
)