Commit Graph
4 Commits
Author SHA1 Message Date
jamiepineandcapy-ai-staging[bot] a2b453afed fix(backend): keep the trimmed clip when a runaway retry cannot split further; track the S3 tokenizer repo in model status and delete
Review follow-ups: with retries_runaway on for an engine that also has a
trim step, a <=100-char chunk flagged as runaway raised instead of
falling back to the trimmed output that already cuts the silence+noise
tail; generate_chunked now returns the trimmed chunk in that terminal
case. ModelConfig gains aux_hf_repo_ids so /models status only reports
the Chatterbox MLX model as downloaded once mlx-community/S3TokenizerV2
is present too (matching the backend's own cache check) and
DELETE /models/{name} removes that repo as well.
2026-10-04 00:25:58 +00:00
jamiepineandcapy-ai-staging[bot] ad6ec3c6ef fix(backend): count the S3TokenizerV2 repo in the Chatterbox MLX cache check; register the backend with PyInstaller
Review follow-ups: mlx-audio's Model.from_pretrained fetches the S3
speech tokenizer from mlx-community/S3TokenizerV2 (~470 MB) separately
from the chatterbox checkout, so _is_model_cached now requires both
repos (same shape as the Hume backend's codec check) and the config's
size_mb reflects the real footprint. backend.backends.chatterbox_mlx_backend
is a function-level import that PyInstaller's graph will not see, so it
is added to the Apple Silicon hidden-import list in build_binary.py and
voicebox-server.spec next to mlx_backend.
2026-10-04 00:25:58 +00:00
jamiepineandcapy-ai-staging[bot] c5cf7436f7 fix(backend): run Chatterbox MLX load and generate on the shared MLX worker thread
With asyncio.to_thread the new backend hits the same thread-local stream
crash as the Qwen MLX backend once the default pool is busy (reproduced
on an M2 Ultra: 'There is no Stream(cpu, 2) in current thread.' on the
first generate after load). Route load and generate through
mlx_backend._run_on_mlx_thread, fold load-if-needed + generate into one
worker submission under the backend's lock, and bind the model locally
inside the closure, matching MLXTTSBackend.
2026-10-04 00:25:58 +00:00
Charles Hasseandcapy-ai-staging[bot] ad5d64c0a3 feat(backend): run Chatterbox multilingual on MLX for Apple Silicon
The Chatterbox backend is pinned to the CPU on macOS, so voice cloning runs at
roughly 4x realtime there. This adds an MLX/Metal backend for the same engine and
selects it on Apple Silicon, mirroring the split the qwen engine already makes
between mlx_backend and pytorch_backend.

Measured on an M4 Max (36 GB) with a cloned pt-BR profile, same API, same profile,
model already loaded:

  short sentence (1.8s of audio):  7.4-9.2s  ->  1.2s
  longer sentence (5.0s of audio): 23.7s     ->  3.1s

The model config for chatterbox-tts is now backend aware, same as the qwen configs,
so the download matches the backend that will consume it.

Nothing changes off Apple Silicon: the PyTorch backend is still selected there, and
the CPU pinning it relies on is untouched.
2026-10-04 00:25:58 +00:00