Commit Graph
2 Commits
Author SHA1 Message Date
Charles Hasseandcapy-ai-staging[bot] 502123a943 fix(backend): retry runaway generations on the Chatterbox MLX backend
The qwen configs already set retries_runaway on MLX, with the reason in a comment:
mlx-audio can continue past an EOS miss and emit silence followed by codec noise.
The Chatterbox MLX backend added in this PR hits the same failure and did not have
the guard wired.

Observed on an M4 Max with a cloned pt-BR profile: a 145 character sentence took
18.0s and returned 5.3s of audio for text worth about 8s, and a short reply came
back as an endless hiss. With retries_runaway enabled the same sentence takes 3.4s
and returns the full 7.4s of speech.
2026-10-04 00:25:58 +00:00
Charles Hasseandcapy-ai-staging[bot] ad5d64c0a3 feat(backend): run Chatterbox multilingual on MLX for Apple Silicon
The Chatterbox backend is pinned to the CPU on macOS, so voice cloning runs at
roughly 4x realtime there. This adds an MLX/Metal backend for the same engine and
selects it on Apple Silicon, mirroring the split the qwen engine already makes
between mlx_backend and pytorch_backend.

Measured on an M4 Max (36 GB) with a cloned pt-BR profile, same API, same profile,
model already loaded:

  short sentence (1.8s of audio):  7.4-9.2s  ->  1.2s
  longer sentence (5.0s of audio): 23.7s     ->  3.1s

The model config for chatterbox-tts is now backend aware, same as the qwen configs,
so the download matches the backend that will consume it.

Nothing changes off Apple Silicon: the PyTorch backend is still selected there, and
the CPU pinning it relies on is untouched.
2026-10-04 00:25:58 +00:00