With asyncio.to_thread the new backend hits the same thread-local stream
crash as the Qwen MLX backend once the default pool is busy (reproduced
on an M2 Ultra: 'There is no Stream(cpu, 2) in current thread.' on the
first generate after load). Route load and generate through
mlx_backend._run_on_mlx_thread, fold load-if-needed + generate into one
worker submission under the backend's lock, and bind the model locally
inside the closure, matching MLXTTSBackend.
The Chatterbox backend is pinned to the CPU on macOS, so voice cloning runs at
roughly 4x realtime there. This adds an MLX/Metal backend for the same engine and
selects it on Apple Silicon, mirroring the split the qwen engine already makes
between mlx_backend and pytorch_backend.
Measured on an M4 Max (36 GB) with a cloned pt-BR profile, same API, same profile,
model already loaded:
short sentence (1.8s of audio): 7.4-9.2s -> 1.2s
longer sentence (5.0s of audio): 23.7s -> 3.1s
The model config for chatterbox-tts is now backend aware, same as the qwen configs,
so the download matches the backend that will consume it.
Nothing changes off Apple Silicon: the PyTorch backend is still selected there, and
the CPU pinning it relies on is untouched.