mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-04 01:25:18 -07:00
The Chatterbox backend is pinned to the CPU on macOS, so voice cloning runs at roughly 4x realtime there. This adds an MLX/Metal backend for the same engine and selects it on Apple Silicon, mirroring the split the qwen engine already makes between mlx_backend and pytorch_backend. Measured on an M4 Max (36 GB) with a cloned pt-BR profile, same API, same profile, model already loaded: short sentence (1.8s of audio): 7.4-9.2s -> 1.2s longer sentence (5.0s of audio): 23.7s -> 3.1s The model config for chatterbox-tts is now backend aware, same as the qwen configs, so the download matches the backend that will consume it. Nothing changes off Apple Silicon: the PyTorch backend is still selected there, and the CPU pinning it relies on is untouched.