mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
Qwen3-TTS (and MLX STT) generation crashed with: "There is no Stream(gpu, N) in current thread." MLXTTSBackend/MLXSTTBackend dispatched model load and generate/ transcribe as separate asyncio.to_thread() calls, which round-robin across Python's default multi-worker executor pool. MLX's Metal backend keeps GPU streams registered per-OS-thread, so a model loaded on one worker thread and then used for generation on a different worker thread hits a missing stream and crashes. Reproduced 100% of the time on macOS/Apple Silicon cloning with both the 1.7B and 0.6B Qwen3-TTS models; Chatterbox/Kokoro were unaffected since they use the PyTorch backend, not this module. Fix: route all four MLX call sites in this file (TTS load, TTS generate, STT load, STT transcribe) through a dedicated single-worker ThreadPoolExecutor instead of asyncio.to_thread's shared pool, so every MLX operation for a given process runs on the same OS thread. Verified: direct /generate API calls against both model sizes completed cleanly after the fix (previously failed every time).