mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-04 01:25:18 -07:00
With asyncio.to_thread the new backend hits the same thread-local stream crash as the Qwen MLX backend once the default pool is busy (reproduced on an M2 Ultra: 'There is no Stream(cpu, 2) in current thread.' on the first generate after load). Route load and generate through mlx_backend._run_on_mlx_thread, fold load-if-needed + generate into one worker submission under the backend's lock, and bind the model locally inside the closure, matching MLXTTSBackend.