mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
The dedicated-MLX-thread fix (#886) covers the TTS and STT backends, but MLXQwenLLMBackend still dispatches load/generate through asyncio.to_thread(), so /llm/generate (personality compose/rewrite, dictation refinement) crashes with the same 'There is no Stream(gpu, 0) in current thread.' — raised from mlx_lm.generate's wired_limit on exit — whenever load and generate land on different pool threads. Route the MLX LLM backend's unload/load/generate through the same _run_on_mlx_thread helper so every MLX call in the process shares one worker thread. Verified on Apple M3 Pro (macOS 26.5, mlx 0.32.0, mlx-lm 0.31.1, mlx-audio 0.4.1): /llm/generate failed 100% before, succeeds after; TTS + STT unaffected. Co-Authored-By: Claude Fable 5 <[email protected]>