Files
voicebox/backend/backends
2693c983f2 fix: extend MLX thread-affinity fix to the Qwen3 LLM backend
The dedicated-MLX-thread fix (#886) covers the TTS and STT backends, but
MLXQwenLLMBackend still dispatches load/generate through
asyncio.to_thread(), so /llm/generate (personality compose/rewrite,
dictation refinement) crashes with the same
'There is no Stream(gpu, 0) in current thread.' — raised from
mlx_lm.generate's wired_limit on exit — whenever load and generate land
on different pool threads.

Route the MLX LLM backend's unload/load/generate through the same
_run_on_mlx_thread helper so every MLX call in the process shares one
worker thread.

Verified on Apple M3 Pro (macOS 26.5, mlx 0.32.0, mlx-lm 0.31.1,
mlx-audio 0.4.1): /llm/generate failed 100% before, succeeds after;
TTS + STT unaffected.

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-10-04 00:01:47 +00:00
..