Commit Graph
1 Commits
Author SHA1 Message Date
99d05b917c fix(backends): run MLX load and inference on one thread
MLX streams are thread-local and mlx-audio caches one on the model at
load time. Model load and inference were each dispatched through
asyncio.to_thread(), which uses a multi-worker pool, so load and
generate/transcribe could land on different OS threads -- the inference
thread then has no Stream(gpu, N) and MLX aborts with
"There is no Stream(gpu, 1) in current thread."

Route every MLX call (load, generate, transcribe; TTS and STT) through a
single dedicated worker thread so a model and its stream always share a
thread. max_workers=1 also serialises the single local GPU. Adds a
regression test covering the thread-affinity invariant.

Fixes #699. Also addresses #675.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-10-04 00:01:47 +00:00