mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
Routing unload through the single MLX worker and waiting on the result blocked the caller (the FastAPI event loop via /models/unload) for the remainder of any in-flight generation — measured 3.7 s stall on a 16 s clip on an M2 Ultra, unbounded for long texts. Dropping the model reference inline is thread-safe (MLX frees buffers through its global allocator) and is what main did before the thread-affinity change; the generation in flight keeps its own reference and the next generate() reloads on the worker.