mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
Two real follow-on issues found by CodeRabbit on PR #989: 1. MLXTTSBackend.load_model_async called self.unload_model() directly on the event-loop thread before dispatching _load_model_sync to the MLX worker thread - model teardown ran on the wrong OS thread, same class of bug the PR itself fixes. Combined unload+load into one _reload_sync callable submitted as a single MLX-thread operation. 2. Public unload_model() (called synchronously from the /models/unload routes) also ran on the caller thread. Now submits to the MLX executor and blocks on the result, so teardown always happens on the worker thread regardless of caller. Same fix applied to MLXSTTBackend. 3. Both backends cache self.model on the instance, but generate()/ transcribe() only locked their own internal load step - a concurrent request for a different model_size could swap self.model between one request's load and its inference (real race: routes/ generations.py calls load_engine_model() and generate_chunked() as separate awaited steps with a gap between them). Added a per-backend asyncio.Lock held across the full load+inference sequence in both generate() and transcribe(). Note: this closes the race for callers using the backend's own public methods consistently; the wider route-level orchestration race (load_engine_model + generate_chunked as two separate calls) is a follow-up outside this file's scope. Verified: same-size and cross-size-switch generations both complete clean after the patch; /models/{name}/unload route returns 200 without deadlocking the (still-responsive) server.