mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
Review follow-ups on the thread-affinity PR: MLXQwenLLMBackend.generate awaited load_model as one worker submission and then submitted _generate_sync as a second, leaving the same load/generate gap the TTS and STT paths close (a concurrent load_model for another size or an unload could swap or null out self.model/self.tokenizer in between). Give the LLM backend the same _op_lock + single _reload_and_generate_sync shape. The TTS/STT/LLM sync closures now bind the model to a local once, so an inline unload_model() from the event loop mid-generation cannot turn a later self.model read (the voice-clone fallback path in particular) into an AttributeError.