mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
CodeRabbit follow-up on the previous round: the asyncio.Lock closed the gap for callers going through generate()/transcribe() consistently, but unload_model() itself never touched that lock and could still land between the load future resolving and the generate/transcribe future being submitted - two separate _run_on_mlx_thread calls, so a real gap existed at the Python level even though the executor is single-worker. Fix: generate() and transcribe() each now submit exactly ONE callable to the MLX executor - a self-healing reload-if-needed-then-infer function - instead of a load submission followed by a separate infer submission. This removes the gap structurally: nothing can observe an intermediate state because there is no intermediate state exposed across an await boundary. unload_model() now also submits unconditionally (loaded-check moved inside _unload_model_sync, which runs atomically with the teardown) rather than racing its own Python-level self.model check against a load in flight. Verified: same-size generate, size-switch generate, unload route, and a post-unload generate all complete clean on the live launchd server.