Files
voicebox/backend/backends
Ron David Ben Ishayandcapy-ai-staging[bot] 57ef6636bc fix(mlx): fold reload+inference into one MLX-worker submission
CodeRabbit follow-up on the previous round: the asyncio.Lock closed the
gap for callers going through generate()/transcribe() consistently, but
unload_model() itself never touched that lock and could still land
between the load future resolving and the generate/transcribe future
being submitted - two separate _run_on_mlx_thread calls, so a real gap
existed at the Python level even though the executor is single-worker.

Fix: generate() and transcribe() each now submit exactly ONE callable
to the MLX executor - a self-healing reload-if-needed-then-infer
function - instead of a load submission followed by a separate infer
submission. This removes the gap structurally: nothing can observe an
intermediate state because there is no intermediate state exposed
across an await boundary. unload_model() now also submits
unconditionally (loaded-check moved inside _unload_model_sync, which
runs atomically with the teardown) rather than racing its own
Python-level self.model check against a load in flight.

Verified: same-size generate, size-switch generate, unload route, and
a post-unload generate all complete clean on the live launchd server.
2026-10-04 00:01:47 +00:00
..