Files
voicebox/backend/backends
jamiepineandcapy-ai-staging[bot] cffe24ffd2 fix(mlx): drain the MLX pool after an in-flight op when unload landed mid-generation
Review follow-up: unload_model() runs inline on the event loop, so when it
lands during a generation it only drops the backend's reference; the
worker's local keeps the model alive and mx.clear_cache() finds nothing to
free. Once the folded load-and-generate (or transcribe) callable finishes
and releases that local, check whether the model was unloaded meanwhile
and drain the pool then. Covered by a test that unloads from inside a
fake model.generate().
2026-10-04 00:25:53 +00:00
..