mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-04 01:25:18 -07:00
Review follow-up: unload_model() runs inline on the event loop, so when it lands during a generation it only drops the backend's reference; the worker's local keeps the model alive and mx.clear_cache() finds nothing to free. Once the folded load-and-generate (or transcribe) callable finishes and releases that local, check whether the model was unloaded meanwhile and drain the pool then. Covered by a test that unloads from inside a fake model.generate().