mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-04 01:25:18 -07:00
Review follow-ups: clear_voice_prompt_memory_cache() now runs before backend.unload_model() at all three call sites so the device-resident prompt tensors are already released when empty_device_cache() / empty_mlx_cache() runs, instead of going back into the caching allocator afterwards. empty_mlx_cache's docstring states why it may run off the MLX worker thread.