mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-04 01:25:18 -07:00
fix(backend): drop the prompt cache before the allocator-empty; note clear_cache is thread-agnostic
Review follow-ups: clear_voice_prompt_memory_cache() now runs before backend.unload_model() at all three call sites so the device-resident prompt tensors are already released when empty_device_cache() / empty_mlx_cache() runs, instead of going back into the caching allocator afterwards. empty_mlx_cache's docstring states why it may run off the MLX worker thread.
This commit is contained in:
committed by
capy-ai-staging[bot]
parent
1a803aa05f
commit
b1323ed0a8
@@ -232,6 +232,12 @@ def empty_mlx_cache() -> None:
|
||||
returning them to the OS. Backends must call this after unloading an
|
||||
MLX model, or the process's memory footprint never shrinks even though
|
||||
the model object itself was dropped.
|
||||
|
||||
Safe from any thread: ``mx.clear_cache`` only drains the global
|
||||
allocator pool and never touches the per-thread stream registry, so
|
||||
unlike load/generate it does not have to run on the MLX worker thread
|
||||
(verified from the FastAPI event loop with a generation in flight on
|
||||
the worker).
|
||||
"""
|
||||
import mlx.core as mx
|
||||
|
||||
|
||||
Reference in New Issue
Block a user