Files
voicebox/backend/services
jamiepineandcapy-ai-staging[bot] b1323ed0a8 fix(backend): drop the prompt cache before the allocator-empty; note clear_cache is thread-agnostic
Review follow-ups: clear_voice_prompt_memory_cache() now runs before
backend.unload_model() at all three call sites so the device-resident
prompt tensors are already released when empty_device_cache() /
empty_mlx_cache() runs, instead of going back into the caching allocator
afterwards. empty_mlx_cache's docstring states why it may run off the
MLX worker thread.
2026-10-04 00:25:53 +00:00
..