Files
voicebox/backend/tests
JnyRoadandcapy-ai-staging[bot] 19f8f51408 fix(backend): release memory when unloading MLX models
Unloading a TTS/Whisper/LLM model on the MLX backend only dropped the
Python reference (`del self.model`). MLX keeps freed array buffers in
its own allocator pool for reuse instead of returning them to the OS,
so the process's memory footprint never actually shrank after unload
on Apple Silicon (the default backend there) until the process exited.

Add empty_mlx_cache() (backend/backends/base.py), wrapping
mx.clear_cache(), and call it from the three MLX unload_model()
implementations: MLXTTSBackend, MLXSTTBackend, MLXQwenLLMBackend.

Separately, the voice-clone prompt cache (backend/utils/cache.py) is a
process-lifetime dict populated by create_voice_prompt() across every
TTS engine, but nothing ever cleared it on model unload — only the
unrelated /tasks/clear-cache endpoint touched it. Add
clear_voice_prompt_memory_cache() (memory only, disk cache untouched
so a later generation still reloads the prompt instead of recomputing
it) and wire it into every TTS unload path (services/tts.py and the
qwen_custom_voice / generic branches of unload_model_by_config).
Whisper and the LLM backends never produce voice prompts, so their
unload paths are left alone.

Testing:
- New unit tests: backend/tests/test_mlx_unload_clears_cache.py,
  backend/tests/test_voice_prompt_cache_unload.py (8 tests, all pass).
- Verified end-to-end on Apple Silicon against real cached models
  (Qwen TTS 1.7B, Whisper Turbo, Qwen3 0.6B): loaded each via the
  running app, unloaded via the real /models/{name}/unload endpoint,
  and confirmed via mx.get_cache_memory()/get_active_memory() that the
  MLX allocator's cache drops to 0 on every cycle. Ran a real
  voice-clone generation end to end and confirmed the in-memory prompt
  cache goes from 1 entry to 0 on unload while the on-disk .prompt
  file is left intact.
2026-10-04 00:25:53 +00:00
..
2026-04-18 16:19:12 -07:00
2026-01-30 16:48:14 -08:00
2026-03-16 01:46:19 -07:00

Backend Tests

Manual test scripts for debugging and validating backend functionality.

Test Files

test_generation_progress.py

Tests TTS generation with SSE progress monitoring to identify UX issues where users see download progress even when the model is already cached.

Usage:

cd backend
python tests/test_generation_progress.py

Prerequisites:

  • Server must be running (python main.py)
  • At least one voice profile must exist

test_real_download.py

Tests real model download with SSE progress monitoring.

Usage:

cd backend
# Delete cache first to force fresh download
rm -rf ~/.cache/huggingface/hub/models--openai--whisper-base
python tests/test_real_download.py

Prerequisites:

  • Server must be running (python main.py)

test_progress.py

Unit tests for ProgressManager and HFProgressTracker functionality.

Usage:

cd backend
python tests/test_progress.py

test_check_progress_state.py

Debugging script to inspect the internal state of ProgressManager and TaskManager.

Usage:

cd backend
python tests/test_check_progress_state.py

Notes

These are manual test scripts, not automated unit tests. They're designed for:

  • Debugging progress tracking issues
  • Validating SSE event streams
  • Monitoring real-time download behavior
  • Inspecting internal state during development