Review feedback on the preprocessor:
1. ``trim_top_db=30`` was labelled "conservative" in the docstring but is
actually *more* aggressive than librosa's default of 60. Normal
speech dynamic range sits around 30 dB, so 30 dB would eat quiet
trailing syllables and soft consonants. Raise the default to 40 dB —
below normal speech dynamic range but still catching obvious edge
silence — and fix the docstring.
2. Unconditional 100 ms edge padding ran even when ``librosa.effects.trim``
removed nothing. For a well-recorded 29.9 s upload that path would
push the waveform past the 30 s ceiling and trigger a spurious "too
long" rejection. Only pad when trimming actually shortened the
audio, and cap the pad so the output never exceeds the input length.
Adds a regression test for the net-neutral length behaviour.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Uploaded/recorded voice samples were rejected outright whenever the peak
exceeded 0.99 ("Audio is clipping (reduce input gain)"). That wasn't
actionable: a recording in the app has no pre-gain control, and an
already-captured file can't be re-taken by the user. The Settings
"Normalize audio" toggle only affects generated TTS output, so users who
enabled it expecting it to help with sample uploads were still blocked.
Replace the hard reject with a small, always-on preprocess step that
runs right after load:
- DC-offset removal
- Conservative edge-silence trim (top_db=30) with 100 ms padding kept
- Peak cap at 0.95 if the input peak exceeds that
Duration and RMS checks now run on the preprocessed waveform, so
samples that were previously rejected for being "hot" are accepted and
stored with safe headroom. True in-waveform clipping artifacts still
can't be repaired — peak scaling only prevents downstream re-clipping
during multi-sample combination and TTS inference.
Adds a unit-test file (previously none existed for audio.py).
Fixes#456.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Move audio validation and saving to thread pool so librosa/ffmpeg decoding
doesn't block the async event loop. Combine validate + load into a single
pass to avoid decoding the file twice. Add 50 MB upload limit and chunked
reads to prevent unbounded memory allocation.
Closes#278
soundfile cannot infer format from .tmp extension, causing all
generations to fail with 'No format specified and unable to get
format from file extension'
- save_audio() now writes to .tmp then os.replace() for atomic writes
- /generate endpoint catches OSError with specific messages for ENOENT, EACCES, ENOSPC, and BrokenPipeError
- New /health/filesystem endpoint checks directory existence, write permissions, and disk space
- New DirectoryCheck and FilesystemHealthResponse models
Cherry-picked and expanded from #178 (@Vaibhavee89)
- New ChatterboxTTSBackend wrapping ChatterboxMultilingualTTS (ResembleAI/chatterbox)
- Supports 23 languages including Hebrew, forces CPU on macOS (MPS issue)
- Monkey-patches torch.load for CPU loading, forces eager attention for compatibility
- trim_tts_output utility cuts trailing silence/hallucination from Chatterbox output
- Full engine integration: /generate, /generate/stream, model status/download/delete
- Hebrew (he) added to supported languages in frontend and backend validation
- Single flat model dropdown extended with Chatterbox option in both generation UIs
- ModelManagement UI groups LuxTTS and Chatterbox under 'Other Voice Models' section