fix(backend): keep the trimmed clip when a runaway retry cannot split further; track the S3 tokenizer repo in model status and delete

Review follow-ups: with retries_runaway on for an engine that also has a
trim step, a <=100-char chunk flagged as runaway raised instead of
falling back to the trimmed output that already cuts the silence+noise
tail; generate_chunked now returns the trimmed chunk in that terminal
case. ModelConfig gains aux_hf_repo_ids so /models status only reports
the Chatterbox MLX model as downloaded once mlx-community/S3TokenizerV2
is present too (matching the backend's own cache check) and
DELETE /models/{name} removes that repo as well.
This commit is contained in:
jamiepine
2026-10-04 00:25:58 +00:00
committed by capy-ai-staging[bot]
parent ad6ec3c6ef
commit a2b453afed
4 changed files with 40 additions and 2 deletions
+9
View File
@@ -265,6 +265,15 @@ async def generate_chunked(
if runaway_detector is not None and runaway_detector(chunk_audio, chunk_sr):
if retry_depth >= MAX_RUNAWAY_RETRIES or len(chunk_text) <= MIN_RUNAWAY_RETRY_CHARS:
if trim_fn is not None:
# Engines with a trim step (Chatterbox) already cut the
# silence-then-noise tail; prefer the trimmed clip over
# failing the whole generation when we cannot split further.
logger.warning(
"Unstable TTS output for %d chars could not be retried further; keeping trimmed output",
len(chunk_text),
)
return np.asarray(trim_fn(chunk_audio, chunk_sr), dtype=np.float32), chunk_sr
raise RuntimeError(
"TTS output remained unstable after retrying smaller text chunks"
)