fix(backend): keep the trimmed clip when a runaway retry cannot split further; track the S3 tokenizer repo in model status and delete

Review follow-ups: with retries_runaway on for an engine that also has a
trim step, a <=100-char chunk flagged as runaway raised instead of
falling back to the trimmed output that already cuts the silence+noise
tail; generate_chunked now returns the trimmed chunk in that terminal
case. ModelConfig gains aux_hf_repo_ids so /models status only reports
the Chatterbox MLX model as downloaded once mlx-community/S3TokenizerV2
is present too (matching the backend's own cache check) and
DELETE /models/{name} removes that repo as well.
This commit is contained in:
jamiepine
2026-10-04 00:25:58 +00:00
committed by capy-ai-staging[bot]
parent ad6ec3c6ef
commit a2b453afed
4 changed files with 40 additions and 2 deletions
+2 -1
View File
@@ -20,6 +20,7 @@ from typing import ClassVar
import numpy as np
from . import CHATTERBOX_MLX_S3_TOKENIZER_REPO
from .base import (
combine_voice_prompts as _combine_voice_prompts,
is_model_cached,
@@ -32,7 +33,7 @@ logger = logging.getLogger(__name__)
CHATTERBOX_MLX_HF_REPO = "mlx-community/chatterbox-multilingual-v3"
# mlx-audio's Model.from_pretrained fetches the S3 speech tokenizer from this
# second repo (~470 MB), so the engine is only "downloaded" once both are cached.
S3_TOKENIZER_HF_REPO = "mlx-community/S3TokenizerV2"
S3_TOKENIZER_HF_REPO = CHATTERBOX_MLX_S3_TOKENIZER_REPO
# Files that must be present for the MLX multilingual model
_MLX_WEIGHT_FILES = ["model.safetensors", "config.json", "tokenizer.json"]