mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
fix(kokoro): trim edges only, keep inter-segment gaps
Registering Kokoro with needs_trim=True routed its output through the generic trim_tts_output, whose 1s internal-silence cut was tuned for Chatterbox hallucinations. KPipeline synthesizes newline- and token-limit-separated segments independently, and the ~0.3s lead plus ~0.7s tail pads at each boundary add up to 1.2s of silence, so with af_sarah a four-paragraph script came back as its first paragraph only (10.9s -> 1.8s) and a 51s text lost half its segments. Drop the needs_trim flag, keep the in-backend trim, and give trim_tts_output a max_internal_silence_ms=None mode that only trims the leading and trailing pads. Tests updated for the new behaviour, with a two-segment fake pipeline and an explicit internal-gap case.
This commit is contained in:
committed by
capy-ai-staging[bot]
parent
615aeaeb35
commit
ae300c5316
@@ -369,7 +369,6 @@ def _get_non_qwen_tts_configs() -> list[ModelConfig]:
|
||||
engine="kokoro",
|
||||
hf_repo_id="hexgrad/Kokoro-82M",
|
||||
size_mb=350,
|
||||
needs_trim=True,
|
||||
languages=["en", "es", "fr", "hi", "it", "pt", "ja", "zh"],
|
||||
),
|
||||
]
|
||||
|
||||
Reference in New Issue
Block a user