fix(kokoro): trim edges only, keep inter-segment gaps

Registering Kokoro with needs_trim=True routed its output through the
generic trim_tts_output, whose 1s internal-silence cut was tuned for
Chatterbox hallucinations. KPipeline synthesizes newline- and
token-limit-separated segments independently, and the ~0.3s lead plus
~0.7s tail pads at each boundary add up to 1.2s of silence, so with
af_sarah a four-paragraph script came back as its first paragraph only
(10.9s -> 1.8s) and a 51s text lost half its segments.

Drop the needs_trim flag, keep the in-backend trim, and give
trim_tts_output a max_internal_silence_ms=None mode that only trims the
leading and trailing pads. Tests updated for the new behaviour, with a
two-segment fake pipeline and an explicit internal-gap case.
This commit is contained in:
jamiepine
2026-10-04 00:01:24 +00:00
committed by capy-ai-staging[bot]
parent 615aeaeb35
commit ae300c5316
4 changed files with 48 additions and 23 deletions
-1
View File
@@ -369,7 +369,6 @@ def _get_non_qwen_tts_configs() -> list[ModelConfig]:
engine="kokoro",
hf_repo_id="hexgrad/Kokoro-82M",
size_mb=350,
needs_trim=True,
languages=["en", "es", "fr", "hi", "it", "pt", "ja", "zh"],
),
]
+6 -1
View File
@@ -291,7 +291,12 @@ class KokoroTTSBackend:
audio = np.concatenate(audio_chunks).astype(np.float32)
from ..utils.audio import trim_tts_output
audio = trim_tts_output(audio, sample_rate=KOKORO_SAMPLE_RATE)
# Edge-only trim: Kokoro pads ~0.3s before and ~0.7s after speech.
# The internal-gap cut is disabled because KPipeline synthesizes
# newline/token-limit segments independently and the pads at each
# segment boundary add up to >1s of silence; with the default cut
# everything after the first segment would be dropped.
audio = trim_tts_output(audio, sample_rate=KOKORO_SAMPLE_RATE, max_internal_silence_ms=None)
return audio, KOKORO_SAMPLE_RATE
return await asyncio.to_thread(_generate_sync)