mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
Registering Kokoro with needs_trim=True routed its output through the generic trim_tts_output, whose 1s internal-silence cut was tuned for Chatterbox hallucinations. KPipeline synthesizes newline- and token-limit-separated segments independently, and the ~0.3s lead plus ~0.7s tail pads at each boundary add up to 1.2s of silence, so with af_sarah a four-paragraph script came back as its first paragraph only (10.9s -> 1.8s) and a 51s text lost half its segments. Drop the needs_trim flag, keep the in-backend trim, and give trim_tts_output a max_internal_silence_ms=None mode that only trims the leading and trailing pads. Tests updated for the new behaviour, with a two-segment fake pipeline and an explicit internal-gap case.