mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
With padding="longest" the encoder rejects inputs shorter than 3000 mel frames when generate() has to detect the language first, so every short clip without a forced language failed with "Whisper expects the mel input features to be of length 3000". Use the long-form feature-extractor and generate() options only when the audio exceeds one 30s window.