Files
voicebox/backend/backends
jamiepineandcapy-ai-staging[bot] 51882b5065 fix(backend): keep the pad-to-30s Whisper path for clips under 30s
With padding="longest" the encoder rejects inputs shorter than 3000 mel
frames when generate() has to detect the language first, so every short
clip without a forced language failed with "Whisper expects the mel input
features to be of length 3000". Use the long-form feature-extractor and
generate() options only when the audio exceeds one 30s window.
2026-10-04 00:01:34 +00:00
..