PR #443 wrapped the model *load* path with `force_offline_if_cached` so
cached models don't phone home at startup. The context manager restores
`HF_HUB_OFFLINE` on exit, which left inference paths (generate,
transcribe, voice-prompt creation) unguarded — and `qwen_tts`,
`mlx_audio`, and `transformers` perform lazy tokenizer/processor/config
lookups during inference. With internet on, those lookups are
near-instant and invisible; with internet off, `requests` hangs on DNS
or connect until the network returns. This is exactly what users in
#462 describe: model shows "Loaded", internet drops, generation
"thinks" forever, internet comes back, generation completes.
Chatterbox and LuxTTS don't exhibit this because their engine libs
resolve everything through already-cached paths at load time.
Fix: wrap each inference-sync body with `force_offline_if_cached(True,
...)`. Since inference only runs after a successful load, weights are
known to be on disk, so `is_cached=True` is unconditional.
Also adds the load-time guard that was missing from
`qwen_custom_voice_backend.py` — CustomVoice previously had no offline
protection at all.
Paths patched:
- PyTorchTTSBackend.create_voice_prompt (create_voice_clone_prompt)
- PyTorchTTSBackend.generate (generate_voice_clone)
- PyTorchSTTBackend.transcribe (Whisper generate + decoder-prompt-ids)
- MLXTTSBackend.generate (mlx_audio generate, all branches)
- MLXSTTBackend.transcribe (mlx_audio whisper generate)
- QwenCustomVoiceBackend._load_model_sync + generate
Does not address the secondary `check_model_inputs() missing 'func'`
error reported in the same issue — that's a `transformers` 5.x
version-skew bug on the install path, separate concern.
Fixes#462.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>