fix(offline-guard): make every cache check cover the files the forced-offline load reads

The load body now runs with HF_HUB_OFFLINE forced when _is_model_cached()
reports True, so a snapshot that holds the weights but not the small files
the loader also opens would fail hard instead of fetching them. Verified
against chatterbox-tts 0.1.7 (mtl_tts.py allow_patterns, tts_turbo.py
from_local) and the TADA loader's unsloth/Llama-3.2-1B tokenizer download;
list those files in the required_files checks. Also note in the changelog
why the 0.4.5 removal of this guard no longer applies.
This commit is contained in:
jamiepine
2026-10-03 09:35:49 +00:00
parent 32ba50cdd2
commit 865c324ff6
4 changed files with 36 additions and 5 deletions
+10 -1
View File
@@ -13,7 +13,16 @@
model now forces offline mode for the duration of the load, so it skips the network HEAD
request (and its 5-retry backoff) for every config file — `config.json`,
`generation_config.json`, and the rest — instead of retrying each one in sequence before the
app becomes ready.
app becomes ready. This reinstates the load-time `force_offline_if_cached` guard that 0.4.5
([#530](https://github.com/jamiepine/voicebox/pull/530)) removed: that removal was a hotfix
for the `_patch_mistral_regex` crash ([#526](https://github.com/jamiepine/voicebox/issues/526)),
which the wrapper installed in the same release now catches at the source, so the guard no
longer trips it. The per-file HEAD retries from
[#434](https://github.com/jamiepine/voicebox/issues/434) were never covered by that wrapper.
Because a load now fails hard offline when any file is missing, every backend's cache check
lists the full set of files its load reads (Chatterbox tokenizer/conds, TADA's Llama
tokenizer mirror), so a partially downloaded snapshot reports "not cached" and downloads
online instead.
### Linux