Fix infinite HF retry storm when loading a cached model offline

model_load_progress() already received is_cached but never forwarded it
to force_offline_if_cached(), which is fully unit-tested but had zero
callers. A fully-cached model still resolved every config file against
huggingface.co, eating the default 5-retry backoff per file offline.
This commit is contained in:
Roman Dolgov
2026-10-03 09:18:41 +00:00
committed by jamiepine
parent 51f49dea19
commit 0517c0687a
3 changed files with 66 additions and 1 deletions
+8
View File
@@ -7,6 +7,14 @@
## [Unreleased]
### Reliability
- **Cached models no longer retry HuggingFace when offline.** Loading a fully-downloaded
model now forces offline mode for the duration of the load, so it skips the network HEAD
request (and its 5-retry backoff) for every config file — `config.json`,
`generation_config.json`, and the rest — instead of retrying each one in sequence before the
app becomes ready.
### Linux
- **ROCm setup works on Linux AMD systems.** Docker ROCm builds now keep PyTorch