mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-09-16 13:20:39 -07:00
fix(build): repair frozen-binary imports for kokoro, chatterbox-multilingual, scipy, transformers (#438)
* fix(build): bundle kokoro source files for transformers runtime introspection transformers opens .py source files at runtime to check attention/MoE implementation via regex (e.g. _can_set_attn_implementation). PyInstaller's --hidden-import only bundles .pyc bytecode, so kokoro/modules.py was missing from the bundle causing a FileNotFoundError on Kokoro model load. Switch from individual --hidden-import entries to --collect-all kokoro in both build_binary.py and voicebox-server.spec. The kokoro package is 172K so no meaningful bundle size impact. Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix(build): use SPECPATH for runtime hook instead of hardcoded absolute path The linter expanded runtime_hooks=[] to an absolute /Users/... path which would break CI and other dev machines. Use os.path.join(SPECPATH, ...) to mirror the relative approach in build_binary.py. Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix(build): runtime hook to work around PyInstaller + Python 3.12 import breakages Four distinct bundling-specific crashes blocked Kokoro and Qwen CustomVoice from loading in the frozen binary: 1. torch._dynamo import triggered via class-body decorators (@torch._dynamo.allow_in_graph on PreTrainedModel, @torch.compiler.disable in flex_attention) pulls in torch._numpy._ufuncs which crashes on module load with NameError: name 'name' is not defined. 2. AlbertModel (Kokoro) triggers @auto_docstring -> modeling_auto -> GenerationMixin -> candidate_generator -> sklearn -> scipy, which hits the same class of bug in scipy.stats._distn_infrastructure (NameError: name 'obj' is not defined). 3. AutoModel (Qwen) pulls the same sklearn -> scipy chain directly. 4. librosa (required by most TTS engines) -> scipy.signal -> scipy.stats hits the _distn_infrastructure crash regardless of the transformers stubs above. The root cause of (1) and (4) is that PyInstaller's frozen importer runs module-level `for X in [<list-comp using dir()>]:` loops with an empty iterable, leaving the loop variable unbound. Trailing `del obj` / unrelated references then crash. Fix: a single runtime hook (pyi_rth_torch_compiler_disable.py) installs: - sys.modules stubs for torch._dynamo and torch._dynamo.config, plus a meta-path finder for torch._dynamo.* submodules — voicebox never uses torch.compile/dynamo for inference, so a permissive no-op stub (callable as decorator, falsey as predicate, context-manager-safe for TransformGetItemToIndex) is drop-in safe. - meta-path finder stubs for transformers.utils.auto_docstring and transformers.generation.candidate_generator — both import-chain short-circuits; docstrings and speculative decoding aren't used for TTS. - meta-path finder for scipy.stats._distn_infrastructure that reads the real .py source via the wrapped loader's get_source(), replaces the bundling-broken `del obj` with `globals().pop('obj', None)`, and compile+exec's the patched source. This keeps the real scipy module intact so librosa and everything downstream works normally. Supporting changes: - backend/pyi_hooks/hook-scipy.stats._distn_infrastructure.py sets module_collection_mode = "pyz+py" so the .py source is actually in the bundle for the runtime patcher to read. - build_binary.py and voicebox-server.spec register the runtime hook and the new hooks dir. Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix(build): force transformers torch<2.6 mask path and bundle spacy_pkuseg - patch transformers.masking_utils to set _is_torch_greater_or_equal_than_2_6 = False, forcing sdpa_mask_older_torch and avoiding the vmap .item() crash that breaks Qwen CustomVoice generation (our torch._dynamo stub can't reproduce TransformGetItemToIndex's graph transform). - add PyInstaller hook to bundle transformers.masking_utils .py source so the runtime finder can source-patch it. - --collect-all spacy_pkuseg so Chatterbox Multilingual can load its Chinese segmenter (dicts/default.pkl + native .so extensions). - add per-finder install diagnostics + _HOOK_VERSION marker to make future bundle-only regressions easier to triage. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * fix(build): pass PyInstaller hook paths relative so .spec is portable Absolute paths ended up in the auto-regenerated voicebox-server.spec because build_binary.py prefixed every --runtime-hook and --additional-hooks-dir with str(backend_dir / ...). That broke builds on any machine whose checkout wasn't at /Users/jamie/... and anyone invoking pyinstaller voicebox-server.spec directly. os.chdir(backend_dir) already runs before PyInstaller (same reason server.py works as a bare filename), so the backend_dir prefix is unnecessary. Drop it so the generated spec references pyi_hooks/, pyi_rth_numpy_compat.py, pyi_rth_torch_compiler_disable.py as repo- relative paths. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> --------- Co-authored-by: Claude Opus 4.6 (1M context) <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 4.6
parent
54a3bf322e
commit
c8cb12f1bc
+23
-12
@@ -55,10 +55,23 @@ def build_server(cuda=False):
|
||||
# numpy 2.x / torch ABI mismatch fix: install memmove fallback for
|
||||
# torch.from_numpy() before the app starts. Runtime hooks run after
|
||||
# FrozenImporter is registered so frozen torch/numpy are importable.
|
||||
# Paths are passed relative to backend_dir because os.chdir(backend_dir)
|
||||
# runs before PyInstaller. Absolute paths would get baked into the
|
||||
# generated .spec, breaking reproducible builds on other machines / CI.
|
||||
args.extend(
|
||||
[
|
||||
"--runtime-hook",
|
||||
str(backend_dir / "pyi_rth_numpy_compat.py"),
|
||||
"pyi_rth_numpy_compat.py",
|
||||
# Stub torch.compiler.disable before transformers imports
|
||||
# flex_attention, which otherwise triggers torch._dynamo →
|
||||
# torch._numpy._ufuncs and crashes at module load under
|
||||
# PyInstaller. See pyi_rth_torch_compiler_disable.py.
|
||||
"--runtime-hook",
|
||||
"pyi_rth_torch_compiler_disable.py",
|
||||
# Per-module collection overrides (e.g. forcing scipy.stats._distn_infrastructure
|
||||
# to bundle .py source alongside .pyc so the runtime hook can source-patch it).
|
||||
"--additional-hooks-dir",
|
||||
"pyi_hooks",
|
||||
]
|
||||
)
|
||||
|
||||
@@ -125,6 +138,11 @@ def build_server(cuda=False):
|
||||
"backend.backends.chatterbox_backend",
|
||||
"--hidden-import",
|
||||
"backend.backends.chatterbox_turbo_backend",
|
||||
# chatterbox multilingual uses spacy_pkuseg for Chinese word
|
||||
# segmentation, which ships pickled dict files (dicts/default.pkl)
|
||||
# and native .so extensions that --hidden-import alone won't bundle.
|
||||
"--collect-all",
|
||||
"spacy_pkuseg",
|
||||
"--hidden-import",
|
||||
"backend.backends.luxtts_backend",
|
||||
"--hidden-import",
|
||||
@@ -241,20 +259,13 @@ def build_server(cuda=False):
|
||||
"--collect-submodules",
|
||||
"tada",
|
||||
# Kokoro 82M — lightweight TTS engine using misaki G2P
|
||||
# collect-all is required because transformers introspects .py source
|
||||
# files at runtime (e.g. _can_set_attn_implementation opens the class
|
||||
# file); hidden-import alone only bundles bytecode.
|
||||
"--hidden-import",
|
||||
"backend.backends.kokoro_backend",
|
||||
"--hidden-import",
|
||||
"--collect-all",
|
||||
"kokoro",
|
||||
"--hidden-import",
|
||||
"kokoro.pipeline",
|
||||
"--hidden-import",
|
||||
"kokoro.model",
|
||||
"--hidden-import",
|
||||
"kokoro.istftnet",
|
||||
"--hidden-import",
|
||||
"kokoro.modules",
|
||||
"--hidden-import",
|
||||
"kokoro.custom_stft",
|
||||
# misaki ships G2P data files (dictionaries, phoneme tables)
|
||||
# that must be bundled for espeak/en/ja/zh G2P to work
|
||||
"--collect-all",
|
||||
|
||||
Reference in New Issue
Block a user