Review follow-ups on the non-fatal actool change: the stub Assets.car /
partial.plist fallback is only acceptable for dev builds, so gate it on
PROFILE != release and keep panicking for release bundles (release.yml
builds signed DMGs where an empty icon catalog must fail the job). The
warning now emits one cargo:warning per non-empty line of actool's
stdout and stderr — actool writes its compile diagnostics to stdout, and
Cargo only reads the first line of a warning directive.
Review follow-ups on the thread-affinity PR: MLXQwenLLMBackend.generate
awaited load_model as one worker submission and then submitted
_generate_sync as a second, leaving the same load/generate gap the TTS
and STT paths close (a concurrent load_model for another size or an
unload could swap or null out self.model/self.tokenizer in between).
Give the LLM backend the same _op_lock + single _reload_and_generate_sync
shape. The TTS/STT/LLM sync closures now bind the model to a local once,
so an inline unload_model() from the event loop mid-generation cannot
turn a later self.model read (the voice-clone fallback path in
particular) into an AttributeError.
Routing unload through the single MLX worker and waiting on the result
blocked the caller (the FastAPI event loop via /models/unload) for the
remainder of any in-flight generation — measured 3.7 s stall on a 16 s
clip on an M2 Ultra, unbounded for long texts. Dropping the model
reference inline is thread-safe (MLX frees buffers through its global
allocator) and is what main did before the thread-affinity change; the
generation in flight keeps its own reference and the next generate()
reloads on the worker.
MLX streams are thread-local and mlx-audio caches one on the model at
load time. Model load and inference were each dispatched through
asyncio.to_thread(), which uses a multi-worker pool, so load and
generate/transcribe could land on different OS threads -- the inference
thread then has no Stream(gpu, N) and MLX aborts with
"There is no Stream(gpu, 1) in current thread."
Route every MLX call (load, generate, transcribe; TTS and STT) through a
single dedicated worker thread so a model and its stream always share a
thread. max_workers=1 also serialises the single local GPU. Adds a
regression test covering the thread-affinity invariant.
Fixes#699. Also addresses #675.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The dedicated-MLX-thread fix (#886) covers the TTS and STT backends, but
MLXQwenLLMBackend still dispatches load/generate through
asyncio.to_thread(), so /llm/generate (personality compose/rewrite,
dictation refinement) crashes with the same
'There is no Stream(gpu, 0) in current thread.' — raised from
mlx_lm.generate's wired_limit on exit — whenever load and generate land
on different pool threads.
Route the MLX LLM backend's unload/load/generate through the same
_run_on_mlx_thread helper so every MLX call in the process shares one
worker thread.
Verified on Apple M3 Pro (macOS 26.5, mlx 0.32.0, mlx-lm 0.31.1,
mlx-audio 0.4.1): /llm/generate failed 100% before, succeeds after;
TTS + STT unaffected.
Co-Authored-By: Claude Fable 5 <[email protected]>
CodeRabbit follow-up on the previous round: the asyncio.Lock closed the
gap for callers going through generate()/transcribe() consistently, but
unload_model() itself never touched that lock and could still land
between the load future resolving and the generate/transcribe future
being submitted - two separate _run_on_mlx_thread calls, so a real gap
existed at the Python level even though the executor is single-worker.
Fix: generate() and transcribe() each now submit exactly ONE callable
to the MLX executor - a self-healing reload-if-needed-then-infer
function - instead of a load submission followed by a separate infer
submission. This removes the gap structurally: nothing can observe an
intermediate state because there is no intermediate state exposed
across an await boundary. unload_model() now also submits
unconditionally (loaded-check moved inside _unload_model_sync, which
runs atomically with the teardown) rather than racing its own
Python-level self.model check against a load in flight.
Verified: same-size generate, size-switch generate, unload route, and
a post-unload generate all complete clean on the live launchd server.
Two real follow-on issues found by CodeRabbit on PR #989:
1. MLXTTSBackend.load_model_async called self.unload_model() directly
on the event-loop thread before dispatching _load_model_sync to the
MLX worker thread - model teardown ran on the wrong OS thread, same
class of bug the PR itself fixes. Combined unload+load into one
_reload_sync callable submitted as a single MLX-thread operation.
2. Public unload_model() (called synchronously from the /models/unload
routes) also ran on the caller thread. Now submits to the MLX
executor and blocks on the result, so teardown always happens on the
worker thread regardless of caller. Same fix applied to
MLXSTTBackend.
3. Both backends cache self.model on the instance, but generate()/
transcribe() only locked their own internal load step - a
concurrent request for a different model_size could swap self.model
between one request's load and its inference (real race: routes/
generations.py calls load_engine_model() and generate_chunked() as
separate awaited steps with a gap between them). Added a per-backend
asyncio.Lock held across the full load+inference sequence in both
generate() and transcribe(). Note: this closes the race for callers
using the backend's own public methods consistently; the wider
route-level orchestration race (load_engine_model + generate_chunked
as two separate calls) is a follow-up outside this file's scope.
Verified: same-size and cross-size-switch generations both complete
clean after the patch; /models/{name}/unload route returns 200 without
deadlocking the (still-responsive) server.
Qwen3-TTS (and MLX STT) generation crashed with:
"There is no Stream(gpu, N) in current thread."
MLXTTSBackend/MLXSTTBackend dispatched model load and generate/
transcribe as separate asyncio.to_thread() calls, which round-robin
across Python's default multi-worker executor pool. MLX's Metal
backend keeps GPU streams registered per-OS-thread, so a model loaded
on one worker thread and then used for generation on a different
worker thread hits a missing stream and crashes.
Reproduced 100% of the time on macOS/Apple Silicon cloning with both
the 1.7B and 0.6B Qwen3-TTS models; Chatterbox/Kokoro were unaffected
since they use the PyTorch backend, not this module.
Fix: route all four MLX call sites in this file (TTS load, TTS
generate, STT load, STT transcribe) through a dedicated single-worker
ThreadPoolExecutor instead of asyncio.to_thread's shared pool, so
every MLX operation for a given process runs on the same OS thread.
Verified: direct /generate API calls against both model sizes
completed cleanly after the fix (previously failed every time).
A plain `docker run -v ./data:/app/data` bind mount owned by the host
user had its uid taken over (chown to 999) and its contents left
unwritable, because the adoption probe only looked at the generations
subdir, which `mkdir -p` had just created root-owned. Probe /app/data
first and fall back to generations; re-own only what the old uid owned;
chown a dir only when root created it; and let a read-only HF cache
mount through (the app only warns about it) instead of aborting on the
bare chown error.
Docker creates /home/voicebox/.cache (parent of the huggingface-cache
volume mountpoint) owned by root on every container create, and a
fresh named volume is root-owned too. The app runs as the voicebox
user (uid 999), so anything that writes a cache outside the HF mount
(torch hub, spacy, ...) fails with Permission denied.
The entrypoint already runs as root before dropping privileges via
gosu — create/chown the cache dirs there (non-recursive, instant).
Co-Authored-By: Claude Fable 5 <[email protected]>
With padding="longest" the encoder rejects inputs shorter than 3000 mel
frames when generate() has to detect the language first, so every short
clip without a forced language failed with "Whisper expects the mel input
features to be of length 3000". Use the long-form feature-extractor and
generate() options only when the audio exceeds one 30s window.
The PyTorch Whisper transcribe() called the HF processor without
truncation=False and model.generate() without return_timestamps=True.
With those defaults, WhisperFeatureExtractor silently truncates inputs
to 30 s (Whisper's native receptive field), so any dictation longer
than ~30 s lost its tail.
Setting truncation=False + padding="longest" + return_attention_mask=True
on the processor, then forwarding the attention mask plus
return_timestamps=True to generate(), flips HF Whisper into long-form
mode: autoregressive decoding over rolling 30 s windows.
Verified by round-tripping a 56.6 s Kokoro TTS sample through
/transcribe — full text returned including content past the 30 s mark;
previously the transcript was cut off roughly halfway through.
MLX backend (mlx_backend.py) is intentionally unchanged: mlx_audio.stt's
generate() already implements rolling-window long-form transcription
with condition_on_previous_text in the upstream library, so it does not
have the same bug. The HF-only kwargs added here would also break the
MLX call signature.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
Registering Kokoro with needs_trim=True routed its output through the
generic trim_tts_output, whose 1s internal-silence cut was tuned for
Chatterbox hallucinations. KPipeline synthesizes newline- and
token-limit-separated segments independently, and the ~0.3s lead plus
~0.7s tail pads at each boundary add up to 1.2s of silence, so with
af_sarah a four-paragraph script came back as its first paragraph only
(10.9s -> 1.8s) and a 51s text lost half its segments.
Drop the needs_trim flag, keep the in-backend trim, and give
trim_tts_output a max_internal_silence_ms=None mode that only trims the
leading and trailing pads. Tests updated for the new behaviour, with a
two-segment fake pipeline and an explicit internal-gap case.
Wrap empty_device_cache in release_generation_memory(), which logs and
swallows cleanup failures so a poisoned CUDA context cannot replace a
finished generation's result, and call it from the streaming endpoint
too, which drives generate_chunked directly.
Initialise tts_model before the try so the finally can read its device
without a locals() probe, import empty_device_cache alongside the other
backend imports instead of inside a bare try/except, and flatten the MPS
branch in empty_device_cache.
The file still used the pre-refactor flat imports and a sys.path hack:
sys.path.insert(0, str(Path(__file__).parent.parent))
from database import Base, VoiceProfile as DBVoiceProfile
from profiles import create_profile, update_profile
`profiles` now lives at backend/services/profiles.py, so collection raised
ImportError. Because pytest aborts the whole run on a collection error, this
one file meant `just test` ran zero tests — duplicate-name validation has had
no coverage since the services refactor.
Switch to package imports like every other test module, and drop
DBVoiceProfile, which was imported but never used.
That exposed a second, latent bug: all 6 tests passed but every one errored in
teardown with PermissionError WinError 32. The fixture closed the session and
then rmtree'd the temp dir, but closing a session does not release
SQLAlchemy's pooled connection, so SQLite still held test.db open on Windows.
Dispose the engine before removing the directory.
6 passed, and full-suite collection goes from aborting to 170 tests.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Collecting backend/tests/ fails outright on a clean checkout:
backend/tests/test_profile_duplicate_names.py:19: in <module>
from database import Base, VoiceProfile as DBVoiceProfile
E ImportError: attempted relative import beyond top-level package
pytest stops at the collection error, so the whole backend suite runs zero
tests rather than the ~142 it otherwise would.
There are two causes stacked on top of each other.
First, a real cycle in production code. routes/profiles.py, routes/history.py
and routes/stories.py each import safe_content_disposition from ..app, while
app.py builds the FastAPI instance at module scope (app = create_app() on
import), which registers those same routers. Importing any of those three
route modules first therefore re-enters a partially initialised app and dies
with "cannot import name 'router' from partially initialized module". It only
works today because app.py always happens to be imported first.
safe_content_disposition is a pure helper over urllib.parse.quote with no
application state, so it moves to backend/utils/http.py. app.py re-exports it
so any external caller importing it from the old location keeps working.
Second, the test reached for modules through a sys.path hack
(sys.path.insert(parent) + "from database import ...") rather than the
"from backend.X import ..." style the rest of the suite uses. That flat import
makes database/models.py's "from ..utils.capture_chords import ..." point
outside the package. It also aimed at the wrong module: it wants the service
layer, which raises ValueError, not the route handler, which converts that
into an HTTPException.
Result: the full suite goes from 0 collected to 148 passed. The one remaining
failure, test_progress.py::test_hf_progress_tracker, is pre-existing and
unrelated (tqdm patching) - it reproduces identically on an unpatched tree.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Update the README, MCP docs, in-app MCP page, i18n strings, landing copy,
docstrings and comments to match the renamed tools. CHANGELOG entries and
docs/plans are historical and left as they were.
Claude Desktop validates tool names against ^[a-zA-Z0-9_-]{1,64}$ and
rejects the whole tool list when any name contains a dot, so the Voicebox
MCP server was unusable there. Rename voicebox.speak, voicebox.transcribe,
voicebox.list_captures and voicebox.list_profiles to voicebox_speak,
voicebox_transcribe, voicebox_list_captures and voicebox_list_profiles.
Fixes#790
Adding seroval as a direct root dependency left the lockfile's
@tanstack/router-core/seroval entry pinned at the vulnerable 1.5.0, so the
only real consumer was still bundling the CVE-2026-59940 version. A bun
'overrides' entry forces every resolution to 1.5.3 and drops the unused
direct dependency.
Queries throughout the codebase filter on generations.profile_id,
generations.status, generations.created_at, story_items.story_id,
story_items.generation_id, generation_versions.generation_id, and
profile_samples.profile_id with every request. Without indexes SQLite
falls back to a full table scan; as history grows (hundreds or thousands
of generations) these scans become the dominant latency.
Changes:
- Add index=True on the most-queried FK and sort columns in models.py so
new installs get them from Base.metadata.create_all
- Add _migrate_add_indexes() called from run_migrations() so existing
installs get the same indexes on next startup (uses CREATE INDEX IF
NOT EXISTS — idempotent, <10 ms on any realistic dataset)
Switch from the default DELETE/ROLLBACK journal to WAL so concurrent
readers (SSE status polls, history queries) are not blocked while the
generation worker holds a write transaction. Set a 5-second busy
timeout to eliminate "database is locked" errors under brief write
contention.
Both PRAGMAs are applied via a custom creator function so every
connection in the pool gets the settings at open time, not just the
first one.
is_model_cached() marked a model as not-cached whenever any .incomplete
blob existed in its cache dir, even when a completed blob with the same
hash already sat next to it. A retried/concurrent download can leave
this orphan behind after the real transfer already finished, which made
the model appear perpetually "downloading" and re-trigger a full
re-download on every load.
Only .incomplete files with no matching completed blob now count as a
genuinely in-progress download.
Both speak surfaces built their GenerationRequest with a hardcoded "en"
fallback and never consulted the resolved profile, so a profile created
with language="fr" was still synthesised as English unless the caller
passed language= explicitly.
This hurts the MCP path most: an agent calling voicebox.speak has no way
to know the bound profile's language, so it cannot pass the argument
either. Every agent-triggered generation on a non-English profile came
out with an English accent.
The fallback chain is now explicit argument -> resolved profile's
language -> "en", which matches how engine and personality already
consult the resolved binding. The "en" backstop is kept so profiles with
no language set behave exactly as before.
Adds backend/tests/test_speak_language.py covering both surfaces: the
fallback, explicit-argument precedence, and the unchanged "en" default.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>