#1137 added tests that call the MCP tool as "voicebox.speak"; #1135 renamed
the tools to underscore names (voicebox_speak) so Claude Desktop accepts them.
Both merged cleanly but the three MCP tests failed on main with NotFoundError.
Test-only change.
Four related fixes uncovered while getting Docker running on a Linux
host without AVX-512 and with an NVIDIA GPU:
- pedalboard>=0.9.21 ships a Linux wheel with AVX-512 instructions
baked into its native extension, which SIGILLs (exit 132) on import
on any CPU without AVX-512 support. Pin below the regression until
upstream fixes it (spotify/pedalboard#454).
- The rocminfo probe in app.py ran unconditionally on every build
variant, always failing and logging on non-ROCm systems. Gate it on
/dev/kfd actually being present.
- sox (required by qwen-tts's X-vector extractor) and a C compiler
(required by Triton to JIT-compile its CUDA driver shim on first
use) were missing from the runtime image, causing hard failures
once those code paths were actually exercised.
- Named volumes and bind mounts (HF cache, app data, generated audio)
are created root-owned by Docker on first use, but the app runs as
the unprivileged voicebox user. chown the mount points in the
entrypoint before dropping privileges so this self-heals on every
start regardless of host UID.
Review follow-ups on the thread-affinity PR: MLXQwenLLMBackend.generate
awaited load_model as one worker submission and then submitted
_generate_sync as a second, leaving the same load/generate gap the TTS
and STT paths close (a concurrent load_model for another size or an
unload could swap or null out self.model/self.tokenizer in between).
Give the LLM backend the same _op_lock + single _reload_and_generate_sync
shape. The TTS/STT/LLM sync closures now bind the model to a local once,
so an inline unload_model() from the event loop mid-generation cannot
turn a later self.model read (the voice-clone fallback path in
particular) into an AttributeError.
Routing unload through the single MLX worker and waiting on the result
blocked the caller (the FastAPI event loop via /models/unload) for the
remainder of any in-flight generation — measured 3.7 s stall on a 16 s
clip on an M2 Ultra, unbounded for long texts. Dropping the model
reference inline is thread-safe (MLX frees buffers through its global
allocator) and is what main did before the thread-affinity change; the
generation in flight keeps its own reference and the next generate()
reloads on the worker.
MLX streams are thread-local and mlx-audio caches one on the model at
load time. Model load and inference were each dispatched through
asyncio.to_thread(), which uses a multi-worker pool, so load and
generate/transcribe could land on different OS threads -- the inference
thread then has no Stream(gpu, N) and MLX aborts with
"There is no Stream(gpu, 1) in current thread."
Route every MLX call (load, generate, transcribe; TTS and STT) through a
single dedicated worker thread so a model and its stream always share a
thread. max_workers=1 also serialises the single local GPU. Adds a
regression test covering the thread-affinity invariant.
Fixes#699. Also addresses #675.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
The dedicated-MLX-thread fix (#886) covers the TTS and STT backends, but
MLXQwenLLMBackend still dispatches load/generate through
asyncio.to_thread(), so /llm/generate (personality compose/rewrite,
dictation refinement) crashes with the same
'There is no Stream(gpu, 0) in current thread.' — raised from
mlx_lm.generate's wired_limit on exit — whenever load and generate land
on different pool threads.
Route the MLX LLM backend's unload/load/generate through the same
_run_on_mlx_thread helper so every MLX call in the process shares one
worker thread.
Verified on Apple M3 Pro (macOS 26.5, mlx 0.32.0, mlx-lm 0.31.1,
mlx-audio 0.4.1): /llm/generate failed 100% before, succeeds after;
TTS + STT unaffected.
Co-Authored-By: Claude Fable 5 <[email protected]>
CodeRabbit follow-up on the previous round: the asyncio.Lock closed the
gap for callers going through generate()/transcribe() consistently, but
unload_model() itself never touched that lock and could still land
between the load future resolving and the generate/transcribe future
being submitted - two separate _run_on_mlx_thread calls, so a real gap
existed at the Python level even though the executor is single-worker.
Fix: generate() and transcribe() each now submit exactly ONE callable
to the MLX executor - a self-healing reload-if-needed-then-infer
function - instead of a load submission followed by a separate infer
submission. This removes the gap structurally: nothing can observe an
intermediate state because there is no intermediate state exposed
across an await boundary. unload_model() now also submits
unconditionally (loaded-check moved inside _unload_model_sync, which
runs atomically with the teardown) rather than racing its own
Python-level self.model check against a load in flight.
Verified: same-size generate, size-switch generate, unload route, and
a post-unload generate all complete clean on the live launchd server.
Two real follow-on issues found by CodeRabbit on PR #989:
1. MLXTTSBackend.load_model_async called self.unload_model() directly
on the event-loop thread before dispatching _load_model_sync to the
MLX worker thread - model teardown ran on the wrong OS thread, same
class of bug the PR itself fixes. Combined unload+load into one
_reload_sync callable submitted as a single MLX-thread operation.
2. Public unload_model() (called synchronously from the /models/unload
routes) also ran on the caller thread. Now submits to the MLX
executor and blocks on the result, so teardown always happens on the
worker thread regardless of caller. Same fix applied to
MLXSTTBackend.
3. Both backends cache self.model on the instance, but generate()/
transcribe() only locked their own internal load step - a
concurrent request for a different model_size could swap self.model
between one request's load and its inference (real race: routes/
generations.py calls load_engine_model() and generate_chunked() as
separate awaited steps with a gap between them). Added a per-backend
asyncio.Lock held across the full load+inference sequence in both
generate() and transcribe(). Note: this closes the race for callers
using the backend's own public methods consistently; the wider
route-level orchestration race (load_engine_model + generate_chunked
as two separate calls) is a follow-up outside this file's scope.
Verified: same-size and cross-size-switch generations both complete
clean after the patch; /models/{name}/unload route returns 200 without
deadlocking the (still-responsive) server.
Qwen3-TTS (and MLX STT) generation crashed with:
"There is no Stream(gpu, N) in current thread."
MLXTTSBackend/MLXSTTBackend dispatched model load and generate/
transcribe as separate asyncio.to_thread() calls, which round-robin
across Python's default multi-worker executor pool. MLX's Metal
backend keeps GPU streams registered per-OS-thread, so a model loaded
on one worker thread and then used for generation on a different
worker thread hits a missing stream and crashes.
Reproduced 100% of the time on macOS/Apple Silicon cloning with both
the 1.7B and 0.6B Qwen3-TTS models; Chatterbox/Kokoro were unaffected
since they use the PyTorch backend, not this module.
Fix: route all four MLX call sites in this file (TTS load, TTS
generate, STT load, STT transcribe) through a dedicated single-worker
ThreadPoolExecutor instead of asyncio.to_thread's shared pool, so
every MLX operation for a given process runs on the same OS thread.
Verified: direct /generate API calls against both model sizes
completed cleanly after the fix (previously failed every time).
With padding="longest" the encoder rejects inputs shorter than 3000 mel
frames when generate() has to detect the language first, so every short
clip without a forced language failed with "Whisper expects the mel input
features to be of length 3000". Use the long-form feature-extractor and
generate() options only when the audio exceeds one 30s window.
The PyTorch Whisper transcribe() called the HF processor without
truncation=False and model.generate() without return_timestamps=True.
With those defaults, WhisperFeatureExtractor silently truncates inputs
to 30 s (Whisper's native receptive field), so any dictation longer
than ~30 s lost its tail.
Setting truncation=False + padding="longest" + return_attention_mask=True
on the processor, then forwarding the attention mask plus
return_timestamps=True to generate(), flips HF Whisper into long-form
mode: autoregressive decoding over rolling 30 s windows.
Verified by round-tripping a 56.6 s Kokoro TTS sample through
/transcribe — full text returned including content past the 30 s mark;
previously the transcript was cut off roughly halfway through.
MLX backend (mlx_backend.py) is intentionally unchanged: mlx_audio.stt's
generate() already implements rolling-window long-form transcription
with condition_on_previous_text in the upstream library, so it does not
have the same bug. The HF-only kwargs added here would also break the
MLX call signature.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
Registering Kokoro with needs_trim=True routed its output through the
generic trim_tts_output, whose 1s internal-silence cut was tuned for
Chatterbox hallucinations. KPipeline synthesizes newline- and
token-limit-separated segments independently, and the ~0.3s lead plus
~0.7s tail pads at each boundary add up to 1.2s of silence, so with
af_sarah a four-paragraph script came back as its first paragraph only
(10.9s -> 1.8s) and a 51s text lost half its segments.
Drop the needs_trim flag, keep the in-backend trim, and give
trim_tts_output a max_internal_silence_ms=None mode that only trims the
leading and trailing pads. Tests updated for the new behaviour, with a
two-segment fake pipeline and an explicit internal-gap case.
Wrap empty_device_cache in release_generation_memory(), which logs and
swallows cleanup failures so a poisoned CUDA context cannot replace a
finished generation's result, and call it from the streaming endpoint
too, which drives generate_chunked directly.
Initialise tts_model before the try so the finally can read its device
without a locals() probe, import empty_device_cache alongside the other
backend imports instead of inside a bare try/except, and flatten the MPS
branch in empty_device_cache.
The file still used the pre-refactor flat imports and a sys.path hack:
sys.path.insert(0, str(Path(__file__).parent.parent))
from database import Base, VoiceProfile as DBVoiceProfile
from profiles import create_profile, update_profile
`profiles` now lives at backend/services/profiles.py, so collection raised
ImportError. Because pytest aborts the whole run on a collection error, this
one file meant `just test` ran zero tests — duplicate-name validation has had
no coverage since the services refactor.
Switch to package imports like every other test module, and drop
DBVoiceProfile, which was imported but never used.
That exposed a second, latent bug: all 6 tests passed but every one errored in
teardown with PermissionError WinError 32. The fixture closed the session and
then rmtree'd the temp dir, but closing a session does not release
SQLAlchemy's pooled connection, so SQLite still held test.db open on Windows.
Dispose the engine before removing the directory.
6 passed, and full-suite collection goes from aborting to 170 tests.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Collecting backend/tests/ fails outright on a clean checkout:
backend/tests/test_profile_duplicate_names.py:19: in <module>
from database import Base, VoiceProfile as DBVoiceProfile
E ImportError: attempted relative import beyond top-level package
pytest stops at the collection error, so the whole backend suite runs zero
tests rather than the ~142 it otherwise would.
There are two causes stacked on top of each other.
First, a real cycle in production code. routes/profiles.py, routes/history.py
and routes/stories.py each import safe_content_disposition from ..app, while
app.py builds the FastAPI instance at module scope (app = create_app() on
import), which registers those same routers. Importing any of those three
route modules first therefore re-enters a partially initialised app and dies
with "cannot import name 'router' from partially initialized module". It only
works today because app.py always happens to be imported first.
safe_content_disposition is a pure helper over urllib.parse.quote with no
application state, so it moves to backend/utils/http.py. app.py re-exports it
so any external caller importing it from the old location keeps working.
Second, the test reached for modules through a sys.path hack
(sys.path.insert(parent) + "from database import ...") rather than the
"from backend.X import ..." style the rest of the suite uses. That flat import
makes database/models.py's "from ..utils.capture_chords import ..." point
outside the package. It also aimed at the wrong module: it wants the service
layer, which raises ValueError, not the route handler, which converts that
into an HTTPException.
Result: the full suite goes from 0 collected to 148 passed. The one remaining
failure, test_progress.py::test_hf_progress_tracker, is pre-existing and
unrelated (tqdm patching) - it reproduces identically on an unpatched tree.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Update the README, MCP docs, in-app MCP page, i18n strings, landing copy,
docstrings and comments to match the renamed tools. CHANGELOG entries and
docs/plans are historical and left as they were.
Claude Desktop validates tool names against ^[a-zA-Z0-9_-]{1,64}$ and
rejects the whole tool list when any name contains a dot, so the Voicebox
MCP server was unusable there. Rename voicebox.speak, voicebox.transcribe,
voicebox.list_captures and voicebox.list_profiles to voicebox_speak,
voicebox_transcribe, voicebox_list_captures and voicebox_list_profiles.
Fixes#790
Queries throughout the codebase filter on generations.profile_id,
generations.status, generations.created_at, story_items.story_id,
story_items.generation_id, generation_versions.generation_id, and
profile_samples.profile_id with every request. Without indexes SQLite
falls back to a full table scan; as history grows (hundreds or thousands
of generations) these scans become the dominant latency.
Changes:
- Add index=True on the most-queried FK and sort columns in models.py so
new installs get them from Base.metadata.create_all
- Add _migrate_add_indexes() called from run_migrations() so existing
installs get the same indexes on next startup (uses CREATE INDEX IF
NOT EXISTS — idempotent, <10 ms on any realistic dataset)
Switch from the default DELETE/ROLLBACK journal to WAL so concurrent
readers (SSE status polls, history queries) are not blocked while the
generation worker holds a write transaction. Set a 5-second busy
timeout to eliminate "database is locked" errors under brief write
contention.
Both PRAGMAs are applied via a custom creator function so every
connection in the pool gets the settings at open time, not just the
first one.
is_model_cached() marked a model as not-cached whenever any .incomplete
blob existed in its cache dir, even when a completed blob with the same
hash already sat next to it. A retried/concurrent download can leave
this orphan behind after the real transfer already finished, which made
the model appear perpetually "downloading" and re-trigger a full
re-download on every load.
Only .incomplete files with no matching completed blob now count as a
genuinely in-progress download.
Both speak surfaces built their GenerationRequest with a hardcoded "en"
fallback and never consulted the resolved profile, so a profile created
with language="fr" was still synthesised as English unless the caller
passed language= explicitly.
This hurts the MCP path most: an agent calling voicebox.speak has no way
to know the bound profile's language, so it cannot pass the argument
either. Every agent-triggered generation on a non-English profile came
out with an English accent.
The fallback chain is now explicit argument -> resolved profile's
language -> "en", which matches how engine and personality already
consult the resolved binding. The "en" backstop is kept so profiles with
no language set behave exactly as before.
Adds backend/tests/test_speak_language.py covering both surfaces: the
fallback, explicit-argument precedence, and the unchanged "en" default.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Without a trace the Settings GPU label and /health keep reporting CUDA
while generation runs on CPU, so the user has no way to confirm the
override engaged.
The override is documented in gpu-acceleration.mdx and listed as step 1 of
the get_torch_device() precedence in tts-generation.mdx, but grepping the
tree for VOICEBOX_FORCE_CPU matched only those two doc files - nothing read
it. Users whose GPU has no compiled kernels in the bundled PyTorch had no
way to fall back to CPU short of renaming the installed CUDA backend
directory.
Resolve it before torch is imported, so it still works when the installed
build is itself the reason CPU is wanted.
list_generations() fetched versions with one SELECT per generation on the
page (50 rows -> 51 queries). Add _get_versions_for_generations() which
loads all versions for the page in a single WHERE generation_id IN (...)
query and groups them in memory; the single-generation helper now
delegates to it so story item details behave identically.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <[email protected]>
When the server outlives the Tauri app that spawned it (keep-running
mode, or a sidecar the next launch reuses), its stdout/stderr pipe has
no reader. Every later print()/tqdm write raises BrokenPipeError, so
POST /captures and /transcribe return "[Errno 32] Broken pipe".
Wrap stdout/stderr so they fall back to devnull on the first failed
write instead of raising.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The load body now runs with HF_HUB_OFFLINE forced when _is_model_cached()
reports True, so a snapshot that holds the weights but not the small files
the loader also opens would fail hard instead of fetching them. Verified
against chatterbox-tts 0.1.7 (mtl_tts.py allow_patterns, tts_turbo.py
from_local) and the TADA loader's unsloth/Llama-3.2-1B tokenizer download;
list those files in the required_files checks. Also note in the changelog
why the 0.4.5 removal of this guard no longer applies.