Commit Graph
344 Commits
Author SHA1 Message Date
jamiepineandcapy-ai-staging[bot] d4a16292c8 fix(routes): drop the dead aux-repo check from the model-status fallback handler 2026-10-04 00:25:58 +00:00
jamiepineandcapy-ai-staging[bot] a2b453afed fix(backend): keep the trimmed clip when a runaway retry cannot split further; track the S3 tokenizer repo in model status and delete
Review follow-ups: with retries_runaway on for an engine that also has a
trim step, a <=100-char chunk flagged as runaway raised instead of
falling back to the trimmed output that already cuts the silence+noise
tail; generate_chunked now returns the trimmed chunk in that terminal
case. ModelConfig gains aux_hf_repo_ids so /models status only reports
the Chatterbox MLX model as downloaded once mlx-community/S3TokenizerV2
is present too (matching the backend's own cache check) and
DELETE /models/{name} removes that repo as well.
2026-10-04 00:25:58 +00:00
jamiepineandcapy-ai-staging[bot] ad6ec3c6ef fix(backend): count the S3TokenizerV2 repo in the Chatterbox MLX cache check; register the backend with PyInstaller
Review follow-ups: mlx-audio's Model.from_pretrained fetches the S3
speech tokenizer from mlx-community/S3TokenizerV2 (~470 MB) separately
from the chatterbox checkout, so _is_model_cached now requires both
repos (same shape as the Hume backend's codec check) and the config's
size_mb reflects the real footprint. backend.backends.chatterbox_mlx_backend
is a function-level import that PyInstaller's graph will not see, so it
is added to the Apple Silicon hidden-import list in build_binary.py and
voicebox-server.spec next to mlx_backend.
2026-10-04 00:25:58 +00:00
jamiepineandcapy-ai-staging[bot] c5cf7436f7 fix(backend): run Chatterbox MLX load and generate on the shared MLX worker thread
With asyncio.to_thread the new backend hits the same thread-local stream
crash as the Qwen MLX backend once the default pool is busy (reproduced
on an M2 Ultra: 'There is no Stream(cpu, 2) in current thread.' on the
first generate after load). Route load and generate through
mlx_backend._run_on_mlx_thread, fold load-if-needed + generate into one
worker submission under the backend's lock, and bind the model locally
inside the closure, matching MLXTTSBackend.
2026-10-04 00:25:58 +00:00
Charles Hasseandcapy-ai-staging[bot] 502123a943 fix(backend): retry runaway generations on the Chatterbox MLX backend
The qwen configs already set retries_runaway on MLX, with the reason in a comment:
mlx-audio can continue past an EOS miss and emit silence followed by codec noise.
The Chatterbox MLX backend added in this PR hits the same failure and did not have
the guard wired.

Observed on an M4 Max with a cloned pt-BR profile: a 145 character sentence took
18.0s and returned 5.3s of audio for text worth about 8s, and a short reply came
back as an endless hiss. With retries_runaway enabled the same sentence takes 3.4s
and returns the full 7.4s of speech.
2026-10-04 00:25:58 +00:00
Charles Hasseandcapy-ai-staging[bot] ad5d64c0a3 feat(backend): run Chatterbox multilingual on MLX for Apple Silicon
The Chatterbox backend is pinned to the CPU on macOS, so voice cloning runs at
roughly 4x realtime there. This adds an MLX/Metal backend for the same engine and
selects it on Apple Silicon, mirroring the split the qwen engine already makes
between mlx_backend and pytorch_backend.

Measured on an M4 Max (36 GB) with a cloned pt-BR profile, same API, same profile,
model already loaded:

  short sentence (1.8s of audio):  7.4-9.2s  ->  1.2s
  longer sentence (5.0s of audio): 23.7s     ->  3.1s

The model config for chatterbox-tts is now backend aware, same as the qwen configs,
so the download matches the backend that will consume it.

Nothing changes off Apple Silicon: the PyTorch backend is still selected there, and
the CPU pinning it relies on is untouched.
2026-10-04 00:25:58 +00:00
jamiepineandcapy-ai-staging[bot] 63899fd865 fix(mlx): release the model binding inside the sync bodies so a failing op's traceback cannot pin it
Review follow-up: the post-op drain ran in a finally, but on the error
path the exception's traceback kept the _generate_sync/_transcribe_sync
frame (and its model local) alive, so mx.clear_cache() had nothing to
return. Each sync body now drops its model (and tokenizer) binding in its
own finally before the exception propagates.
2026-10-04 00:25:53 +00:00
jamiepineandcapy-ai-staging[bot] fb67ad436c fix(mlx): drain the pool in a finally so a failing in-flight op releases it too; fix the mid-generation test's call count
Review follow-ups: the post-op drain now runs in a finally block for the
TTS, STT and LLM folded callables, so an exception escaping the generation
(e.g. the voice-clone fallback failing) still releases the buffers the
worker held. The mid-generation test asserted one drain call but the
path legitimately produces two (unload_model's own and the post-op one);
a second test covers the error path.
2026-10-04 00:25:53 +00:00
jamiepineandcapy-ai-staging[bot] cffe24ffd2 fix(mlx): drain the MLX pool after an in-flight op when unload landed mid-generation
Review follow-up: unload_model() runs inline on the event loop, so when it
lands during a generation it only drops the backend's reference; the
worker's local keeps the model alive and mx.clear_cache() finds nothing to
free. Once the folded load-and-generate (or transcribe) callable finishes
and releases that local, check whether the model was unloaded meanwhile
and drain the pool then. Covered by a test that unloads from inside a
fake model.generate().
2026-10-04 00:25:53 +00:00
jamiepineandcapy-ai-staging[bot] b1323ed0a8 fix(backend): drop the prompt cache before the allocator-empty; note clear_cache is thread-agnostic
Review follow-ups: clear_voice_prompt_memory_cache() now runs before
backend.unload_model() at all three call sites so the device-resident
prompt tensors are already released when empty_device_cache() /
empty_mlx_cache() runs, instead of going back into the caching allocator
afterwards. empty_mlx_cache's docstring states why it may run off the
MLX worker thread.
2026-10-04 00:25:53 +00:00
JnyRoadandcapy-ai-staging[bot] 1a803aa05f fix(backend): suppress E402 for intentionally-late MLX imports
CodeRabbit flagged that the imports touched in the previous commit
(TTSBackend, .base, ..utils.cache) trigger Ruff's E402 check because
they must come after patch_huggingface_hub_offline() /
ensure_original_qwen_config_cached() run — reordering them would
defeat the offline-patch-before-import guarantee the comment above
describes. Add narrow `# noqa: E402` to the four import statements
that intentionally follow those calls, without touching unrelated
pre-existing lint findings in the file.
2026-10-04 00:25:53 +00:00
JnyRoadandcapy-ai-staging[bot] 19f8f51408 fix(backend): release memory when unloading MLX models
Unloading a TTS/Whisper/LLM model on the MLX backend only dropped the
Python reference (`del self.model`). MLX keeps freed array buffers in
its own allocator pool for reuse instead of returning them to the OS,
so the process's memory footprint never actually shrank after unload
on Apple Silicon (the default backend there) until the process exited.

Add empty_mlx_cache() (backend/backends/base.py), wrapping
mx.clear_cache(), and call it from the three MLX unload_model()
implementations: MLXTTSBackend, MLXSTTBackend, MLXQwenLLMBackend.

Separately, the voice-clone prompt cache (backend/utils/cache.py) is a
process-lifetime dict populated by create_voice_prompt() across every
TTS engine, but nothing ever cleared it on model unload — only the
unrelated /tasks/clear-cache endpoint touched it. Add
clear_voice_prompt_memory_cache() (memory only, disk cache untouched
so a later generation still reloads the prompt instead of recomputing
it) and wire it into every TTS unload path (services/tts.py and the
qwen_custom_voice / generic branches of unload_model_by_config).
Whisper and the LLM backends never produce voice prompts, so their
unload paths are left alone.

Testing:
- New unit tests: backend/tests/test_mlx_unload_clears_cache.py,
  backend/tests/test_voice_prompt_cache_unload.py (8 tests, all pass).
- Verified end-to-end on Apple Silicon against real cached models
  (Qwen TTS 1.7B, Whisper Turbo, Qwen3 0.6B): loaded each via the
  running app, unloaded via the real /models/{name}/unload endpoint,
  and confirmed via mx.get_cache_memory()/get_active_memory() that the
  MLX allocator's cache drops to 0 on every cycle. Ran a real
  voice-clone generation end to end and confirmed the in-memory prompt
  cache goes from 1 entry to 0 on unload while the on-disk .prompt
  file is left intact.
2026-10-04 00:25:53 +00:00
jamiepineandcapy-ai-staging[bot] 5803aaaa91 docs(history): fix stale comment about rollback on delete 2026-10-04 00:19:41 +00:00
jamiepineandcapy-ai-staging[bot] 7e9e88de18 fix(history): don't roll back a delete after version files are already unlinked
Review follow-up: a rollback after a failed main-audio unlink restored
version rows whose files were already removed. Treat a locked main file
like the failed-generation sweep does: log and continue, leaving at worst
an orphaned audio file rather than rows that point at nothing.
2026-10-04 00:19:41 +00:00
jamiepineandcapy-ai-staging[bot] 755c664e66 fix(history): commit generation delete cascade once, roll back on unlink failure
Review follow-up: _delete_generation_children defaulted to committing the
story-item and version deletes before the main audio unlink, so a locked
file left the generation row in place with its story items already gone.
Both single-delete and the failed-generation sweep now pass commit=False
and commit once with the generation row; a failed unlink rolls back.
2026-10-04 00:19:41 +00:00
jamiepineandcapy-ai-staging[bot] 8e0060cccf test(profiles): reword the locked-file test docstring to the guarantee it checks 2026-10-04 00:19:41 +00:00
jamiepineandcapy-ai-staging[bot] f0f3802b79 fix(profiles): commit the whole profile delete cascade once
delete_generations_by_profile (and the version sweep under it) committed on
their own, so a failure between the generation sweep and the profile row
left the profile listed with its history already gone. Both helpers take a
commit flag now, and delete_profile passes commit=False so the single
db.commit() at its end publishes the entire cascade. Other callers keep the
default and are unchanged.
2026-10-04 00:19:41 +00:00
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] 64b4b976d6 fix(versions): don't let a locked version file abort the delete cascade 2026-10-04 00:19:41 +00:00
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] 4750b47b04 fix(profiles): delete a profile's generations instead of orphaning them 2026-10-04 00:19:41 +00:00
jamiepineandcapy-ai-staging[bot] 4f54e16824 test(speak): call the MCP tool by its underscore name
#1137 added tests that call the MCP tool as "voicebox.speak"; #1135 renamed
the tools to underscore names (voicebox_speak) so Claude Desktop accepts them.
Both merged cleanly but the three MCP tests failed on main with NotFoundError.
Test-only change.
2026-10-04 00:15:34 +00:00
harryandcapy-ai-staging[bot] ab9a19790c fix(docker): fix startup crash and permission errors in the container
Four related fixes uncovered while getting Docker running on a Linux
host without AVX-512 and with an NVIDIA GPU:

- pedalboard>=0.9.21 ships a Linux wheel with AVX-512 instructions
  baked into its native extension, which SIGILLs (exit 132) on import
  on any CPU without AVX-512 support. Pin below the regression until
  upstream fixes it (spotify/pedalboard#454).
- The rocminfo probe in app.py ran unconditionally on every build
  variant, always failing and logging on non-ROCm systems. Gate it on
  /dev/kfd actually being present.
- sox (required by qwen-tts's X-vector extractor) and a C compiler
  (required by Triton to JIT-compile its CUDA driver shim on first
  use) were missing from the runtime image, causing hard failures
  once those code paths were actually exercised.
- Named volumes and bind mounts (HF cache, app data, generated audio)
  are created root-owned by Docker on first use, but the app runs as
  the unprivileged voicebox user. chown the mount points in the
  entrypoint before dropping privileges so this self-heals on every
  start regardless of host UID.
2026-10-04 00:12:29 +00:00
capy-ai-staging[bot]andGitHub fee218ef0b Merge pull request #1130 from jamiepine/prep/pr-1112
Fix infinite HF retry storm when loading a cached model offline
2026-10-04 00:05:46 +00:00
capy-ai-staging[bot]andGitHub 39b1c91291 Merge pull request #1144 from jamiepine/prep/pr-1031
fix(setup): pin dev venv to Python 3.12
2026-10-04 00:05:41 +00:00
capy-ai-staging[bot]andGitHub 3221b2796b Merge pull request #1146 from jamiepine/prep/pr-659
fix: replace deprecated datetime.utcnow() with datetime.now(UTC) throughout
2026-10-04 00:05:36 +00:00
jamiepine ba942ec1fd Merge remote-tracking branch 'origin/main' into prep/pr-659
# Conflicts:
#	backend/database/models.py
2026-10-04 00:02:23 +00:00
jamiepineandcapy-ai-staging[bot] dd8ab5cb20 fix(mlx): fold Qwen3 LLM load+generate into one worker submission; bind model locally in generate closures
Review follow-ups on the thread-affinity PR: MLXQwenLLMBackend.generate
awaited load_model as one worker submission and then submitted
_generate_sync as a second, leaving the same load/generate gap the TTS
and STT paths close (a concurrent load_model for another size or an
unload could swap or null out self.model/self.tokenizer in between).
Give the LLM backend the same _op_lock + single _reload_and_generate_sync
shape. The TTS/STT/LLM sync closures now bind the model to a local once,
so an inline unload_model() from the event loop mid-generation cannot
turn a later self.model read (the voice-clone fallback path in
particular) into an AttributeError.
2026-10-04 00:01:47 +00:00
jamiepineandcapy-ai-staging[bot] cf984885e9 fix(mlx): keep unload_model off the MLX worker so it cannot stall the event loop
Routing unload through the single MLX worker and waiting on the result
blocked the caller (the FastAPI event loop via /models/unload) for the
remainder of any in-flight generation — measured 3.7 s stall on a 16 s
clip on an M2 Ultra, unbounded for long texts. Dropping the model
reference inline is thread-safe (MLX frees buffers through its global
allocator) and is what main did before the thread-affinity change; the
generation in flight keeps its own reference and the next generate()
reloads on the worker.
2026-10-04 00:01:47 +00:00
99d05b917c fix(backends): run MLX load and inference on one thread
MLX streams are thread-local and mlx-audio caches one on the model at
load time. Model load and inference were each dispatched through
asyncio.to_thread(), which uses a multi-worker pool, so load and
generate/transcribe could land on different OS threads -- the inference
thread then has no Stream(gpu, N) and MLX aborts with
"There is no Stream(gpu, 1) in current thread."

Route every MLX call (load, generate, transcribe; TTS and STT) through a
single dedicated worker thread so a model and its stream always share a
thread. max_workers=1 also serialises the single local GPU. Adds a
regression test covering the thread-affinity invariant.

Fixes #699. Also addresses #675.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-10-04 00:01:47 +00:00
2693c983f2 fix: extend MLX thread-affinity fix to the Qwen3 LLM backend
The dedicated-MLX-thread fix (#886) covers the TTS and STT backends, but
MLXQwenLLMBackend still dispatches load/generate through
asyncio.to_thread(), so /llm/generate (personality compose/rewrite,
dictation refinement) crashes with the same
'There is no Stream(gpu, 0) in current thread.' — raised from
mlx_lm.generate's wired_limit on exit — whenever load and generate land
on different pool threads.

Route the MLX LLM backend's unload/load/generate through the same
_run_on_mlx_thread helper so every MLX call in the process shares one
worker thread.

Verified on Apple M3 Pro (macOS 26.5, mlx 0.32.0, mlx-lm 0.31.1,
mlx-audio 0.4.1): /llm/generate failed 100% before, succeeds after;
TTS + STT unaffected.

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-10-04 00:01:47 +00:00
Ron David Ben Ishayandcapy-ai-staging[bot] 57ef6636bc fix(mlx): fold reload+inference into one MLX-worker submission
CodeRabbit follow-up on the previous round: the asyncio.Lock closed the
gap for callers going through generate()/transcribe() consistently, but
unload_model() itself never touched that lock and could still land
between the load future resolving and the generate/transcribe future
being submitted - two separate _run_on_mlx_thread calls, so a real gap
existed at the Python level even though the executor is single-worker.

Fix: generate() and transcribe() each now submit exactly ONE callable
to the MLX executor - a self-healing reload-if-needed-then-infer
function - instead of a load submission followed by a separate infer
submission. This removes the gap structurally: nothing can observe an
intermediate state because there is no intermediate state exposed
across an await boundary. unload_model() now also submits
unconditionally (loaded-check moved inside _unload_model_sync, which
runs atomically with the teardown) rather than racing its own
Python-level self.model check against a load in flight.

Verified: same-size generate, size-switch generate, unload route, and
a post-unload generate all complete clean on the live launchd server.
2026-10-04 00:01:47 +00:00
Ron David Ben Ishayandcapy-ai-staging[bot] fa8db820d7 fix(mlx): address review feedback on the thread-affinity fix
Two real follow-on issues found by CodeRabbit on PR #989:

1. MLXTTSBackend.load_model_async called self.unload_model() directly
   on the event-loop thread before dispatching _load_model_sync to the
   MLX worker thread - model teardown ran on the wrong OS thread, same
   class of bug the PR itself fixes. Combined unload+load into one
   _reload_sync callable submitted as a single MLX-thread operation.

2. Public unload_model() (called synchronously from the /models/unload
   routes) also ran on the caller thread. Now submits to the MLX
   executor and blocks on the result, so teardown always happens on the
   worker thread regardless of caller. Same fix applied to
   MLXSTTBackend.

3. Both backends cache self.model on the instance, but generate()/
   transcribe() only locked their own internal load step - a
   concurrent request for a different model_size could swap self.model
   between one request's load and its inference (real race: routes/
   generations.py calls load_engine_model() and generate_chunked() as
   separate awaited steps with a gap between them). Added a per-backend
   asyncio.Lock held across the full load+inference sequence in both
   generate() and transcribe(). Note: this closes the race for callers
   using the backend's own public methods consistently; the wider
   route-level orchestration race (load_engine_model + generate_chunked
   as two separate calls) is a follow-up outside this file's scope.

Verified: same-size and cross-size-switch generations both complete
clean after the patch; /models/{name}/unload route returns 200 without
deadlocking the (still-responsive) server.
2026-10-04 00:01:47 +00:00
Ron David Ben Ishayandcapy-ai-staging[bot] f135a471ae fix(mlx): pin all MLX ops to a single worker thread
Qwen3-TTS (and MLX STT) generation crashed with:
"There is no Stream(gpu, N) in current thread."

MLXTTSBackend/MLXSTTBackend dispatched model load and generate/
transcribe as separate asyncio.to_thread() calls, which round-robin
across Python's default multi-worker executor pool. MLX's Metal
backend keeps GPU streams registered per-OS-thread, so a model loaded
on one worker thread and then used for generation on a different
worker thread hits a missing stream and crashes.

Reproduced 100% of the time on macOS/Apple Silicon cloning with both
the 1.7B and 0.6B Qwen3-TTS models; Chatterbox/Kokoro were unaffected
since they use the PyTorch backend, not this module.

Fix: route all four MLX call sites in this file (TTS load, TTS
generate, STT load, STT transcribe) through a dedicated single-worker
ThreadPoolExecutor instead of asyncio.to_thread's shared pool, so
every MLX operation for a given process runs on the same OS thread.

Verified: direct /generate API calls against both model sizes
completed cleanly after the fix (previously failed every time).
2026-10-04 00:01:47 +00:00
jamiepineandcapy-ai-staging[bot] 86dc46b930 refactor(stt): name the Whisper sample rate and window constants 2026-10-04 00:01:34 +00:00
jamiepineandcapy-ai-staging[bot] 51882b5065 fix(backend): keep the pad-to-30s Whisper path for clips under 30s
With padding="longest" the encoder rejects inputs shorter than 3000 mel
frames when generate() has to detect the language first, so every short
clip without a forced language failed with "Whisper expects the mel input
features to be of length 3000". Use the long-form feature-extractor and
generate() options only when the audio exceeds one 30s window.
2026-10-04 00:01:34 +00:00
noxandcapy-ai-staging[bot] 7575a65e7f fix(backend): pass language/task directly to Whisper generate 2026-10-04 00:01:34 +00:00
e61c85365b fix(backend): enable long-form Whisper transcription on PyTorch path
The PyTorch Whisper transcribe() called the HF processor without
truncation=False and model.generate() without return_timestamps=True.
With those defaults, WhisperFeatureExtractor silently truncates inputs
to 30 s (Whisper's native receptive field), so any dictation longer
than ~30 s lost its tail.

Setting truncation=False + padding="longest" + return_attention_mask=True
on the processor, then forwarding the attention mask plus
return_timestamps=True to generate(), flips HF Whisper into long-form
mode: autoregressive decoding over rolling 30 s windows.

Verified by round-tripping a 56.6 s Kokoro TTS sample through
/transcribe — full text returned including content past the 30 s mark;
previously the transcript was cut off roughly halfway through.

MLX backend (mlx_backend.py) is intentionally unchanged: mlx_audio.stt's
generate() already implements rolling-window long-form transcription
with condition_on_previous_text in the upstream library, so it does not
have the same bug. The HF-only kwargs added here would also break the
MLX call signature.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
2026-10-04 00:01:34 +00:00
jamiepineandcapy-ai-staging[bot] ae300c5316 fix(kokoro): trim edges only, keep inter-segment gaps
Registering Kokoro with needs_trim=True routed its output through the
generic trim_tts_output, whose 1s internal-silence cut was tuned for
Chatterbox hallucinations. KPipeline synthesizes newline- and
token-limit-separated segments independently, and the ~0.3s lead plus
~0.7s tail pads at each boundary add up to 1.2s of silence, so with
af_sarah a four-paragraph script came back as its first paragraph only
(10.9s -> 1.8s) and a 51s text lost half its segments.

Drop the needs_trim flag, keep the in-backend trim, and give
trim_tts_output a max_internal_silence_ms=None mode that only trims the
leading and trailing pads. Tests updated for the new behaviour, with a
two-segment fake pipeline and an explicit internal-gap case.
2026-10-04 00:01:24 +00:00
devangkanthariaandcapy-ai-staging[bot] 615aeaeb35 fix(kokoro): trim trailing silence and run-on noise on short prompt synthesis (#960) 2026-10-04 00:01:24 +00:00
jamiepineandcapy-ai-staging[bot] b788dc383c fix(generation): make post-generation cleanup best-effort and cover /generate/stream
Wrap empty_device_cache in release_generation_memory(), which logs and
swallows cleanup failures so a poisoned CUDA context cannot replace a
finished generation's result, and call it from the streaming endpoint
too, which drives generate_chunked directly.
2026-10-04 00:01:18 +00:00
jamiepineandcapy-ai-staging[bot] 17fd1ddd1b refactor(generation): simplify post-generation cache cleanup
Initialise tts_model before the try so the finally can read its device
without a locals() probe, import empty_device_cache alongside the other
backend imports instead of inside a bare try/except, and flatten the MPS
branch in empty_device_cache.
2026-10-04 00:01:18 +00:00
devangkanthariaandcapy-ai-staging[bot] 38389db82b fix(backend): prevent unbounded memory accumulation over consecutive TTS generations (#923) 2026-10-04 00:01:18 +00:00
63f09455ef fix(tests): repair test_profile_duplicate_names so the suite can run
The file still used the pre-refactor flat imports and a sys.path hack:

    sys.path.insert(0, str(Path(__file__).parent.parent))
    from database import Base, VoiceProfile as DBVoiceProfile
    from profiles import create_profile, update_profile

`profiles` now lives at backend/services/profiles.py, so collection raised
ImportError. Because pytest aborts the whole run on a collection error, this
one file meant `just test` ran zero tests — duplicate-name validation has had
no coverage since the services refactor.

Switch to package imports like every other test module, and drop
DBVoiceProfile, which was imported but never used.

That exposed a second, latent bug: all 6 tests passed but every one errored in
teardown with PermissionError WinError 32. The fixture closed the session and
then rmtree'd the temp dir, but closing a session does not release
SQLAlchemy's pooled connection, so SQLite still held test.db open on Windows.
Dispose the engine before removing the directory.

6 passed, and full-suite collection goes from aborting to 170 tests.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:01:11 +00:00
3e240a9ef1 fix(backend): break the app<->routes import cycle that aborts the test suite
Collecting backend/tests/ fails outright on a clean checkout:

    backend/tests/test_profile_duplicate_names.py:19: in <module>
        from database import Base, VoiceProfile as DBVoiceProfile
    E   ImportError: attempted relative import beyond top-level package

pytest stops at the collection error, so the whole backend suite runs zero
tests rather than the ~142 it otherwise would.

There are two causes stacked on top of each other.

First, a real cycle in production code. routes/profiles.py, routes/history.py
and routes/stories.py each import safe_content_disposition from ..app, while
app.py builds the FastAPI instance at module scope (app = create_app() on
import), which registers those same routers. Importing any of those three
route modules first therefore re-enters a partially initialised app and dies
with "cannot import name 'router' from partially initialized module". It only
works today because app.py always happens to be imported first.

safe_content_disposition is a pure helper over urllib.parse.quote with no
application state, so it moves to backend/utils/http.py. app.py re-exports it
so any external caller importing it from the old location keeps working.

Second, the test reached for modules through a sys.path hack
(sys.path.insert(parent) + "from database import ...") rather than the
"from backend.X import ..." style the rest of the suite uses. That flat import
makes database/models.py's "from ..utils.capture_chords import ..." point
outside the package. It also aimed at the wrong module: it wants the service
layer, which raises ValueError, not the route handler, which converts that
into an HTTPException.

Result: the full suite goes from 0 collected to 148 passed. The one remaining
failure, test_progress.py::test_hf_progress_tracker, is pre-existing and
unrelated (tqdm patching) - it reproduces identically on an unpatched tree.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:01:11 +00:00
jamiepine 3f933d6341 Merge origin/main into prep/pr-1031 2026-10-04 00:01:09 +00:00
jamiepine 99bcc7019b Merge remote-tracking branch 'origin/main' into prep/pr-1112
# Conflicts:
#	CHANGELOG.md
2026-10-04 00:01:04 +00:00
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] db2e7797b9 fix(models): reject an explicitly empty model_name instead of loading Qwen 2026-10-04 00:01:00 +00:00
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] cf0da2e6a3 fix(models): honor model_name on /models/load and /models/unload (fixes #977) 2026-10-04 00:01:00 +00:00
jamiepineandcapy-ai-staging[bot] d09c5c39e3 docs: update the last two comment references to the dotted MCP tool names 2026-10-04 00:00:54 +00:00
jamiepineandcapy-ai-staging[bot] 37af180887 docs: refer to the MCP tools by their new underscore names
Update the README, MCP docs, in-app MCP page, i18n strings, landing copy,
docstrings and comments to match the renamed tools. CHANGELOG entries and
docs/plans are historical and left as they were.
2026-10-04 00:00:54 +00:00
Jashwanthandcapy-ai-staging[bot] 8990104081 fix: use underscore-separated MCP tool names so Claude Desktop accepts them
Claude Desktop validates tool names against ^[a-zA-Z0-9_-]{1,64}$ and
rejects the whole tool list when any name contains a dot, so the Voicebox
MCP server was unusable there. Rename voicebox.speak, voicebox.transcribe,
voicebox.list_captures and voicebox.list_profiles to voicebox_speak,
voicebox_transcribe, voicebox_list_captures and voicebox_list_profiles.

Fixes #790
2026-10-04 00:00:54 +00:00