Commit Graph
327 Commits
Author SHA1 Message Date
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] 64b4b976d6 fix(versions): don't let a locked version file abort the delete cascade 2026-10-04 00:19:41 +00:00
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] 4750b47b04 fix(profiles): delete a profile's generations instead of orphaning them 2026-10-04 00:19:41 +00:00
jamiepineandcapy-ai-staging[bot] 4f54e16824 test(speak): call the MCP tool by its underscore name
#1137 added tests that call the MCP tool as "voicebox.speak"; #1135 renamed
the tools to underscore names (voicebox_speak) so Claude Desktop accepts them.
Both merged cleanly but the three MCP tests failed on main with NotFoundError.
Test-only change.
2026-10-04 00:15:34 +00:00
harryandcapy-ai-staging[bot] ab9a19790c fix(docker): fix startup crash and permission errors in the container
Four related fixes uncovered while getting Docker running on a Linux
host without AVX-512 and with an NVIDIA GPU:

- pedalboard>=0.9.21 ships a Linux wheel with AVX-512 instructions
  baked into its native extension, which SIGILLs (exit 132) on import
  on any CPU without AVX-512 support. Pin below the regression until
  upstream fixes it (spotify/pedalboard#454).
- The rocminfo probe in app.py ran unconditionally on every build
  variant, always failing and logging on non-ROCm systems. Gate it on
  /dev/kfd actually being present.
- sox (required by qwen-tts's X-vector extractor) and a C compiler
  (required by Triton to JIT-compile its CUDA driver shim on first
  use) were missing from the runtime image, causing hard failures
  once those code paths were actually exercised.
- Named volumes and bind mounts (HF cache, app data, generated audio)
  are created root-owned by Docker on first use, but the app runs as
  the unprivileged voicebox user. chown the mount points in the
  entrypoint before dropping privileges so this self-heals on every
  start regardless of host UID.
2026-10-04 00:12:29 +00:00
capy-ai-staging[bot]andGitHub fee218ef0b Merge pull request #1130 from jamiepine/prep/pr-1112
Fix infinite HF retry storm when loading a cached model offline
2026-10-04 00:05:46 +00:00
capy-ai-staging[bot]andGitHub 39b1c91291 Merge pull request #1144 from jamiepine/prep/pr-1031
fix(setup): pin dev venv to Python 3.12
2026-10-04 00:05:41 +00:00
capy-ai-staging[bot]andGitHub 3221b2796b Merge pull request #1146 from jamiepine/prep/pr-659
fix: replace deprecated datetime.utcnow() with datetime.now(UTC) throughout
2026-10-04 00:05:36 +00:00
jamiepine ba942ec1fd Merge remote-tracking branch 'origin/main' into prep/pr-659
# Conflicts:
#	backend/database/models.py
2026-10-04 00:02:23 +00:00
jamiepineandcapy-ai-staging[bot] dd8ab5cb20 fix(mlx): fold Qwen3 LLM load+generate into one worker submission; bind model locally in generate closures
Review follow-ups on the thread-affinity PR: MLXQwenLLMBackend.generate
awaited load_model as one worker submission and then submitted
_generate_sync as a second, leaving the same load/generate gap the TTS
and STT paths close (a concurrent load_model for another size or an
unload could swap or null out self.model/self.tokenizer in between).
Give the LLM backend the same _op_lock + single _reload_and_generate_sync
shape. The TTS/STT/LLM sync closures now bind the model to a local once,
so an inline unload_model() from the event loop mid-generation cannot
turn a later self.model read (the voice-clone fallback path in
particular) into an AttributeError.
2026-10-04 00:01:47 +00:00
jamiepineandcapy-ai-staging[bot] cf984885e9 fix(mlx): keep unload_model off the MLX worker so it cannot stall the event loop
Routing unload through the single MLX worker and waiting on the result
blocked the caller (the FastAPI event loop via /models/unload) for the
remainder of any in-flight generation — measured 3.7 s stall on a 16 s
clip on an M2 Ultra, unbounded for long texts. Dropping the model
reference inline is thread-safe (MLX frees buffers through its global
allocator) and is what main did before the thread-affinity change; the
generation in flight keeps its own reference and the next generate()
reloads on the worker.
2026-10-04 00:01:47 +00:00
99d05b917c fix(backends): run MLX load and inference on one thread
MLX streams are thread-local and mlx-audio caches one on the model at
load time. Model load and inference were each dispatched through
asyncio.to_thread(), which uses a multi-worker pool, so load and
generate/transcribe could land on different OS threads -- the inference
thread then has no Stream(gpu, N) and MLX aborts with
"There is no Stream(gpu, 1) in current thread."

Route every MLX call (load, generate, transcribe; TTS and STT) through a
single dedicated worker thread so a model and its stream always share a
thread. max_workers=1 also serialises the single local GPU. Adds a
regression test covering the thread-affinity invariant.

Fixes #699. Also addresses #675.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
2026-10-04 00:01:47 +00:00
2693c983f2 fix: extend MLX thread-affinity fix to the Qwen3 LLM backend
The dedicated-MLX-thread fix (#886) covers the TTS and STT backends, but
MLXQwenLLMBackend still dispatches load/generate through
asyncio.to_thread(), so /llm/generate (personality compose/rewrite,
dictation refinement) crashes with the same
'There is no Stream(gpu, 0) in current thread.' — raised from
mlx_lm.generate's wired_limit on exit — whenever load and generate land
on different pool threads.

Route the MLX LLM backend's unload/load/generate through the same
_run_on_mlx_thread helper so every MLX call in the process shares one
worker thread.

Verified on Apple M3 Pro (macOS 26.5, mlx 0.32.0, mlx-lm 0.31.1,
mlx-audio 0.4.1): /llm/generate failed 100% before, succeeds after;
TTS + STT unaffected.

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-10-04 00:01:47 +00:00
Ron David Ben Ishayandcapy-ai-staging[bot] 57ef6636bc fix(mlx): fold reload+inference into one MLX-worker submission
CodeRabbit follow-up on the previous round: the asyncio.Lock closed the
gap for callers going through generate()/transcribe() consistently, but
unload_model() itself never touched that lock and could still land
between the load future resolving and the generate/transcribe future
being submitted - two separate _run_on_mlx_thread calls, so a real gap
existed at the Python level even though the executor is single-worker.

Fix: generate() and transcribe() each now submit exactly ONE callable
to the MLX executor - a self-healing reload-if-needed-then-infer
function - instead of a load submission followed by a separate infer
submission. This removes the gap structurally: nothing can observe an
intermediate state because there is no intermediate state exposed
across an await boundary. unload_model() now also submits
unconditionally (loaded-check moved inside _unload_model_sync, which
runs atomically with the teardown) rather than racing its own
Python-level self.model check against a load in flight.

Verified: same-size generate, size-switch generate, unload route, and
a post-unload generate all complete clean on the live launchd server.
2026-10-04 00:01:47 +00:00
Ron David Ben Ishayandcapy-ai-staging[bot] fa8db820d7 fix(mlx): address review feedback on the thread-affinity fix
Two real follow-on issues found by CodeRabbit on PR #989:

1. MLXTTSBackend.load_model_async called self.unload_model() directly
   on the event-loop thread before dispatching _load_model_sync to the
   MLX worker thread - model teardown ran on the wrong OS thread, same
   class of bug the PR itself fixes. Combined unload+load into one
   _reload_sync callable submitted as a single MLX-thread operation.

2. Public unload_model() (called synchronously from the /models/unload
   routes) also ran on the caller thread. Now submits to the MLX
   executor and blocks on the result, so teardown always happens on the
   worker thread regardless of caller. Same fix applied to
   MLXSTTBackend.

3. Both backends cache self.model on the instance, but generate()/
   transcribe() only locked their own internal load step - a
   concurrent request for a different model_size could swap self.model
   between one request's load and its inference (real race: routes/
   generations.py calls load_engine_model() and generate_chunked() as
   separate awaited steps with a gap between them). Added a per-backend
   asyncio.Lock held across the full load+inference sequence in both
   generate() and transcribe(). Note: this closes the race for callers
   using the backend's own public methods consistently; the wider
   route-level orchestration race (load_engine_model + generate_chunked
   as two separate calls) is a follow-up outside this file's scope.

Verified: same-size and cross-size-switch generations both complete
clean after the patch; /models/{name}/unload route returns 200 without
deadlocking the (still-responsive) server.
2026-10-04 00:01:47 +00:00
Ron David Ben Ishayandcapy-ai-staging[bot] f135a471ae fix(mlx): pin all MLX ops to a single worker thread
Qwen3-TTS (and MLX STT) generation crashed with:
"There is no Stream(gpu, N) in current thread."

MLXTTSBackend/MLXSTTBackend dispatched model load and generate/
transcribe as separate asyncio.to_thread() calls, which round-robin
across Python's default multi-worker executor pool. MLX's Metal
backend keeps GPU streams registered per-OS-thread, so a model loaded
on one worker thread and then used for generation on a different
worker thread hits a missing stream and crashes.

Reproduced 100% of the time on macOS/Apple Silicon cloning with both
the 1.7B and 0.6B Qwen3-TTS models; Chatterbox/Kokoro were unaffected
since they use the PyTorch backend, not this module.

Fix: route all four MLX call sites in this file (TTS load, TTS
generate, STT load, STT transcribe) through a dedicated single-worker
ThreadPoolExecutor instead of asyncio.to_thread's shared pool, so
every MLX operation for a given process runs on the same OS thread.

Verified: direct /generate API calls against both model sizes
completed cleanly after the fix (previously failed every time).
2026-10-04 00:01:47 +00:00
jamiepineandcapy-ai-staging[bot] 86dc46b930 refactor(stt): name the Whisper sample rate and window constants 2026-10-04 00:01:34 +00:00
jamiepineandcapy-ai-staging[bot] 51882b5065 fix(backend): keep the pad-to-30s Whisper path for clips under 30s
With padding="longest" the encoder rejects inputs shorter than 3000 mel
frames when generate() has to detect the language first, so every short
clip without a forced language failed with "Whisper expects the mel input
features to be of length 3000". Use the long-form feature-extractor and
generate() options only when the audio exceeds one 30s window.
2026-10-04 00:01:34 +00:00
noxandcapy-ai-staging[bot] 7575a65e7f fix(backend): pass language/task directly to Whisper generate 2026-10-04 00:01:34 +00:00
e61c85365b fix(backend): enable long-form Whisper transcription on PyTorch path
The PyTorch Whisper transcribe() called the HF processor without
truncation=False and model.generate() without return_timestamps=True.
With those defaults, WhisperFeatureExtractor silently truncates inputs
to 30 s (Whisper's native receptive field), so any dictation longer
than ~30 s lost its tail.

Setting truncation=False + padding="longest" + return_attention_mask=True
on the processor, then forwarding the attention mask plus
return_timestamps=True to generate(), flips HF Whisper into long-form
mode: autoregressive decoding over rolling 30 s windows.

Verified by round-tripping a 56.6 s Kokoro TTS sample through
/transcribe — full text returned including content past the 30 s mark;
previously the transcript was cut off roughly halfway through.

MLX backend (mlx_backend.py) is intentionally unchanged: mlx_audio.stt's
generate() already implements rolling-window long-form transcription
with condition_on_previous_text in the upstream library, so it does not
have the same bug. The HF-only kwargs added here would also break the
MLX call signature.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
2026-10-04 00:01:34 +00:00
jamiepineandcapy-ai-staging[bot] ae300c5316 fix(kokoro): trim edges only, keep inter-segment gaps
Registering Kokoro with needs_trim=True routed its output through the
generic trim_tts_output, whose 1s internal-silence cut was tuned for
Chatterbox hallucinations. KPipeline synthesizes newline- and
token-limit-separated segments independently, and the ~0.3s lead plus
~0.7s tail pads at each boundary add up to 1.2s of silence, so with
af_sarah a four-paragraph script came back as its first paragraph only
(10.9s -> 1.8s) and a 51s text lost half its segments.

Drop the needs_trim flag, keep the in-backend trim, and give
trim_tts_output a max_internal_silence_ms=None mode that only trims the
leading and trailing pads. Tests updated for the new behaviour, with a
two-segment fake pipeline and an explicit internal-gap case.
2026-10-04 00:01:24 +00:00
devangkanthariaandcapy-ai-staging[bot] 615aeaeb35 fix(kokoro): trim trailing silence and run-on noise on short prompt synthesis (#960) 2026-10-04 00:01:24 +00:00
jamiepineandcapy-ai-staging[bot] b788dc383c fix(generation): make post-generation cleanup best-effort and cover /generate/stream
Wrap empty_device_cache in release_generation_memory(), which logs and
swallows cleanup failures so a poisoned CUDA context cannot replace a
finished generation's result, and call it from the streaming endpoint
too, which drives generate_chunked directly.
2026-10-04 00:01:18 +00:00
jamiepineandcapy-ai-staging[bot] 17fd1ddd1b refactor(generation): simplify post-generation cache cleanup
Initialise tts_model before the try so the finally can read its device
without a locals() probe, import empty_device_cache alongside the other
backend imports instead of inside a bare try/except, and flatten the MPS
branch in empty_device_cache.
2026-10-04 00:01:18 +00:00
devangkanthariaandcapy-ai-staging[bot] 38389db82b fix(backend): prevent unbounded memory accumulation over consecutive TTS generations (#923) 2026-10-04 00:01:18 +00:00
63f09455ef fix(tests): repair test_profile_duplicate_names so the suite can run
The file still used the pre-refactor flat imports and a sys.path hack:

    sys.path.insert(0, str(Path(__file__).parent.parent))
    from database import Base, VoiceProfile as DBVoiceProfile
    from profiles import create_profile, update_profile

`profiles` now lives at backend/services/profiles.py, so collection raised
ImportError. Because pytest aborts the whole run on a collection error, this
one file meant `just test` ran zero tests — duplicate-name validation has had
no coverage since the services refactor.

Switch to package imports like every other test module, and drop
DBVoiceProfile, which was imported but never used.

That exposed a second, latent bug: all 6 tests passed but every one errored in
teardown with PermissionError WinError 32. The fixture closed the session and
then rmtree'd the temp dir, but closing a session does not release
SQLAlchemy's pooled connection, so SQLite still held test.db open on Windows.
Dispose the engine before removing the directory.

6 passed, and full-suite collection goes from aborting to 170 tests.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:01:11 +00:00
3e240a9ef1 fix(backend): break the app<->routes import cycle that aborts the test suite
Collecting backend/tests/ fails outright on a clean checkout:

    backend/tests/test_profile_duplicate_names.py:19: in <module>
        from database import Base, VoiceProfile as DBVoiceProfile
    E   ImportError: attempted relative import beyond top-level package

pytest stops at the collection error, so the whole backend suite runs zero
tests rather than the ~142 it otherwise would.

There are two causes stacked on top of each other.

First, a real cycle in production code. routes/profiles.py, routes/history.py
and routes/stories.py each import safe_content_disposition from ..app, while
app.py builds the FastAPI instance at module scope (app = create_app() on
import), which registers those same routers. Importing any of those three
route modules first therefore re-enters a partially initialised app and dies
with "cannot import name 'router' from partially initialized module". It only
works today because app.py always happens to be imported first.

safe_content_disposition is a pure helper over urllib.parse.quote with no
application state, so it moves to backend/utils/http.py. app.py re-exports it
so any external caller importing it from the old location keeps working.

Second, the test reached for modules through a sys.path hack
(sys.path.insert(parent) + "from database import ...") rather than the
"from backend.X import ..." style the rest of the suite uses. That flat import
makes database/models.py's "from ..utils.capture_chords import ..." point
outside the package. It also aimed at the wrong module: it wants the service
layer, which raises ValueError, not the route handler, which converts that
into an HTTPException.

Result: the full suite goes from 0 collected to 148 passed. The one remaining
failure, test_progress.py::test_hf_progress_tracker, is pre-existing and
unrelated (tqdm patching) - it reproduces identically on an unpatched tree.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:01:11 +00:00
jamiepine 3f933d6341 Merge origin/main into prep/pr-1031 2026-10-04 00:01:09 +00:00
jamiepine 99bcc7019b Merge remote-tracking branch 'origin/main' into prep/pr-1112
# Conflicts:
#	CHANGELOG.md
2026-10-04 00:01:04 +00:00
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] db2e7797b9 fix(models): reject an explicitly empty model_name instead of loading Qwen 2026-10-04 00:01:00 +00:00
SEPURI-SAI-KRISHNAandcapy-ai-staging[bot] cf0da2e6a3 fix(models): honor model_name on /models/load and /models/unload (fixes #977) 2026-10-04 00:01:00 +00:00
jamiepineandcapy-ai-staging[bot] d09c5c39e3 docs: update the last two comment references to the dotted MCP tool names 2026-10-04 00:00:54 +00:00
jamiepineandcapy-ai-staging[bot] 37af180887 docs: refer to the MCP tools by their new underscore names
Update the README, MCP docs, in-app MCP page, i18n strings, landing copy,
docstrings and comments to match the renamed tools. CHANGELOG entries and
docs/plans are historical and left as they were.
2026-10-04 00:00:54 +00:00
Jashwanthandcapy-ai-staging[bot] 8990104081 fix: use underscore-separated MCP tool names so Claude Desktop accepts them
Claude Desktop validates tool names against ^[a-zA-Z0-9_-]{1,64}$ and
rejects the whole tool list when any name contains a dot, so the Voicebox
MCP server was unusable there. Rename voicebox.speak, voicebox.transcribe,
voicebox.list_captures and voicebox.list_profiles to voicebox_speak,
voicebox_transcribe, voicebox_list_captures and voicebox_list_profiles.

Fixes #790
2026-10-04 00:00:54 +00:00
jamiepineandcapy-ai-staging[bot] dd1ada8239 fix(db): log which indexes the migration actually created 2026-10-04 00:00:41 +00:00
jamiepineandcapy-ai-staging[bot] 22a3f6b3dd fix(db): drop unused Index import 2026-10-04 00:00:41 +00:00
Will Andersonandcapy-ai-staging[bot] b97ffbfba9 Add indexes on high-traffic foreign keys and sort columns
Queries throughout the codebase filter on generations.profile_id,
generations.status, generations.created_at, story_items.story_id,
story_items.generation_id, generation_versions.generation_id, and
profile_samples.profile_id with every request. Without indexes SQLite
falls back to a full table scan; as history grows (hundreds or thousands
of generations) these scans become the dominant latency.

Changes:
- Add index=True on the most-queried FK and sort columns in models.py so
  new installs get them from Base.metadata.create_all
- Add _migrate_add_indexes() called from run_migrations() so existing
  installs get the same indexes on next startup (uses CREATE INDEX IF
  NOT EXISTS — idempotent, <10 ms on any realistic dataset)
2026-10-04 00:00:41 +00:00
jamiepineandcapy-ai-staging[bot] ccc092bad0 perf(db): set the SQLite pragmas from a connect hook and clear WAL sidecars on db reset 2026-10-04 00:00:35 +00:00
Will Andersonandcapy-ai-staging[bot] 36bb6c1ad7 Enable SQLite WAL journal mode and 5 s busy timeout
Switch from the default DELETE/ROLLBACK journal to WAL so concurrent
readers (SSE status polls, history queries) are not blocked while the
generation worker holds a write transaction.  Set a 5-second busy
timeout to eliminate "database is locked" errors under brief write
contention.

Both PRAGMAs are applied via a custom creator function so every
connection in the pool gets the settings at open time, not just the
first one.
2026-10-04 00:00:35 +00:00
jamiepineandcapy-ai-staging[bot] e86a2dcaa5 fix(cache): share the orphaned-.incomplete check with /models/status 2026-10-04 00:00:28 +00:00
Alejandro Gaston Alvarezandcapy-ai-staging[bot] f62aff0809 fix(cache): don't treat orphaned .incomplete blobs as an in-progress download
is_model_cached() marked a model as not-cached whenever any .incomplete
blob existed in its cache dir, even when a completed blob with the same
hash already sat next to it. A retried/concurrent download can leave
this orphan behind after the real transfer already finished, which made
the model appear perpetually "downloading" and re-trigger a full
re-download on every load.

Only .incomplete files with no matching completed blob now count as a
genuinely in-progress download.
2026-10-04 00:00:28 +00:00
jamiepineandcapy-ai-staging[bot] 6ad47dda89 test(speak): build the MCP test server from fastmcp, the package production imports 2026-10-04 00:00:22 +00:00
jamiepineandcapy-ai-staging[bot] 1a3942f3b3 style: ruff-format the new speak language tests 2026-10-04 00:00:22 +00:00
c339d2c324 fix(speak): honour the voice profile's language instead of forcing English
Both speak surfaces built their GenerationRequest with a hardcoded "en"
fallback and never consulted the resolved profile, so a profile created
with language="fr" was still synthesised as English unless the caller
passed language= explicitly.

This hurts the MCP path most: an agent calling voicebox.speak has no way
to know the bound profile's language, so it cannot pass the argument
either. Every agent-triggered generation on a non-English profile came
out with an English accent.

The fallback chain is now explicit argument -> resolved profile's
language -> "en", which matches how engine and personality already
consult the resolved binding. The "en" backstop is kept so profiles with
no language set behave exactly as before.

Adds backend/tests/test_speak_language.py covering both surfaces: the
fallback, explicit-argument precedence, and the unchanged "en" default.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:00:22 +00:00
jamiepineandcapy-ai-staging[bot] 4a03cd7e77 fix(backend): log when VOICEBOX_FORCE_CPU forces the CPU device
Without a trace the Settings GPU label and /health keep reporting CUDA
while generation runs on CPU, so the user has no way to confirm the
override engaged.
2026-10-04 00:00:03 +00:00
Ousama Ben Younesandcapy-ai-staging[bot] 8b8c1429db fix(backend): honour VOICEBOX_FORCE_CPU in device selection
The override is documented in gpu-acceleration.mdx and listed as step 1 of
the get_torch_device() precedence in tts-generation.mdx, but grepping the
tree for VOICEBOX_FORCE_CPU matched only those two doc files - nothing read
it. Users whose GPU has no compiled kernels in the bundled PyTorch had no
way to fall back to CPU short of renaming the installed CUDA backend
directory.

Resolve it before torch is imported, so it still works when the installed
build is itself the reason CPU is wanted.
2026-10-04 00:00:03 +00:00
youtsuhoandcapy-ai-staging[bot] 029d4d3378 fix(backend): batch generation-version queries to eliminate N+1 in history listing
list_generations() fetched versions with one SELECT per generation on the
page (50 rows -> 51 queries). Add _get_versions_for_generations() which
loads all versions for the page in a single WHERE generation_id IN (...)
query and groups them in memory; the single-generation helper now
delegates to it so story item details behave identically.

Generated with Codebuff 🤖
Co-Authored-By: Codebuff <[email protected]>
2026-10-03 23:59:56 +00:00
jamiepineandcapy-ai-staging[bot] c9f5f2c3e7 fix(server): route writelines through the pipe-safe write
writelines() was forwarded straight to the wrapped stream by __getattr__,
so it could still raise BrokenPipeError after the app's pipe closed.
2026-10-03 23:59:43 +00:00
848bb2e4c6 fix(server): stop failing requests after the app's stdout pipe closes
When the server outlives the Tauri app that spawned it (keep-running
mode, or a sidecar the next launch reuses), its stdout/stderr pipe has
no reader. Every later print()/tqdm write raises BrokenPipeError, so
POST /captures and /transcribe return "[Errno 32] Broken pipe".

Wrap stdout/stderr so they fall back to devnull on the first failed
write instead of raising.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-10-03 23:59:43 +00:00
Nikhil Jangidandcapy-ai-staging[bot] cdeeb38be4 test: cover generation engine defaults
Verify omitted engines remain unset while explicit values are validated.
2026-10-03 23:59:33 +00:00
Nikhil Jangidandcapy-ai-staging[bot] d4face494f fix: honor profile engine for generation requests
Allow omitted engines to fall through to the profile default.
2026-10-03 23:59:33 +00:00