Collecting backend/tests/ fails outright on a clean checkout:
backend/tests/test_profile_duplicate_names.py:19: in <module>
from database import Base, VoiceProfile as DBVoiceProfile
E ImportError: attempted relative import beyond top-level package
pytest stops at the collection error, so the whole backend suite runs zero
tests rather than the ~142 it otherwise would.
There are two causes stacked on top of each other.
First, a real cycle in production code. routes/profiles.py, routes/history.py
and routes/stories.py each import safe_content_disposition from ..app, while
app.py builds the FastAPI instance at module scope (app = create_app() on
import), which registers those same routers. Importing any of those three
route modules first therefore re-enters a partially initialised app and dies
with "cannot import name 'router' from partially initialized module". It only
works today because app.py always happens to be imported first.
safe_content_disposition is a pure helper over urllib.parse.quote with no
application state, so it moves to backend/utils/http.py. app.py re-exports it
so any external caller importing it from the old location keeps working.
Second, the test reached for modules through a sys.path hack
(sys.path.insert(parent) + "from database import ...") rather than the
"from backend.X import ..." style the rest of the suite uses. That flat import
makes database/models.py's "from ..utils.capture_chords import ..." point
outside the package. It also aimed at the wrong module: it wants the service
layer, which raises ValueError, not the route handler, which converts that
into an HTTPException.
Result: the full suite goes from 0 collected to 148 passed. The one remaining
failure, test_progress.py::test_hf_progress_tracker, is pre-existing and
unrelated (tqdm patching) - it reproduces identically on an unpatched tree.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Update the README, MCP docs, in-app MCP page, i18n strings, landing copy,
docstrings and comments to match the renamed tools. CHANGELOG entries and
docs/plans are historical and left as they were.
Claude Desktop validates tool names against ^[a-zA-Z0-9_-]{1,64}$ and
rejects the whole tool list when any name contains a dot, so the Voicebox
MCP server was unusable there. Rename voicebox.speak, voicebox.transcribe,
voicebox.list_captures and voicebox.list_profiles to voicebox_speak,
voicebox_transcribe, voicebox_list_captures and voicebox_list_profiles.
Fixes#790
Adding seroval as a direct root dependency left the lockfile's
@tanstack/router-core/seroval entry pinned at the vulnerable 1.5.0, so the
only real consumer was still bundling the CVE-2026-59940 version. A bun
'overrides' entry forces every resolution to 1.5.3 and drops the unused
direct dependency.
Queries throughout the codebase filter on generations.profile_id,
generations.status, generations.created_at, story_items.story_id,
story_items.generation_id, generation_versions.generation_id, and
profile_samples.profile_id with every request. Without indexes SQLite
falls back to a full table scan; as history grows (hundreds or thousands
of generations) these scans become the dominant latency.
Changes:
- Add index=True on the most-queried FK and sort columns in models.py so
new installs get them from Base.metadata.create_all
- Add _migrate_add_indexes() called from run_migrations() so existing
installs get the same indexes on next startup (uses CREATE INDEX IF
NOT EXISTS — idempotent, <10 ms on any realistic dataset)
Switch from the default DELETE/ROLLBACK journal to WAL so concurrent
readers (SSE status polls, history queries) are not blocked while the
generation worker holds a write transaction. Set a 5-second busy
timeout to eliminate "database is locked" errors under brief write
contention.
Both PRAGMAs are applied via a custom creator function so every
connection in the pool gets the settings at open time, not just the
first one.
is_model_cached() marked a model as not-cached whenever any .incomplete
blob existed in its cache dir, even when a completed blob with the same
hash already sat next to it. A retried/concurrent download can leave
this orphan behind after the real transfer already finished, which made
the model appear perpetually "downloading" and re-trigger a full
re-download on every load.
Only .incomplete files with no matching completed blob now count as a
genuinely in-progress download.
Both speak surfaces built their GenerationRequest with a hardcoded "en"
fallback and never consulted the resolved profile, so a profile created
with language="fr" was still synthesised as English unless the caller
passed language= explicitly.
This hurts the MCP path most: an agent calling voicebox.speak has no way
to know the bound profile's language, so it cannot pass the argument
either. Every agent-triggered generation on a non-English profile came
out with an English accent.
The fallback chain is now explicit argument -> resolved profile's
language -> "en", which matches how engine and personality already
consult the resolved binding. The "en" backstop is kept so profiles with
no language set behave exactly as before.
Adds backend/tests/test_speak_language.py covering both surfaces: the
fallback, explicit-argument precedence, and the unchanged "en" default.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Addresses the budget comment on #1058 by the second route the review
offered -- defining the constant as the head budget rather than reserving
the suffix inside it.
TOAST_ERROR_BUDGET read as though it bounded `display`, but `display` is
the head plus " …", so it could be 402. Renamed to HEAD_BUDGET and
documented as bounding the message rather than the rendered string, with
the suffix now a named constant instead of a literal in the template.
Reserving the two characters was the alternative, but nothing downstream
has a hard limit -- the description box scrolls -- so it would have
shortened the message to satisfy a round number.
Verified: head <= 400 and display <= 402 on every truncating input,
including no-space text, a short first line, sentence-boundary backoff
and the real 4795-char error, with the untouched-when-not-truncated and
exact-`full` invariants still holding.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Follow-up to discussion_r3831361300 on #1058.
Both non-truncated return paths handed back the trimmed working copy, so
an error like " request timed out\n" came back altered even though
nothing had been omitted. That contradicted the stated intent that short
errors pass through untouched, and left `display` differing from `full`
for no reason.
`display` is now byte-identical to `full` whenever `truncated` is false —
the trimmed copy is only used for measuring against the budget and for
building the shortened head. Documented on the field.
Verified across padded short errors, clean short errors, empty and
whitespace-only input, and either side of the threshold: display === full
on every untruncated case, and `full` matches the input exactly in all of
them.
One visible consequence: with whitespace-pre-wrap on the description, an
error carrying leading or trailing newlines now renders with that blank
space. Trivial for the messages this sees in practice, and the
alternative was silently editing the text.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Two CodeRabbit findings on #1058.
condenseError trimmed the input before storing it in `full`, which is
documented as the untouched original and is what the Copy action hands
over. The trim now applies only to the working copy used for measuring
and cutting, so `full` is byte-for-byte what the server sent while
`display` and `omitted` still ignore surrounding blank space.
The Copy handler called navigator.clipboard.writeText with no guard.
Outside a secure context the property access itself throws, and
writeText rejects when permission is denied; neither was handled, so a
click could become an unhandled rejection with no sign that nothing was
copied. Both paths are now caught and reported, pointing at Settings ->
Logs as the fallback.
Not taken: aligning MIN_TO_CONDENSE with the 400-char budget. The gap is
deliberate -- cutting a 450-char error to 400 saves 50 characters in a
description that already scrolls, and no Copy action is needed there
because `display` holds the whole message. Documented in place.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A failed generation put the server's error straight into a toast. The
transformers "Unrecognized model" error is ~4.8KB, of which the first
sentence carries the meaning and the remaining 4.7KB is an alphabetical
list of every architecture it knows. In a 420px toast with
overflow-hidden and no scroll, that clipped the text at both ends and
pushed the close button off-screen: unreadable and undismissable.
- ToastDescription is capped at 40vh and scrolls, wraps on whitespace
and breaks long unspaced tokens so a path cannot widen the toast.
- The toast aligns to the start rather than centring, so the title
stays visible next to a tall description.
- condenseError() keeps the head of an oversized error, cutting at the
first newline or the last sentence end inside a 400-char budget, and
reports how many characters it dropped. Short errors pass through
untouched.
- When it does truncate, the toast offers a Copy action for the full
text and points at Settings -> Logs.
Verified against the real 4795-char error: 4795 -> 400 chars keeping
both meaningful sentences.
The hook moves to .tsx to render ToastAction, matching useAutoUpdater.tsx
which is a .tsx hook for the same reason. createElement was tried first
but this repo's ToastActionElement type is the older shadcn definition
(ReactElement<typeof ToastAction>) which only accepts JSX-constructed
elements.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
When the user disables auto_refine (LLM polish) in Settings, the Qwen
refinement model is no longer required for dictation to arm. Previously,
canRecord checked llmReady unconditionally, so useChordSync called
disable_hotkey whenever Qwen was not downloaded -- even if the user
never intended to use refinement.
Changes:
- useDictationReadiness: gate llmReady behind autoRefine in canRecord,
missing, and the polling predicates so hotkeys arm with just Whisper
STT when refinement is off.
- DictationReadinessChecklist: hide the LLM row when autoRefine is
false -- the checklist now shows only Whisper STT.
Closes#753
Without a trace the Settings GPU label and /health keep reporting CUDA
while generation runs on CPU, so the user has no way to confirm the
override engaged.
The override is documented in gpu-acceleration.mdx and listed as step 1 of
the get_torch_device() precedence in tts-generation.mdx, but grepping the
tree for VOICEBOX_FORCE_CPU matched only those two doc files - nothing read
it. Users whose GPU has no compiled kernels in the bundled PyTorch had no
way to fall back to CPU short of renaming the installed CUDA backend
directory.
Resolve it before torch is imported, so it still works when the installed
build is itself the reason CPU is wanted.
list_generations() fetched versions with one SELECT per generation on the
page (50 rows -> 51 queries). Add _get_versions_for_generations() which
loads all versions for the page in a single WHERE generation_id IN (...)
query and groups them in memory; the single-generation helper now
delegates to it so story item details behave identically.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <[email protected]>
The 19-row menu is taller than its 280px max-height, so arrow-key
navigation past the fold lost its highlight. Scroll the active row into
view on index change, list the delivery tags in the README, and give
[sarcastic] and [whispering] emoji that are not already used by
[chuckle] and [shush].
When the server outlives the Tauri app that spawned it (keep-running
mode, or a sidecar the next launch reuses), its stdout/stderr pipe has
no reader. Every later print()/tqdm write raises BrokenPipeError, so
POST /captures and /transcribe return "[Errno 32] Broken pipe".
Wrap stdout/stderr so they fall back to devnull on the first failed
write instead of raising.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The floating generate box is fixed at the bottom of the viewport, so
all of its Select dropdowns (voice profile, language, engine, effects)
opened downward into — or beyond — the window edge. Add side="top" to
each SelectContent so the menus appear above their trigger instead.
Co-authored-by: Claude Sonnet 4.6 <[email protected]>
key_from_str() had no arm for "Function", so it fell through to
None. Since build_chord propagates that as a hard Err via ?, binding
any chord containing fn made build_chord_bindings fail entirely —
HotkeyMonitor was never spawned, silently killing both push-to-talk
and toggle-to-talk until the chord was reverted.
Every other layer (keytap's macOS key tap, Key::Function itself, the
frontend's canonicalKeyFromEvent/displayLabelForKey) already handles
fn — only this string-to-Key bridge was missing the arm.
Fixes#941
Export filenames were derived from only the first 30 characters of the
generation text. Generations with similar wording (a common workflow when
iterating on the same line) produced identical filenames, so exports
collided on disk — the browser appended " (1)"/" (2)" and users ended up
opening audio that didn't match the expected filename.
Append the first 8 chars of the generation id to the .wav and .voicebox.zip
export filenames, in both the backend Content-Disposition headers and the
frontend save-file hooks.
Co-authored-by: Claude Opus 4.8 <[email protected]>
`UploadFile.filename` can be None, and `Path(None)` raises TypeError. On the
avatar endpoint this happens before the try/except, so a filename-less upload
surfaces as an unhandled 500 instead of a clean response. Every other upload
handler already guards this with `file.filename or ""` (add_profile_sample,
transcription, generations); apply the same guard here.
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
* fix(ui): parse naive-UTC timestamps consistently in formatAbsoluteDate
Backend timestamps are naive UTC (Python `datetime.utcnow()`) and are
serialized without a timezone suffix. `formatDate` already normalizes
these by appending `Z` before parsing, but `formatAbsoluteDate` called
`new Date(date)` directly. Per the ES spec, a timezone-less date-time
string is parsed as local time, so absolute timestamps were shown off by
the viewer's UTC offset (e.g. +9h in JST) — and disagreed with the
relative time rendered by `formatDate` for the same value (visible in the
Captures detail panel, which uses both on `capture.created_at`).
Extract the normalization into a shared `parseServerDate` helper and use
it in both formatters.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* docs(format): clarify parseServerDate comment on date-only vs date-time parsing
ECMAScript parses date-only strings ("2026-07-23") as UTC but timezone-less
date-time strings ("2026-07-23T10:00:00") as local time. The backend emits the
latter, which is the case this helper normalizes. Corrects the comment per PR
review feedback.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* docs(format): trim parseServerDate comment to match surrounding style
Reduce the multi-line explanation to a single why-comment consistent with
other utils comments.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
The MCP voicebox.speak tool built its GenerationRequest without a
model_size, so every agent-triggered generation fell back to the schema
default ("1.7B"). There was no way to reach the 0.6B Qwen variant (or
TADA's 1B/3B) through MCP, and callers paid a model reload whenever the
requested size differed from what was already loaded.
Thread an optional model_size through voicebox.speak and the _speak
helper into GenerationRequest, mirroring the REST /generate surface.
Omitting it passes None, which generate_speech normalizes to the engine
default, so existing callers are unaffected.
Add backend/tests/test_mcp_speak.py covering the forwarded value, the
omitted-default path, and rejection of an invalid size.
Fixes#884
Add three environment variables to prevent miopenStatusUnknownError and
system stuttering during inference on RDNA4 GPUs:
- MIOPEN_USER_DB_PATH: redirect MIOpen kernel cache to writable, persistent dir
- MIOPEN_CUSTOM_CACHE_DIR: same, for custom operator cache
- MIOPEN_FIND_MODE=FAST: use heuristic kernel selection instead of exhaustive
benchmarking, which fails on RDNA4 with ptr: 0 size: 0 workspace warnings
MIOPEN_FIND_MODE=FAST does not affect output quality. All MIOpen kernel
variants produce the same numerical result; fast mode selects a known-good
kernel using heuristics instead of benchmarking every variant on the GPU.
Tested on RX 9070 (gfx1201) with ROCm 7.2 and PyTorch 2.12.1+rocm7.2.
Hardware note: tested on Ryzen 7 9800X3D + RX 9070 with Gigabyte B650M DS3H
motherboard. The exhaustive benchmarking failures may be related to IOMMU
behavior on this platform. This system was affected by an IOMMU bug patched
upstream in kernel 6.19.10, which may be a contributing factor. May not
affect all RDNA4 systems. MIOPEN_FIND_MODE=FAST is a safe default regardless.
Depends on PR #862 which fixes the broken ROCm Docker build.
A Windows Git checkout with checkout-time CRLF conversion enabled
produces CRLF working-tree copies of package.json and
scripts/rocm-entrypoint.sh, breaking the Docker build two ways:
- The frontend stage's `sed -i -z 's/,\n ]/…/'` is LF-anchored, so
it doesn't match against \r\n and leaves an invalid trailing comma
in package.json, which then fails JSON parsing in the vite build.
- The final stage copies rocm-entrypoint.sh straight from the build
context; with a CRLF shebang the container reports the misleading
"no such file or directory" for an entrypoint that plainly exists,
because Linux can't resolve "/bin/sh\r" as an interpreter.
Add .gitattributes forcing LF for both files at checkout time, plus a
sed normalization step in each Dockerfile stage for resilience with
clones that predate the .gitattributes rule.
Fixes#915
std::env::set_var is not thread-safe on Unix (unsafe as of Rust 2024
edition) and calling it from a spawned capture thread while other
threads (tokio runtime, webview, Tauri plugins) may read the
environment is a data race risk. It also never got unset, so the
monitor source would leak into any later cpal/ALSA init in the same
process.
Replace the env-var indirection with direct device selection: when
pactl reports a monitor source name, search cpal's input device
enumeration for an exact match. Fall back to a substring match on
'monitor' (the original pactl-unavailable path), then the host's
default input device. This is the 'pass the source name directly to
cpal' option from the issue - no env mutation, no leakage between
capture sessions, and it still re-detects the current default sink's
monitor on every start_capture call.
Fixes#471
The /transcribe endpoint passed the raw uploaded file straight to the STT
backend (mlx_audio.stt -> miniaudio), which only decodes WAV/FLAC/MP3/Vorbis.
Browser recordings arrive as WebM/Opus (Chrome/Firefox MediaRecorder), so
web-mode dictation failed with 500 "unsupported file format". The Tauri app
was unaffected because WebKit produces MP4.
librosa already fully decodes the upload to compute duration (falling back to
audioread/ffmpeg for exotic containers), so re-encode that PCM to a temp WAV
and hand it to Whisper. WAV inputs pass through unchanged; the temp file is
cleaned up in the finally block.
Co-authored-by: Claude Opus 4.8 <[email protected]>