With padding="longest" the encoder rejects inputs shorter than 3000 mel
frames when generate() has to detect the language first, so every short
clip without a forced language failed with "Whisper expects the mel input
features to be of length 3000". Use the long-form feature-extractor and
generate() options only when the audio exceeds one 30s window.
The PyTorch Whisper transcribe() called the HF processor without
truncation=False and model.generate() without return_timestamps=True.
With those defaults, WhisperFeatureExtractor silently truncates inputs
to 30 s (Whisper's native receptive field), so any dictation longer
than ~30 s lost its tail.
Setting truncation=False + padding="longest" + return_attention_mask=True
on the processor, then forwarding the attention mask plus
return_timestamps=True to generate(), flips HF Whisper into long-form
mode: autoregressive decoding over rolling 30 s windows.
Verified by round-tripping a 56.6 s Kokoro TTS sample through
/transcribe — full text returned including content past the 30 s mark;
previously the transcript was cut off roughly halfway through.
MLX backend (mlx_backend.py) is intentionally unchanged: mlx_audio.stt's
generate() already implements rolling-window long-form transcription
with condition_on_previous_text in the upstream library, so it does not
have the same bug. The HF-only kwargs added here would also break the
MLX call signature.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
Registering Kokoro with needs_trim=True routed its output through the
generic trim_tts_output, whose 1s internal-silence cut was tuned for
Chatterbox hallucinations. KPipeline synthesizes newline- and
token-limit-separated segments independently, and the ~0.3s lead plus
~0.7s tail pads at each boundary add up to 1.2s of silence, so with
af_sarah a four-paragraph script came back as its first paragraph only
(10.9s -> 1.8s) and a 51s text lost half its segments.
Drop the needs_trim flag, keep the in-backend trim, and give
trim_tts_output a max_internal_silence_ms=None mode that only trims the
leading and trailing pads. Tests updated for the new behaviour, with a
two-segment fake pipeline and an explicit internal-gap case.
Wrap empty_device_cache in release_generation_memory(), which logs and
swallows cleanup failures so a poisoned CUDA context cannot replace a
finished generation's result, and call it from the streaming endpoint
too, which drives generate_chunked directly.
Initialise tts_model before the try so the finally can read its device
without a locals() probe, import empty_device_cache alongside the other
backend imports instead of inside a bare try/except, and flatten the MPS
branch in empty_device_cache.
The file still used the pre-refactor flat imports and a sys.path hack:
sys.path.insert(0, str(Path(__file__).parent.parent))
from database import Base, VoiceProfile as DBVoiceProfile
from profiles import create_profile, update_profile
`profiles` now lives at backend/services/profiles.py, so collection raised
ImportError. Because pytest aborts the whole run on a collection error, this
one file meant `just test` ran zero tests — duplicate-name validation has had
no coverage since the services refactor.
Switch to package imports like every other test module, and drop
DBVoiceProfile, which was imported but never used.
That exposed a second, latent bug: all 6 tests passed but every one errored in
teardown with PermissionError WinError 32. The fixture closed the session and
then rmtree'd the temp dir, but closing a session does not release
SQLAlchemy's pooled connection, so SQLite still held test.db open on Windows.
Dispose the engine before removing the directory.
6 passed, and full-suite collection goes from aborting to 170 tests.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Collecting backend/tests/ fails outright on a clean checkout:
backend/tests/test_profile_duplicate_names.py:19: in <module>
from database import Base, VoiceProfile as DBVoiceProfile
E ImportError: attempted relative import beyond top-level package
pytest stops at the collection error, so the whole backend suite runs zero
tests rather than the ~142 it otherwise would.
There are two causes stacked on top of each other.
First, a real cycle in production code. routes/profiles.py, routes/history.py
and routes/stories.py each import safe_content_disposition from ..app, while
app.py builds the FastAPI instance at module scope (app = create_app() on
import), which registers those same routers. Importing any of those three
route modules first therefore re-enters a partially initialised app and dies
with "cannot import name 'router' from partially initialized module". It only
works today because app.py always happens to be imported first.
safe_content_disposition is a pure helper over urllib.parse.quote with no
application state, so it moves to backend/utils/http.py. app.py re-exports it
so any external caller importing it from the old location keeps working.
Second, the test reached for modules through a sys.path hack
(sys.path.insert(parent) + "from database import ...") rather than the
"from backend.X import ..." style the rest of the suite uses. That flat import
makes database/models.py's "from ..utils.capture_chords import ..." point
outside the package. It also aimed at the wrong module: it wants the service
layer, which raises ValueError, not the route handler, which converts that
into an HTTPException.
Result: the full suite goes from 0 collected to 148 passed. The one remaining
failure, test_progress.py::test_hf_progress_tracker, is pre-existing and
unrelated (tqdm patching) - it reproduces identically on an unpatched tree.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Update the README, MCP docs, in-app MCP page, i18n strings, landing copy,
docstrings and comments to match the renamed tools. CHANGELOG entries and
docs/plans are historical and left as they were.
Claude Desktop validates tool names against ^[a-zA-Z0-9_-]{1,64}$ and
rejects the whole tool list when any name contains a dot, so the Voicebox
MCP server was unusable there. Rename voicebox.speak, voicebox.transcribe,
voicebox.list_captures and voicebox.list_profiles to voicebox_speak,
voicebox_transcribe, voicebox_list_captures and voicebox_list_profiles.
Fixes#790
Adding seroval as a direct root dependency left the lockfile's
@tanstack/router-core/seroval entry pinned at the vulnerable 1.5.0, so the
only real consumer was still bundling the CVE-2026-59940 version. A bun
'overrides' entry forces every resolution to 1.5.3 and drops the unused
direct dependency.
Queries throughout the codebase filter on generations.profile_id,
generations.status, generations.created_at, story_items.story_id,
story_items.generation_id, generation_versions.generation_id, and
profile_samples.profile_id with every request. Without indexes SQLite
falls back to a full table scan; as history grows (hundreds or thousands
of generations) these scans become the dominant latency.
Changes:
- Add index=True on the most-queried FK and sort columns in models.py so
new installs get them from Base.metadata.create_all
- Add _migrate_add_indexes() called from run_migrations() so existing
installs get the same indexes on next startup (uses CREATE INDEX IF
NOT EXISTS — idempotent, <10 ms on any realistic dataset)
Switch from the default DELETE/ROLLBACK journal to WAL so concurrent
readers (SSE status polls, history queries) are not blocked while the
generation worker holds a write transaction. Set a 5-second busy
timeout to eliminate "database is locked" errors under brief write
contention.
Both PRAGMAs are applied via a custom creator function so every
connection in the pool gets the settings at open time, not just the
first one.
is_model_cached() marked a model as not-cached whenever any .incomplete
blob existed in its cache dir, even when a completed blob with the same
hash already sat next to it. A retried/concurrent download can leave
this orphan behind after the real transfer already finished, which made
the model appear perpetually "downloading" and re-trigger a full
re-download on every load.
Only .incomplete files with no matching completed blob now count as a
genuinely in-progress download.
Both speak surfaces built their GenerationRequest with a hardcoded "en"
fallback and never consulted the resolved profile, so a profile created
with language="fr" was still synthesised as English unless the caller
passed language= explicitly.
This hurts the MCP path most: an agent calling voicebox.speak has no way
to know the bound profile's language, so it cannot pass the argument
either. Every agent-triggered generation on a non-English profile came
out with an English accent.
The fallback chain is now explicit argument -> resolved profile's
language -> "en", which matches how engine and personality already
consult the resolved binding. The "en" backstop is kept so profiles with
no language set behave exactly as before.
Adds backend/tests/test_speak_language.py covering both surfaces: the
fallback, explicit-argument precedence, and the unchanged "en" default.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Addresses the budget comment on #1058 by the second route the review
offered -- defining the constant as the head budget rather than reserving
the suffix inside it.
TOAST_ERROR_BUDGET read as though it bounded `display`, but `display` is
the head plus " …", so it could be 402. Renamed to HEAD_BUDGET and
documented as bounding the message rather than the rendered string, with
the suffix now a named constant instead of a literal in the template.
Reserving the two characters was the alternative, but nothing downstream
has a hard limit -- the description box scrolls -- so it would have
shortened the message to satisfy a round number.
Verified: head <= 400 and display <= 402 on every truncating input,
including no-space text, a short first line, sentence-boundary backoff
and the real 4795-char error, with the untouched-when-not-truncated and
exact-`full` invariants still holding.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Follow-up to discussion_r3831361300 on #1058.
Both non-truncated return paths handed back the trimmed working copy, so
an error like " request timed out\n" came back altered even though
nothing had been omitted. That contradicted the stated intent that short
errors pass through untouched, and left `display` differing from `full`
for no reason.
`display` is now byte-identical to `full` whenever `truncated` is false —
the trimmed copy is only used for measuring against the budget and for
building the shortened head. Documented on the field.
Verified across padded short errors, clean short errors, empty and
whitespace-only input, and either side of the threshold: display === full
on every untruncated case, and `full` matches the input exactly in all of
them.
One visible consequence: with whitespace-pre-wrap on the description, an
error carrying leading or trailing newlines now renders with that blank
space. Trivial for the messages this sees in practice, and the
alternative was silently editing the text.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Two CodeRabbit findings on #1058.
condenseError trimmed the input before storing it in `full`, which is
documented as the untouched original and is what the Copy action hands
over. The trim now applies only to the working copy used for measuring
and cutting, so `full` is byte-for-byte what the server sent while
`display` and `omitted` still ignore surrounding blank space.
The Copy handler called navigator.clipboard.writeText with no guard.
Outside a secure context the property access itself throws, and
writeText rejects when permission is denied; neither was handled, so a
click could become an unhandled rejection with no sign that nothing was
copied. Both paths are now caught and reported, pointing at Settings ->
Logs as the fallback.
Not taken: aligning MIN_TO_CONDENSE with the 400-char budget. The gap is
deliberate -- cutting a 450-char error to 400 saves 50 characters in a
description that already scrolls, and no Copy action is needed there
because `display` holds the whole message. Documented in place.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A failed generation put the server's error straight into a toast. The
transformers "Unrecognized model" error is ~4.8KB, of which the first
sentence carries the meaning and the remaining 4.7KB is an alphabetical
list of every architecture it knows. In a 420px toast with
overflow-hidden and no scroll, that clipped the text at both ends and
pushed the close button off-screen: unreadable and undismissable.
- ToastDescription is capped at 40vh and scrolls, wraps on whitespace
and breaks long unspaced tokens so a path cannot widen the toast.
- The toast aligns to the start rather than centring, so the title
stays visible next to a tall description.
- condenseError() keeps the head of an oversized error, cutting at the
first newline or the last sentence end inside a 400-char budget, and
reports how many characters it dropped. Short errors pass through
untouched.
- When it does truncate, the toast offers a Copy action for the full
text and points at Settings -> Logs.
Verified against the real 4795-char error: 4795 -> 400 chars keeping
both meaningful sentences.
The hook moves to .tsx to render ToastAction, matching useAutoUpdater.tsx
which is a .tsx hook for the same reason. createElement was tried first
but this repo's ToastActionElement type is the older shadcn definition
(ReactElement<typeof ToastAction>) which only accepts JSX-constructed
elements.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
When the user disables auto_refine (LLM polish) in Settings, the Qwen
refinement model is no longer required for dictation to arm. Previously,
canRecord checked llmReady unconditionally, so useChordSync called
disable_hotkey whenever Qwen was not downloaded -- even if the user
never intended to use refinement.
Changes:
- useDictationReadiness: gate llmReady behind autoRefine in canRecord,
missing, and the polling predicates so hotkeys arm with just Whisper
STT when refinement is off.
- DictationReadinessChecklist: hide the LLM row when autoRefine is
false -- the checklist now shows only Whisper STT.
Closes#753
Without a trace the Settings GPU label and /health keep reporting CUDA
while generation runs on CPU, so the user has no way to confirm the
override engaged.
The override is documented in gpu-acceleration.mdx and listed as step 1 of
the get_torch_device() precedence in tts-generation.mdx, but grepping the
tree for VOICEBOX_FORCE_CPU matched only those two doc files - nothing read
it. Users whose GPU has no compiled kernels in the bundled PyTorch had no
way to fall back to CPU short of renaming the installed CUDA backend
directory.
Resolve it before torch is imported, so it still works when the installed
build is itself the reason CPU is wanted.
list_generations() fetched versions with one SELECT per generation on the
page (50 rows -> 51 queries). Add _get_versions_for_generations() which
loads all versions for the page in a single WHERE generation_id IN (...)
query and groups them in memory; the single-generation helper now
delegates to it so story item details behave identically.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <[email protected]>
The 19-row menu is taller than its 280px max-height, so arrow-key
navigation past the fold lost its highlight. Scroll the active row into
view on index change, list the delivery tags in the README, and give
[sarcastic] and [whispering] emoji that are not already used by
[chuckle] and [shush].
When the server outlives the Tauri app that spawned it (keep-running
mode, or a sidecar the next launch reuses), its stdout/stderr pipe has
no reader. Every later print()/tqdm write raises BrokenPipeError, so
POST /captures and /transcribe return "[Errno 32] Broken pipe".
Wrap stdout/stderr so they fall back to devnull on the first failed
write instead of raising.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The floating generate box is fixed at the bottom of the viewport, so
all of its Select dropdowns (voice profile, language, engine, effects)
opened downward into — or beyond — the window edge. Add side="top" to
each SelectContent so the menus appear above their trigger instead.
Co-authored-by: Claude Sonnet 4.6 <[email protected]>
key_from_str() had no arm for "Function", so it fell through to
None. Since build_chord propagates that as a hard Err via ?, binding
any chord containing fn made build_chord_bindings fail entirely —
HotkeyMonitor was never spawned, silently killing both push-to-talk
and toggle-to-talk until the chord was reverted.
Every other layer (keytap's macOS key tap, Key::Function itself, the
frontend's canonicalKeyFromEvent/displayLabelForKey) already handles
fn — only this string-to-Key bridge was missing the arm.
Fixes#941