Addresses the budget comment on #1058 by the second route the review
offered -- defining the constant as the head budget rather than reserving
the suffix inside it.
TOAST_ERROR_BUDGET read as though it bounded `display`, but `display` is
the head plus " …", so it could be 402. Renamed to HEAD_BUDGET and
documented as bounding the message rather than the rendered string, with
the suffix now a named constant instead of a literal in the template.
Reserving the two characters was the alternative, but nothing downstream
has a hard limit -- the description box scrolls -- so it would have
shortened the message to satisfy a round number.
Verified: head <= 400 and display <= 402 on every truncating input,
including no-space text, a short first line, sentence-boundary backoff
and the real 4795-char error, with the untouched-when-not-truncated and
exact-`full` invariants still holding.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Follow-up to discussion_r3831361300 on #1058.
Both non-truncated return paths handed back the trimmed working copy, so
an error like " request timed out\n" came back altered even though
nothing had been omitted. That contradicted the stated intent that short
errors pass through untouched, and left `display` differing from `full`
for no reason.
`display` is now byte-identical to `full` whenever `truncated` is false —
the trimmed copy is only used for measuring against the budget and for
building the shortened head. Documented on the field.
Verified across padded short errors, clean short errors, empty and
whitespace-only input, and either side of the threshold: display === full
on every untruncated case, and `full` matches the input exactly in all of
them.
One visible consequence: with whitespace-pre-wrap on the description, an
error carrying leading or trailing newlines now renders with that blank
space. Trivial for the messages this sees in practice, and the
alternative was silently editing the text.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Two CodeRabbit findings on #1058.
condenseError trimmed the input before storing it in `full`, which is
documented as the untouched original and is what the Copy action hands
over. The trim now applies only to the working copy used for measuring
and cutting, so `full` is byte-for-byte what the server sent while
`display` and `omitted` still ignore surrounding blank space.
The Copy handler called navigator.clipboard.writeText with no guard.
Outside a secure context the property access itself throws, and
writeText rejects when permission is denied; neither was handled, so a
click could become an unhandled rejection with no sign that nothing was
copied. Both paths are now caught and reported, pointing at Settings ->
Logs as the fallback.
Not taken: aligning MIN_TO_CONDENSE with the 400-char budget. The gap is
deliberate -- cutting a 450-char error to 400 saves 50 characters in a
description that already scrolls, and no Copy action is needed there
because `display` holds the whole message. Documented in place.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A failed generation put the server's error straight into a toast. The
transformers "Unrecognized model" error is ~4.8KB, of which the first
sentence carries the meaning and the remaining 4.7KB is an alphabetical
list of every architecture it knows. In a 420px toast with
overflow-hidden and no scroll, that clipped the text at both ends and
pushed the close button off-screen: unreadable and undismissable.
- ToastDescription is capped at 40vh and scrolls, wraps on whitespace
and breaks long unspaced tokens so a path cannot widen the toast.
- The toast aligns to the start rather than centring, so the title
stays visible next to a tall description.
- condenseError() keeps the head of an oversized error, cutting at the
first newline or the last sentence end inside a 400-char budget, and
reports how many characters it dropped. Short errors pass through
untouched.
- When it does truncate, the toast offers a Copy action for the full
text and points at Settings -> Logs.
Verified against the real 4795-char error: 4795 -> 400 chars keeping
both meaningful sentences.
The hook moves to .tsx to render ToastAction, matching useAutoUpdater.tsx
which is a .tsx hook for the same reason. createElement was tried first
but this repo's ToastActionElement type is the older shadcn definition
(ReactElement<typeof ToastAction>) which only accepts JSX-constructed
elements.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
When the user disables auto_refine (LLM polish) in Settings, the Qwen
refinement model is no longer required for dictation to arm. Previously,
canRecord checked llmReady unconditionally, so useChordSync called
disable_hotkey whenever Qwen was not downloaded -- even if the user
never intended to use refinement.
Changes:
- useDictationReadiness: gate llmReady behind autoRefine in canRecord,
missing, and the polling predicates so hotkeys arm with just Whisper
STT when refinement is off.
- DictationReadinessChecklist: hide the LLM row when autoRefine is
false -- the checklist now shows only Whisper STT.
Closes#753
Without a trace the Settings GPU label and /health keep reporting CUDA
while generation runs on CPU, so the user has no way to confirm the
override engaged.
The override is documented in gpu-acceleration.mdx and listed as step 1 of
the get_torch_device() precedence in tts-generation.mdx, but grepping the
tree for VOICEBOX_FORCE_CPU matched only those two doc files - nothing read
it. Users whose GPU has no compiled kernels in the bundled PyTorch had no
way to fall back to CPU short of renaming the installed CUDA backend
directory.
Resolve it before torch is imported, so it still works when the installed
build is itself the reason CPU is wanted.
list_generations() fetched versions with one SELECT per generation on the
page (50 rows -> 51 queries). Add _get_versions_for_generations() which
loads all versions for the page in a single WHERE generation_id IN (...)
query and groups them in memory; the single-generation helper now
delegates to it so story item details behave identically.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <[email protected]>
The 19-row menu is taller than its 280px max-height, so arrow-key
navigation past the fold lost its highlight. Scroll the active row into
view on index change, list the delivery tags in the README, and give
[sarcastic] and [whispering] emoji that are not already used by
[chuckle] and [shush].
When the server outlives the Tauri app that spawned it (keep-running
mode, or a sidecar the next launch reuses), its stdout/stderr pipe has
no reader. Every later print()/tqdm write raises BrokenPipeError, so
POST /captures and /transcribe return "[Errno 32] Broken pipe".
Wrap stdout/stderr so they fall back to devnull on the first failed
write instead of raising.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The floating generate box is fixed at the bottom of the viewport, so
all of its Select dropdowns (voice profile, language, engine, effects)
opened downward into — or beyond — the window edge. Add side="top" to
each SelectContent so the menus appear above their trigger instead.
Co-authored-by: Claude Sonnet 4.6 <[email protected]>
key_from_str() had no arm for "Function", so it fell through to
None. Since build_chord propagates that as a hard Err via ?, binding
any chord containing fn made build_chord_bindings fail entirely —
HotkeyMonitor was never spawned, silently killing both push-to-talk
and toggle-to-talk until the chord was reverted.
Every other layer (keytap's macOS key tap, Key::Function itself, the
frontend's canonicalKeyFromEvent/displayLabelForKey) already handles
fn — only this string-to-Key bridge was missing the arm.
Fixes#941
Export filenames were derived from only the first 30 characters of the
generation text. Generations with similar wording (a common workflow when
iterating on the same line) produced identical filenames, so exports
collided on disk — the browser appended " (1)"/" (2)" and users ended up
opening audio that didn't match the expected filename.
Append the first 8 chars of the generation id to the .wav and .voicebox.zip
export filenames, in both the backend Content-Disposition headers and the
frontend save-file hooks.
Co-authored-by: Claude Opus 4.8 <[email protected]>
`UploadFile.filename` can be None, and `Path(None)` raises TypeError. On the
avatar endpoint this happens before the try/except, so a filename-less upload
surfaces as an unhandled 500 instead of a clean response. Every other upload
handler already guards this with `file.filename or ""` (add_profile_sample,
transcription, generations); apply the same guard here.
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
* fix(ui): parse naive-UTC timestamps consistently in formatAbsoluteDate
Backend timestamps are naive UTC (Python `datetime.utcnow()`) and are
serialized without a timezone suffix. `formatDate` already normalizes
these by appending `Z` before parsing, but `formatAbsoluteDate` called
`new Date(date)` directly. Per the ES spec, a timezone-less date-time
string is parsed as local time, so absolute timestamps were shown off by
the viewer's UTC offset (e.g. +9h in JST) — and disagreed with the
relative time rendered by `formatDate` for the same value (visible in the
Captures detail panel, which uses both on `capture.created_at`).
Extract the normalization into a shared `parseServerDate` helper and use
it in both formatters.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* docs(format): clarify parseServerDate comment on date-only vs date-time parsing
ECMAScript parses date-only strings ("2026-07-23") as UTC but timezone-less
date-time strings ("2026-07-23T10:00:00") as local time. The backend emits the
latter, which is the case this helper normalizes. Corrects the comment per PR
review feedback.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
* docs(format): trim parseServerDate comment to match surrounding style
Reduce the multi-line explanation to a single why-comment consistent with
other utils comments.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
The MCP voicebox.speak tool built its GenerationRequest without a
model_size, so every agent-triggered generation fell back to the schema
default ("1.7B"). There was no way to reach the 0.6B Qwen variant (or
TADA's 1B/3B) through MCP, and callers paid a model reload whenever the
requested size differed from what was already loaded.
Thread an optional model_size through voicebox.speak and the _speak
helper into GenerationRequest, mirroring the REST /generate surface.
Omitting it passes None, which generate_speech normalizes to the engine
default, so existing callers are unaffected.
Add backend/tests/test_mcp_speak.py covering the forwarded value, the
omitted-default path, and rejection of an invalid size.
Fixes#884
Add three environment variables to prevent miopenStatusUnknownError and
system stuttering during inference on RDNA4 GPUs:
- MIOPEN_USER_DB_PATH: redirect MIOpen kernel cache to writable, persistent dir
- MIOPEN_CUSTOM_CACHE_DIR: same, for custom operator cache
- MIOPEN_FIND_MODE=FAST: use heuristic kernel selection instead of exhaustive
benchmarking, which fails on RDNA4 with ptr: 0 size: 0 workspace warnings
MIOPEN_FIND_MODE=FAST does not affect output quality. All MIOpen kernel
variants produce the same numerical result; fast mode selects a known-good
kernel using heuristics instead of benchmarking every variant on the GPU.
Tested on RX 9070 (gfx1201) with ROCm 7.2 and PyTorch 2.12.1+rocm7.2.
Hardware note: tested on Ryzen 7 9800X3D + RX 9070 with Gigabyte B650M DS3H
motherboard. The exhaustive benchmarking failures may be related to IOMMU
behavior on this platform. This system was affected by an IOMMU bug patched
upstream in kernel 6.19.10, which may be a contributing factor. May not
affect all RDNA4 systems. MIOPEN_FIND_MODE=FAST is a safe default regardless.
Depends on PR #862 which fixes the broken ROCm Docker build.
A Windows Git checkout with checkout-time CRLF conversion enabled
produces CRLF working-tree copies of package.json and
scripts/rocm-entrypoint.sh, breaking the Docker build two ways:
- The frontend stage's `sed -i -z 's/,\n ]/…/'` is LF-anchored, so
it doesn't match against \r\n and leaves an invalid trailing comma
in package.json, which then fails JSON parsing in the vite build.
- The final stage copies rocm-entrypoint.sh straight from the build
context; with a CRLF shebang the container reports the misleading
"no such file or directory" for an entrypoint that plainly exists,
because Linux can't resolve "/bin/sh\r" as an interpreter.
Add .gitattributes forcing LF for both files at checkout time, plus a
sed normalization step in each Dockerfile stage for resilience with
clones that predate the .gitattributes rule.
Fixes#915
std::env::set_var is not thread-safe on Unix (unsafe as of Rust 2024
edition) and calling it from a spawned capture thread while other
threads (tokio runtime, webview, Tauri plugins) may read the
environment is a data race risk. It also never got unset, so the
monitor source would leak into any later cpal/ALSA init in the same
process.
Replace the env-var indirection with direct device selection: when
pactl reports a monitor source name, search cpal's input device
enumeration for an exact match. Fall back to a substring match on
'monitor' (the original pactl-unavailable path), then the host's
default input device. This is the 'pass the source name directly to
cpal' option from the issue - no env mutation, no leakage between
capture sessions, and it still re-detects the current default sink's
monitor on every start_capture call.
Fixes#471
The /transcribe endpoint passed the raw uploaded file straight to the STT
backend (mlx_audio.stt -> miniaudio), which only decodes WAV/FLAC/MP3/Vorbis.
Browser recordings arrive as WebM/Opus (Chrome/Firefox MediaRecorder), so
web-mode dictation failed with 500 "unsupported file format". The Tauri app
was unaffected because WebKit produces MP4.
librosa already fully decodes the upload to compute duration (falling back to
audioread/ffmpeg for exotic containers), so re-encode that PCM to a temp WAV
and hand it to Whisper. WAV inputs pass through unchanged; the temp file is
cleaned up in the finally block.
Co-authored-by: Claude Opus 4.8 <[email protected]>
Encoder.eval() alone still builds an autograd graph because parameters
require grad by default. On 8GB GPUs that ballooned TADA encode VRAM far
past the model footprint (issue 890). Wrap the encode forward in
inference_mode and add a unit test that asserts the flag is set.
Co-authored-by: fooSynaptic <[email protected]>
The macOS auto-paste sequence in `send_paste` posted the Cmd-down
CGEvent with flags = 0, setting the Command flag only on the V events.
On real hardware the Cmd keyDown (a flagsChanged event) already carries
kCGEventFlagMaskCommand, and Chromium/Electron builds its tracked
modifier state from that flag. With flags = 0 the tracker stays at
"Command up", so the following V matches neither the Cmd+V accelerator
(tracker says no modifier) nor plain-text insertion (the V event's own
flags say Command is held) — Electron drops it silently, producing no
paste and no stray "v". AppKit reads the V event's own modifier flags
and pastes regardless, which is why native apps (Notes, TextEdit,
Warp) worked while Electron targets (Slack, VS Code, VS Code Insiders)
silently no-op'd.
Setting kCGEventFlagMaskCommand on the Cmd-down event makes the
flagsChanged event well-formed; Chromium then registers Command=down
and Cmd+V matches. Likely fixes#762 and #643.
* Fix voice sample validation on Python 3.13
Python 3.13 removed audioop from the standard library, which broke reference
audio validation when adding voice samples. Add the audioop-lts backport for
3.13+ installs and bundle audioop in PyInstaller builds on the same versions.
* style(tests): satisfy Ruff import ordering
---------
Co-authored-by: Jamie Pine <[email protected]>
* fix(backend): return 404 instead of 500 for audio of failed generations
A failed generation stores an empty audio_path. resolve_storage_path("")
resolved to the data directory itself, which exists, so the route's 404
guard passed and FileResponse raised RuntimeError ("File at path .../data
is not a file"), surfacing as a 500.
- resolve_storage_path now returns None for empty paths
- audio routes check is_file() instead of exists() so directories never
reach FileResponse
- GET /audio/{generation_id} reports "Generation failed; no audio
available" when the generation status is failed
Co-Authored-By: Claude Fable 5 <[email protected]>
* fix(backend): reject empty Path objects in resolve_storage_path
Path("") is truthy, so the previous `if not path` guard only caught
None and empty strings. Callers such as database/migrations.py pass
Path objects, so an empty Path could still resolve to the data dir.
Check None separately and reject paths with no parts.
Also add regression tests asserting the version and sample audio
endpoints 404 when a stored path resolves to an existing directory
(guards the is_file() checks against regressing to exists()).
Addresses CodeRabbit review on PR #893.
Co-Authored-By: Claude Fable 5 <[email protected]>
* style(tests): drop parentheses on pytest.fixture decorator (ruff PT001)
Co-Authored-By: Claude Fable 5 <[email protected]>
* style(tests): satisfy Ruff naming rule
---------
Co-authored-by: Claude Fable 5 <[email protected]>
Co-authored-by: Jamie Pine <[email protected]>
* fix(setup): install mlx-lm and mlx-audio in setup-python on Apple Silicon
The dev setup installed requirements-mlx.txt but not mlx-audio/mlx-lm
themselves, so POST /transcribe failed on a fresh Apple Silicon setup
with "No module named 'mlx_audio'" (then "No module named 'mlx_lm'").
The release workflow already installs both with --no-deps (they declare
transformers>=5.x, conflicting with our <=4.57.x cap); mirror that in
the setup-python recipe with the same pins.
Co-Authored-By: Claude Fable 5 <[email protected]>
* test: add MLX smoke test for the --no-deps mlx-audio/mlx-lm install
mlx-audio and mlx-lm are installed --no-deps, so a missing transitive
dependency only surfaces at import time. Add a pytest-discoverable
smoke test (skipped off Apple Silicon) covering the exact entry points
the backend uses: mlx_audio.tts.load, mlx_audio.stt.load (which also
exercises the miniaudio dep from issue #505), mlx_lm.load/generate,
and a basic mlx.core op.
Co-Authored-By: Claude Fable 5 <[email protected]>
---------
Co-authored-by: Claude Fable 5 <[email protected]>
Docker compose sets HSA_OVERRIDE_GFX_VERSION=${HSA_OVERRIDE_GFX_VERSION:-}
which results in an empty string when not provided. An empty string is
not the same as unset - ROCm treats it as 'force-empty' and no GPU is
detected, even natively supported ones (e.g. gfx1201 / RX 9070 on ROCm 7.2).
Pop the env var when it is empty, before torch loads, so ROCm auto-detects
the GPU correctly.
Tested on RX 9070 (gfx1201) with ROCm 7.2 and PyTorch 2.12.1+rocm7.2.
The Windows `build-server` just recipe only built and copied the
voicebox-server sidecar, omitting the voicebox-mcp stdio shim that the
Unix scripts/build-server.sh builds via `build_binary.py --shim`.
As a result `just build` on Windows produced only one sidecar and the
Tauri bundle step failed with:
resource path `binaries\voicebox-mcp-<triple>.exe` doesn't exist
Build and copy the shim sidecar after the server, mirroring
build-server.sh. Hoist the triple/binaries-dir setup ahead of both
builds so the shim step reuses them.
Co-authored-by: namu.shin <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
list_stories() previously executed one COUNT(story_items) query per
story in a Python loop. With N stories that is N+1 round-trips to
SQLite regardless of list length. Replace with a single aggregated
GROUP BY query that fetches all counts at once, then populate each
StoryResponse from a dict lookup.
force_offline_if_cached flips HF_HUB_OFFLINE (env + huggingface_hub
constant + transformers._is_offline_mode) process-wide for the duration
of a cached LLM load, silently switching every concurrent model
download/load on other threads to offline mode. With default capture
settings (whisper-turbo STT + Qwen3 refinement + auto_refine) a first
run downloads several models concurrently, and a poisoned fetch
surfaces as "Can't load feature extractor..." (whisper) or
"Unrecognized model ... model_type" (Qwen3) rather than anything
mentioning offline mode.
These are the last two call sites of the guard — the same pattern was
deliberately removed app-wide in #524/#530 after identical failures,
and the 0.5.0 LLM backend reintroduced it. LLM loads now run with the
process's default HF_HUB_OFFLINE state, matching every other backend
(issue #462 precedent).
Fixes#841
Claude-Session: https://claude.ai/code/session_011iwL9AyeAWgz2jpgcHxJpC
Co-authored-by: Claude Fable 5 <[email protected]>
* feat(i18n): add Korean (ko) locale with 559 translation keys
* fix(i18n): complete Korean translations for current UI
---------
Co-authored-by: Jamie Pine <[email protected]>
Adds Spanish as a UI display language, matching the existing
4-locale pattern (en, ja, zh-CN, zh-TW) with full key parity.
- app/src/i18n/locales/es/translation.json: 832 strings across 18
namespaces, translated from the en master. Keys, {{interpolation}}
placeholders, <code>/<path>/<link>/<strong> tags and _one/_other
plurals preserved. Brand/model names (Whisper, Qwen3, CUDA, MCP…)
left untranslated by design.
- app/src/i18n/index.ts: register `es` in SUPPORTED_LANGUAGES and
resources; the language switcher and LanguageCode derive automatically.
- app/src/lib/utils/format.ts: wire the date-fns `es` locale for
relative-date formatting.
Verified: key parity 832/832 (no missing/extra, placeholders & tags
intact), biome check clean, app+web typecheck pass, build:web succeeds.
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Co-authored-by: Jamie Pine <[email protected]>
The dictate pill window is built hidden at setup and the frontend emits
dictate:hide as soon as it mounts. The handler calls
set_ignore_cursor_events(true) on a window GTK has never realized, and
tao's CursorIgnoreEvents path unwraps the missing GdkWindow
(tao-0.34.5 event_loop.rs:449), panicking inside a glib dispatch that
cannot unwind — the process aborts within seconds of launch on Linux.
The click-through toggle exists as a macOS workaround for transparent
always-on-top NSWindows lingering as invisible click targets; it was
never needed on Linux. Gate all three call sites so Linux never toggles
it: the true/false pair stays balanced (never set, never unset), and
macOS/Windows builds are unchanged.
Co-authored-by: Claude Fable 5 <[email protected]>
TaskManager.error_download() intentionally keeps a failed task in the
active list (status="error") so /tasks/active can surface the error
and retry UI — but /models/status derived its "downloading" flag from
the same unfiltered list. One failed download therefore showed the
model as downloading:true / downloaded:false for the life of the
process, masking the model's real cache state (even a fully valid
on-disk cache) until an app restart. Likely behind endless-spinner
reports like #181 and the restart-fixes-it pattern in #883.
Add TaskManager.get_pending_downloads() (downloading/extracting only)
and use it in /models/status; /tasks/active behavior is unchanged.
Fixes#925
Claude-Session: https://claude.ai/code/session_011iwL9AyeAWgz2jpgcHxJpC
Co-authored-by: Claude Fable 5 <[email protected]>