Commit Graph
660 Commits
Author SHA1 Message Date
jamiepineandcapy-ai-staging[bot] 1a3942f3b3 style: ruff-format the new speak language tests 2026-10-04 00:00:22 +00:00
c339d2c324 fix(speak): honour the voice profile's language instead of forcing English
Both speak surfaces built their GenerationRequest with a hardcoded "en"
fallback and never consulted the resolved profile, so a profile created
with language="fr" was still synthesised as English unless the caller
passed language= explicitly.

This hurts the MCP path most: an agent calling voicebox.speak has no way
to know the bound profile's language, so it cannot pass the argument
either. Every agent-triggered generation on a non-English profile came
out with an English accent.

The fallback chain is now explicit argument -> resolved profile's
language -> "en", which matches how engine and personality already
consult the resolved binding. The "en" backstop is kept so profiles with
no language set behave exactly as before.

Adds backend/tests/test_speak_language.py covering both surfaces: the
fallback, explicit-argument precedence, and the unchanged "en" default.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:00:22 +00:00
jamiepineandcapy-ai-staging[bot] e440b44ae7 fix(ui): keep a first line that fits whole instead of re-cutting it at a sentence end 2026-10-04 00:00:15 +00:00
jamiepineandcapy-ai-staging[bot] 872121c0f7 style: apply biome import order and formatting to useGenerationProgress 2026-10-04 00:00:15 +00:00
8086c818b6 refactor(ui): name the head budget and the truncation suffix
Addresses the budget comment on #1058 by the second route the review
offered -- defining the constant as the head budget rather than reserving
the suffix inside it.

TOAST_ERROR_BUDGET read as though it bounded `display`, but `display` is
the head plus " …", so it could be 402. Renamed to HEAD_BUDGET and
documented as bounding the message rather than the rendered string, with
the suffix now a named constant instead of a literal in the template.

Reserving the two characters was the alternative, but nothing downstream
has a hard limit -- the description box scrolls -- so it would have
shortened the message to satisfy a round number.

Verified: head <= 400 and display <= 402 on every truncating input,
including no-space text, a short first line, sentence-boundary backoff
and the real 4795-char error, with the untouched-when-not-truncated and
exact-`full` invariants still holding.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:00:15 +00:00
e3b9c98977 fix(ui): return untruncated errors exactly as received
Follow-up to discussion_r3831361300 on #1058.

Both non-truncated return paths handed back the trimmed working copy, so
an error like "  request timed out\n" came back altered even though
nothing had been omitted. That contradicted the stated intent that short
errors pass through untouched, and left `display` differing from `full`
for no reason.

`display` is now byte-identical to `full` whenever `truncated` is false —
the trimmed copy is only used for measuring against the budget and for
building the shortened head. Documented on the field.

Verified across padded short errors, clean short errors, empty and
whitespace-only input, and either side of the threshold: display === full
on every untruncated case, and `full` matches the input exactly in all of
them.

One visible consequence: with whitespace-pre-wrap on the description, an
error carrying leading or trailing newlines now renders with that blank
space. Trivial for the messages this sees in practice, and the
alternative was silently editing the text.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:00:15 +00:00
96c6f9cad9 fix(ui): preserve the original error text and handle clipboard failures
Two CodeRabbit findings on #1058.

condenseError trimmed the input before storing it in `full`, which is
documented as the untouched original and is what the Copy action hands
over. The trim now applies only to the working copy used for measuring
and cutting, so `full` is byte-for-byte what the server sent while
`display` and `omitted` still ignore surrounding blank space.

The Copy handler called navigator.clipboard.writeText with no guard.
Outside a secure context the property access itself throws, and
writeText rejects when permission is denied; neither was handled, so a
click could become an unhandled rejection with no sign that nothing was
copied. Both paths are now caught and reported, pointing at Settings ->
Logs as the fallback.

Not taken: aligning MIN_TO_CONDENSE with the 400-char budget. The gap is
deliberate -- cutting a 450-char error to 400 saves 50 characters in a
description that already scrolls, and no Copy action is needed there
because `display` holds the whole message. Documented in place.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:00:15 +00:00
9ba2b33069 fix(ui): make long error toasts readable and copyable
A failed generation put the server's error straight into a toast. The
transformers "Unrecognized model" error is ~4.8KB, of which the first
sentence carries the meaning and the remaining 4.7KB is an alphabetical
list of every architecture it knows. In a 420px toast with
overflow-hidden and no scroll, that clipped the text at both ends and
pushed the close button off-screen: unreadable and undismissable.

- ToastDescription is capped at 40vh and scrolls, wraps on whitespace
  and breaks long unspaced tokens so a path cannot widen the toast.
- The toast aligns to the start rather than centring, so the title
  stays visible next to a tall description.
- condenseError() keeps the head of an oversized error, cutting at the
  first newline or the last sentence end inside a 400-char budget, and
  reports how many characters it dropped. Short errors pass through
  untouched.
- When it does truncate, the toast offers a Copy action for the full
  text and points at Settings -> Logs.

Verified against the real 4795-char error: 4795 -> 400 chars keeping
both meaningful sentences.

The hook moves to .tsx to render ToastAction, matching useAutoUpdater.tsx
which is a .tsx hook for the same reason. createElement was tried first
but this repo's ToastActionElement type is the older shadcn definition
(ReactElement<typeof ToastAction>) which only accepts JSX-constructed
elements.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-10-04 00:00:15 +00:00
devangkanthariaandcapy-ai-staging[bot] ca56137ca0 fix(settings): invalidate capture-readiness on auto_refine toggle (#753) 2026-10-04 00:00:08 +00:00
devangkanthariaandcapy-ai-staging[bot] 0e1e9922a8 fix(dictation): skip LLM readiness gate when auto_refine is off (#753)
When the user disables auto_refine (LLM polish) in Settings, the Qwen
refinement model is no longer required for dictation to arm. Previously,
canRecord checked llmReady unconditionally, so useChordSync called
disable_hotkey whenever Qwen was not downloaded -- even if the user
never intended to use refinement.

Changes:
- useDictationReadiness: gate llmReady behind autoRefine in canRecord,
  missing, and the polling predicates so hotkeys arm with just Whisper
  STT when refinement is off.
- DictationReadinessChecklist: hide the LLM row when autoRefine is
  false -- the checklist now shows only Whisper STT.

Closes #753
2026-10-04 00:00:08 +00:00
jamiepineandcapy-ai-staging[bot] 4a03cd7e77 fix(backend): log when VOICEBOX_FORCE_CPU forces the CPU device
Without a trace the Settings GPU label and /health keep reporting CUDA
while generation runs on CPU, so the user has no way to confirm the
override engaged.
2026-10-04 00:00:03 +00:00
Ousama Ben Younesandcapy-ai-staging[bot] 8b8c1429db fix(backend): honour VOICEBOX_FORCE_CPU in device selection
The override is documented in gpu-acceleration.mdx and listed as step 1 of
the get_torch_device() precedence in tts-generation.mdx, but grepping the
tree for VOICEBOX_FORCE_CPU matched only those two doc files - nothing read
it. Users whose GPU has no compiled kernels in the bundled PyTorch had no
way to fall back to CPU short of renaming the installed CUDA backend
directory.

Resolve it before torch is imported, so it still works when the installed
build is itself the reason CPU is wanted.
2026-10-04 00:00:03 +00:00
youtsuhoandcapy-ai-staging[bot] 029d4d3378 fix(backend): batch generation-version queries to eliminate N+1 in history listing
list_generations() fetched versions with one SELECT per generation on the
page (50 rows -> 51 queries). Add _get_versions_for_generations() which
loads all versions for the page in a single WHERE generation_id IN (...)
query and groups them in memory; the single-generation helper now
delegates to it so story item details behave identically.

Generated with Codebuff 🤖
Co-Authored-By: Codebuff <[email protected]>
2026-10-03 23:59:56 +00:00
jamiepineandcapy-ai-staging[bot] 214adc5ff2 fix: keep the active tag row in view and document the full tag set
The 19-row menu is taller than its 280px max-height, so arrow-key
navigation past the fold lost its highlight. Scroll the active row into
view on index change, list the delivery tags in the README, and give
[sarcastic] and [whispering] emoji that are not already used by
[chuckle] and [shush].
2026-10-03 23:59:50 +00:00
jamiepineandcapy-ai-staging[bot] 50d8d6a34f style: wrap TAG_REGEX to satisfy biome formatter 2026-10-03 23:59:50 +00:00
Margalitandcapy-ai-staging[bot] 4cb8ad3b1f Add 10 missing paralinguistic delivery tags to ParalinguisticInput 2026-10-03 23:59:50 +00:00
jamiepineandcapy-ai-staging[bot] c9f5f2c3e7 fix(server): route writelines through the pipe-safe write
writelines() was forwarded straight to the wrapped stream by __getattr__,
so it could still raise BrokenPipeError after the app's pipe closed.
2026-10-03 23:59:43 +00:00
848bb2e4c6 fix(server): stop failing requests after the app's stdout pipe closes
When the server outlives the Tauri app that spawned it (keep-running
mode, or a sidecar the next launch reuses), its stdout/stderr pipe has
no reader. Every later print()/tqdm write raises BrokenPipeError, so
POST /captures and /transcribe return "[Errno 32] Broken pipe".

Wrap stdout/stderr so they fall back to devnull on the first failed
write instead of raising.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-10-03 23:59:43 +00:00
Nikhil Jangidandcapy-ai-staging[bot] f5d63540a6 docs: tidy changelog spacing
Keep the unreleased changelog formatting consistent.
2026-10-03 23:59:33 +00:00
Nikhil Jangidandcapy-ai-staging[bot] f2271ae845 docs: note profile engine default fix
Document the generation engine precedence correction.
2026-10-03 23:59:33 +00:00
Nikhil Jangidandcapy-ai-staging[bot] cdeeb38be4 test: cover generation engine defaults
Verify omitted engines remain unset while explicit values are validated.
2026-10-03 23:59:33 +00:00
Nikhil Jangidandcapy-ai-staging[bot] d4face494f fix: honor profile engine for generation requests
Allow omitted engines to fall through to the profile default.
2026-10-03 23:59:33 +00:00
Alex SummerandGitHub 51f49dea19 fix(docs): update quick start guide to reflect correct terminology for voice profiles (#963) 2026-07-26 23:32:02 -07:00
80610d880e fix(ui): open FloatingGenerateBox selects upward to prevent clipping (fixes #928) (#936)
The floating generate box is fixed at the bottom of the viewport, so
all of its Select dropdowns (voice profile, language, engine, effects)
opened downward into — or beyond — the window edge. Add side="top" to
each SelectContent so the menus appear above their trigger instead.

Co-authored-by: Claude Sonnet 4.6 <[email protected]>
2026-07-26 23:31:59 -07:00
Sai Sridhar TarraandGitHub 397051ba44 fix(key_codes): add Function key arm so macOS fn can be bound to a chord (#950)
key_from_str() had no arm for "Function", so it fell through to
None. Since build_chord propagates that as a hard Err via ?, binding
any chord containing fn made build_chord_bindings fail entirely —
HotkeyMonitor was never spawned, silently killing both push-to-talk
and toggle-to-talk until the chord was reverted.

Every other layer (keytap's macOS key tap, Key::Function itself, the
frontend's canonicalKeyFromEvent/displayLabelForKey) already handles
fn — only this string-to-Key bridge was missing the arm.

Fixes #941
2026-07-26 23:31:54 -07:00
1ba935e83b fix(export): disambiguate export filenames with generation id (#956)
Export filenames were derived from only the first 30 characters of the
generation text. Generations with similar wording (a common workflow when
iterating on the same line) produced identical filenames, so exports
collided on disk — the browser appended " (1)"/" (2)" and users ended up
opening audio that didn't match the expected filename.

Append the first 8 chars of the generation id to the .wav and .voicebox.zip
export filenames, in both the backend Content-Disposition headers and the
frontend save-file hooks.

Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-07-26 23:31:51 -07:00
6a6f4643da fix(backend): guard avatar upload against a missing filename (#954)
`UploadFile.filename` can be None, and `Path(None)` raises TypeError. On the
avatar endpoint this happens before the try/except, so a filename-less upload
surfaces as an unhandled 500 instead of a clean response. Every other upload
handler already guards this with `file.filename or ""` (add_profile_sample,
transcription, generations); apply the same guard here.

Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-26 23:31:48 -07:00
44ef8daba3 fix(ui): parse naive-UTC timestamps consistently in formatAbsoluteDate (#953)
* fix(ui): parse naive-UTC timestamps consistently in formatAbsoluteDate

Backend timestamps are naive UTC (Python `datetime.utcnow()`) and are
serialized without a timezone suffix. `formatDate` already normalizes
these by appending `Z` before parsing, but `formatAbsoluteDate` called
`new Date(date)` directly. Per the ES spec, a timezone-less date-time
string is parsed as local time, so absolute timestamps were shown off by
the viewer's UTC offset (e.g. +9h in JST) — and disagreed with the
relative time rendered by `formatDate` for the same value (visible in the
Captures detail panel, which uses both on `capture.created_at`).

Extract the normalization into a shared `parseServerDate` helper and use
it in both formatters.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* docs(format): clarify parseServerDate comment on date-only vs date-time parsing

ECMAScript parses date-only strings ("2026-07-23") as UTC but timezone-less
date-time strings ("2026-07-23T10:00:00") as local time. The backend emits the
latter, which is the case this helper normalizes. Corrects the comment per PR
review feedback.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* docs(format): trim parseServerDate comment to match surrounding style

Reduce the multi-line explanation to a single why-comment consistent with
other utils comments.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-26 23:31:45 -07:00
AhmedIrfanandGitHub 68ece25a80 fix(mcp): add model_size parameter to voicebox.speak (#895)
The MCP voicebox.speak tool built its GenerationRequest without a
model_size, so every agent-triggered generation fell back to the schema
default ("1.7B"). There was no way to reach the 0.6B Qwen variant (or
TADA's 1B/3B) through MCP, and callers paid a model reload whenever the
requested size differed from what was already loaded.

Thread an optional model_size through voicebox.speak and the _speak
helper into GenerationRequest, mirroring the REST /generate surface.
Omitting it passes None, which generate_speech normalizes to the engine
default, so existing callers are unaffected.

Add backend/tests/test_mcp_speak.py covering the forwarded value, the
omitted-default path, and rejection of an invalid size.

Fixes #884
2026-07-26 23:31:41 -07:00
XariannandGitHub 1db0fdf645 fix(rocm): add MIOpen stability env vars to docker-compose.rocm.yml (#865)
Add three environment variables to prevent miopenStatusUnknownError and
system stuttering during inference on RDNA4 GPUs:

- MIOPEN_USER_DB_PATH: redirect MIOpen kernel cache to writable, persistent dir
- MIOPEN_CUSTOM_CACHE_DIR: same, for custom operator cache
- MIOPEN_FIND_MODE=FAST: use heuristic kernel selection instead of exhaustive
  benchmarking, which fails on RDNA4 with ptr: 0 size: 0 workspace warnings

MIOPEN_FIND_MODE=FAST does not affect output quality. All MIOpen kernel
variants produce the same numerical result; fast mode selects a known-good
kernel using heuristics instead of benchmarking every variant on the GPU.

Tested on RX 9070 (gfx1201) with ROCm 7.2 and PyTorch 2.12.1+rocm7.2.

Hardware note: tested on Ryzen 7 9800X3D + RX 9070 with Gigabyte B650M DS3H
motherboard. The exhaustive benchmarking failures may be related to IOMMU
behavior on this platform. This system was affected by an IOMMU bug patched
upstream in kernel 6.19.10, which may be a contributing factor. May not
affect all RDNA4 systems. MIOPEN_FIND_MODE=FAST is a safe default regardless.

Depends on PR #862 which fixes the broken ROCm Docker build.
2026-07-26 23:31:37 -07:00
Sai Sridhar TarraandGitHub 2a001fd63f fix(docker): normalize CRLF line endings on Windows checkouts (#951)
A Windows Git checkout with checkout-time CRLF conversion enabled
produces CRLF working-tree copies of package.json and
scripts/rocm-entrypoint.sh, breaking the Docker build two ways:

- The frontend stage's `sed -i -z 's/,\n  ]/…/'` is LF-anchored, so
  it doesn't match against \r\n and leaves an invalid trailing comma
  in package.json, which then fails JSON parsing in the vite build.
- The final stage copies rocm-entrypoint.sh straight from the build
  context; with a CRLF shebang the container reports the misleading
  "no such file or directory" for an entrypoint that plainly exists,
  because Linux can't resolve "/bin/sh\r" as an interpreter.

Add .gitattributes forcing LF for both files at checkout time, plus a
sed normalization step in each Dockerfile stage for resilience with
clones that predate the .gitattributes rule.

Fixes #915
2026-07-26 23:31:33 -07:00
Sai Sridhar TarraandGitHub e5813304ef fix(linux-audio): select monitor device by name instead of setting PULSE_SOURCE (#949)
std::env::set_var is not thread-safe on Unix (unsafe as of Rust 2024
edition) and calling it from a spawned capture thread while other
threads (tokio runtime, webview, Tauri plugins) may read the
environment is a data race risk. It also never got unset, so the
monitor source would leak into any later cpal/ALSA init in the same
process.

Replace the env-var indirection with direct device selection: when
pactl reports a monitor source name, search cpal's input device
enumeration for an exact match. Fall back to a substring match on
'monitor' (the original pactl-unavailable path), then the host's
default input device. This is the 'pass the source name directly to
cpal' option from the issue - no env mutation, no leakage between
capture sessions, and it still re-detects the current default sink's
monitor on every start_capture call.

Fixes #471
2026-07-26 23:31:30 -07:00
a5773807a5 fix(transcription): transcode uploads to WAV before STT (#957)
The /transcribe endpoint passed the raw uploaded file straight to the STT
backend (mlx_audio.stt -> miniaudio), which only decodes WAV/FLAC/MP3/Vorbis.
Browser recordings arrive as WebM/Opus (Chrome/Firefox MediaRecorder), so
web-mode dictation failed with 500 "unsupported file format". The Tauri app
was unaffected because WebKit produces MP4.

librosa already fully decodes the upload to compute duration (falling back to
audioread/ffmpeg for exotic containers), so re-encode that PCM to a temp WAV
and hand it to Whisper. WAV inputs pass through unchanged; the temp file is
cleaned up in the finally block.

Co-authored-by: Claude Opus 4.8 <[email protected]>
2026-07-26 23:31:27 -07:00
ed54347e81 Fix runaway MLX Qwen audio chunks (#964)
* fix runaway MLX Qwen audio chunks

* test: tighten runaway retry coverage

---------

Co-authored-by: huanghua01 <[email protected]>
2026-07-26 23:31:23 -07:00
624f6a2140 fix(tada): run voice-prompt encode under torch.inference_mode (#955)
Encoder.eval() alone still builds an autograd graph because parameters
require grad by default. On 8GB GPUs that ballooned TADA encode VRAM far
past the model footprint (issue 890). Wrap the encode forward in
inference_mode and add a unit test that asserts the flag is set.

Co-authored-by: fooSynaptic <[email protected]>
2026-07-26 23:31:20 -07:00
Kyle BuxtonandGitHub 669f85024f fix(macos): set Command flag on Cmd-down event so Electron apps paste (#952)
The macOS auto-paste sequence in `send_paste` posted the Cmd-down
CGEvent with flags = 0, setting the Command flag only on the V events.

On real hardware the Cmd keyDown (a flagsChanged event) already carries
kCGEventFlagMaskCommand, and Chromium/Electron builds its tracked
modifier state from that flag. With flags = 0 the tracker stays at
"Command up", so the following V matches neither the Cmd+V accelerator
(tracker says no modifier) nor plain-text insertion (the V event's own
flags say Command is held) — Electron drops it silently, producing no
paste and no stray "v". AppKit reads the V event's own modifier flags
and pastes regardless, which is why native apps (Notes, TextEdit,
Warp) worked while Electron targets (Slack, VS Code, VS Code Insiders)
silently no-op'd.

Setting kCGEventFlagMaskCommand on the Cmd-down event makes the
flagsChanged event well-formed; Chromium then registers Command=down
and Cmd+V matches. Likely fixes #762 and #643.
2026-07-26 23:31:17 -07:00
52f8d8dd38 Fix voice sample validation on Python 3.13 (fixes #852) (#853)
* Fix voice sample validation on Python 3.13

Python 3.13 removed audioop from the standard library, which broke reference
audio validation when adding voice samples. Add the audioop-lts backport for
3.13+ installs and bundle audioop in PyInstaller builds on the same versions.

* style(tests): satisfy Ruff import ordering

---------

Co-authored-by: Jamie Pine <[email protected]>
2026-07-20 22:35:23 -07:00
fb1e16d2ce fix(backend): return 404 instead of 500 for audio of failed generations (#893)
* fix(backend): return 404 instead of 500 for audio of failed generations

A failed generation stores an empty audio_path. resolve_storage_path("")
resolved to the data directory itself, which exists, so the route's 404
guard passed and FileResponse raised RuntimeError ("File at path .../data
is not a file"), surfacing as a 500.

- resolve_storage_path now returns None for empty paths
- audio routes check is_file() instead of exists() so directories never
  reach FileResponse
- GET /audio/{generation_id} reports "Generation failed; no audio
  available" when the generation status is failed

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(backend): reject empty Path objects in resolve_storage_path

Path("") is truthy, so the previous `if not path` guard only caught
None and empty strings. Callers such as database/migrations.py pass
Path objects, so an empty Path could still resolve to the data dir.
Check None separately and reject paths with no parts.

Also add regression tests asserting the version and sample audio
endpoints 404 when a stored path resolves to an existing directory
(guards the is_file() checks against regressing to exists()).

Addresses CodeRabbit review on PR #893.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style(tests): drop parentheses on pytest.fixture decorator (ruff PT001)

Co-Authored-By: Claude Fable 5 <[email protected]>

* style(tests): satisfy Ruff naming rule

---------

Co-authored-by: Claude Fable 5 <[email protected]>
Co-authored-by: Jamie Pine <[email protected]>
2026-07-20 22:35:04 -07:00
f750596364 fix(setup): install mlx-lm and mlx-audio in setup-python on Apple Silicon (#892)
* fix(setup): install mlx-lm and mlx-audio in setup-python on Apple Silicon

The dev setup installed requirements-mlx.txt but not mlx-audio/mlx-lm
themselves, so POST /transcribe failed on a fresh Apple Silicon setup
with "No module named 'mlx_audio'" (then "No module named 'mlx_lm'").
The release workflow already installs both with --no-deps (they declare
transformers>=5.x, conflicting with our <=4.57.x cap); mirror that in
the setup-python recipe with the same pins.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: add MLX smoke test for the --no-deps mlx-audio/mlx-lm install

mlx-audio and mlx-lm are installed --no-deps, so a missing transitive
dependency only surfaces at import time. Add a pytest-discoverable
smoke test (skipped off Apple Silicon) covering the exact entry points
the backend uses: mlx_audio.tts.load, mlx_audio.stt.load (which also
exercises the miniaudio dep from issue #505), mlx_lm.load/generate,
and a basic mlx.core op.

Co-Authored-By: Claude Fable 5 <[email protected]>

---------

Co-authored-by: Claude Fable 5 <[email protected]>
2026-07-20 22:26:46 -07:00
XariannandGitHub 91cd6df108 fix(rocm): unset empty HSA_OVERRIDE_GFX_VERSION before torch loads (#864)
Docker compose sets HSA_OVERRIDE_GFX_VERSION=${HSA_OVERRIDE_GFX_VERSION:-}
which results in an empty string when not provided. An empty string is
not the same as unset - ROCm treats it as 'force-empty' and no GPU is
detected, even natively supported ones (e.g. gfx1201 / RX 9070 on ROCm 7.2).

Pop the env var when it is empty, before torch loads, so ROCm auto-detects
the GPU correctly.

Tested on RX 9070 (gfx1201) with ROCm 7.2 and PyTorch 2.12.1+rocm7.2.
2026-07-20 22:26:25 -07:00
484a39ad9f fix(build): build voicebox-mcp shim sidecar on Windows (#794)
The Windows `build-server` just recipe only built and copied the
voicebox-server sidecar, omitting the voicebox-mcp stdio shim that the
Unix scripts/build-server.sh builds via `build_binary.py --shim`.

As a result `just build` on Windows produced only one sidecar and the
Tauri bundle step failed with:

    resource path `binaries\voicebox-mcp-<triple>.exe` doesn't exist

Build and copy the shim sidecar after the server, mirroring
build-server.sh. Hoist the triple/binaries-dir setup ahead of both
builds so the shim step reuses them.

Co-authored-by: namu.shin <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
2026-07-20 22:26:04 -07:00
3bfcbdc819 fix: the justfile syntax errror, GPU information cannot be read (#669)
Co-authored-by: xor_s <[email protected]>
2026-07-20 22:25:43 -07:00
Shabeer VPKandGitHub 190bc5e8a8 Update .dockerignore (#861)
whitelist ROCm entrypoint
2026-07-20 22:25:20 -07:00
jitendra kumar sainiandGitHub 80af641b61 docs: fix incorrect app identifier in CONTRIBUTING.md (#863)
The CUDA backend path used com.voicebox.app, but the actual Tauri identifier is sh.voicebox.app (as in tauri.conf.json and all other docs).
2026-07-20 21:19:03 -07:00
neuron-tech-aiandGitHub 6936789a88 Batch story item counts in list_stories to eliminate N+1 (#663)
list_stories() previously executed one COUNT(story_items) query per
story in a Python loop. With N stories that is N+1 round-trips to
SQLite regardless of list length. Replace with a single aggregated
GROUP BY query that fetches all counts at once, then populate each
StoryResponse from a dict lookup.
2026-07-20 21:14:46 -07:00
youtsuhoandGitHub f3eca34d33 fix: gate macOS-only keyboard_layout symbols behind cfg (#831)
* fix: gate macOS-only keyboard_layout symbols behind cfg to suppress dead_code warnings

* chore: sync bun.lock with package.json
2026-07-20 21:07:33 -07:00
30db291b01 fix(kokoro): add missing male Mandarin voices (#788)
Co-authored-by: Siddharth Chintawar <[email protected]>
Co-authored-by: Cursor <[email protected]>
2026-07-20 20:51:06 -07:00
Andrew BarnesandGitHub 71b51366bc Fix CUDA downloads on unsupported platforms (#770)
* Fix CUDA downloads on unsupported platforms

* fix: align CUDA status nullability

* fix: require CUDA download support flag
2026-07-20 19:22:18 -07:00
258b92c9c0 fix(offline): remove process-global offline guard from Qwen3 LLM loads (#924)
force_offline_if_cached flips HF_HUB_OFFLINE (env + huggingface_hub
constant + transformers._is_offline_mode) process-wide for the duration
of a cached LLM load, silently switching every concurrent model
download/load on other threads to offline mode. With default capture
settings (whisper-turbo STT + Qwen3 refinement + auto_refine) a first
run downloads several models concurrently, and a poisoned fetch
surfaces as "Can't load feature extractor..." (whisper) or
"Unrecognized model ... model_type" (Qwen3) rather than anything
mentioning offline mode.

These are the last two call sites of the guard — the same pattern was
deliberately removed app-wide in #524/#530 after identical failures,
and the 0.5.0 LLM backend reintroduced it. LLM loads now run with the
process's default HF_HUB_OFFLINE state, matching every other backend
(issue #462 precedent).

Fixes #841


Claude-Session: https://claude.ai/code/session_011iwL9AyeAWgz2jpgcHxJpC

Co-authored-by: Claude Fable 5 <[email protected]>
2026-07-20 18:47:10 -07:00
e6cf50c7f7 feat(i18n): add Korean (ko) locale with 559 translation keys (#814)
* feat(i18n): add Korean (ko) locale with 559 translation keys

* fix(i18n): complete Korean translations for current UI

---------

Co-authored-by: Jamie Pine <[email protected]>
2026-07-20 15:19:39 -07:00