mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
Merge origin/main into prep/pr-1031
This commit is contained in:
@@ -121,7 +121,7 @@ POST /generate
|
||||
Shipped 2026-04-25 (PR #544). Voicebox went from a voice-cloning studio to a full voice studio — dictation in, agent speech out, a local LLM in the middle.
|
||||
|
||||
- **Dictation** — global hotkey capture (push-to-talk + toggle chords), on-screen pill with live state, auto-paste into the focused field with clipboard save/restore, chord-picker UI. Scoped Accessibility permission (transcripts still land if paste is denied).
|
||||
- **MCP server** at `http://127.0.0.1:17493/mcp` — `voicebox.speak` / `.transcribe` / `.list_captures` / `.list_profiles`. Streamable HTTP primary transport, stdio sidecar shim, per-client voice binding via `X-Voicebox-Client-Id`. Speaking pill always shows agent-initiated output.
|
||||
- **MCP server** at `http://127.0.0.1:17493/mcp` — `voicebox_speak` / `voicebox_transcribe` / `voicebox_list_captures` / `voicebox_list_profiles`. Streamable HTTP primary transport, stdio sidecar shim, per-client voice binding via `X-Voicebox-Client-Id`. Speaking pill always shows agent-initiated output.
|
||||
- **Personality** — voice profiles carry an optional ≤2000-char persona. Compose (shuffle an in-character line) and Speak-in-character (rewrite input before TTS), both on a local Qwen3 LLM that doubles as the refinement model.
|
||||
- **Refinement** — on-device Qwen3 strips fillers, fixes punctuation, optional self-correction rewrites; Whisper hallucination-loop stripping at a 6-token threshold; per-capture flag snapshots; model picker (0.6B / 1.7B / 4B).
|
||||
- **`POST /speak` REST wrapper** and **i18next foundation** (English + zh-CN) also landed.
|
||||
|
||||
@@ -105,15 +105,15 @@ for the shim to connect.
|
||||
|
||||
| Tool | Use |
|
||||
|---|---|
|
||||
| `voicebox.speak` | Speak text in a voice profile. Returns a `generation_id` to poll. |
|
||||
| `voicebox.transcribe` | Whisper transcription of base64 audio or an absolute local path. |
|
||||
| `voicebox.list_captures` | Recent captures with transcripts, paginated. |
|
||||
| `voicebox.list_profiles` | Available voice profiles (cloned + preset). |
|
||||
| `voicebox_speak` | Speak text in a voice profile. Returns a `generation_id` to poll. |
|
||||
| `voicebox_transcribe` | Whisper transcription of base64 audio or an absolute local path. |
|
||||
| `voicebox_list_captures` | Recent captures with transcripts, paginated. |
|
||||
| `voicebox_list_profiles` | Available voice profiles (cloned + preset). |
|
||||
|
||||
### `voicebox.speak`
|
||||
### `voicebox_speak`
|
||||
|
||||
```ts
|
||||
voicebox.speak({
|
||||
voicebox_speak({
|
||||
text: "Deploy complete.",
|
||||
profile?: "Morgan", // name or id; falls back to per-client binding, then default
|
||||
engine?: "qwen", // qwen | qwen_custom_voice | luxtts | chatterbox | chatterbox_turbo | tada | kokoro
|
||||
@@ -138,10 +138,10 @@ Returns:
|
||||
- **Persona mode** — `personality: true` and the profile must have a personality prompt set.
|
||||
The LLM rewrites the text in character before TTS. See [Voice Personalities](/overview/voice-personalities).
|
||||
|
||||
### `voicebox.transcribe`
|
||||
### `voicebox_transcribe`
|
||||
|
||||
```ts
|
||||
voicebox.transcribe({
|
||||
voicebox_transcribe({
|
||||
audio_base64?: "<base64>", // exactly one of these two
|
||||
audio_path?: "/absolute/path/to/file.wav",
|
||||
language?: "en",
|
||||
@@ -151,18 +151,18 @@ voicebox.transcribe({
|
||||
|
||||
Returns `{ text, duration, language, model }`. 200 MB ceiling on either path.
|
||||
|
||||
### `voicebox.list_captures`
|
||||
### `voicebox_list_captures`
|
||||
|
||||
`{ limit?: 20, offset?: 0 }` → `{ captures: [...], total }`. `limit` is
|
||||
clamped to `1..=200`.
|
||||
|
||||
### `voicebox.list_profiles`
|
||||
### `voicebox_list_profiles`
|
||||
|
||||
No args → `{ profiles: [{ id, name, voice_type, language, has_personality }] }`.
|
||||
|
||||
## Voice resolution
|
||||
|
||||
Every call to `voicebox.speak` (and `POST /speak`) resolves the voice profile
|
||||
Every call to `voicebox_speak` (and `POST /speak`) resolves the voice profile
|
||||
in this order:
|
||||
|
||||
<Steps>
|
||||
@@ -195,7 +195,7 @@ carries:
|
||||
| `label` | Display name in the Settings UI (e.g. "Claude Code"). |
|
||||
| `profile_id` | The voice this client uses when `profile` isn't passed. |
|
||||
| `default_engine` | Override the TTS engine for this client. |
|
||||
| `default_personality` | When true, `voicebox.speak` routes through the profile's personality LLM (rewrite) by default. |
|
||||
| `default_personality` | When true, `voicebox_speak` routes through the profile's personality LLM (rewrite) by default. |
|
||||
| `last_seen_at` | Last time the server saw a request from this client. |
|
||||
|
||||
`last_seen_at` is stamped automatically by middleware on every `/mcp/*`
|
||||
@@ -239,8 +239,8 @@ agent:
|
||||
npx @modelcontextprotocol/inspector http://127.0.0.1:17493/mcp
|
||||
```
|
||||
|
||||
Start with `voicebox.list_profiles` to confirm wiring, then
|
||||
`voicebox.speak` for end-to-end — you should hear audio and see the
|
||||
Start with `voicebox_list_profiles` to confirm wiring, then
|
||||
`voicebox_speak` for end-to-end — you should hear audio and see the
|
||||
generation land in the Captures tab.
|
||||
|
||||
<Callout type="info">
|
||||
@@ -262,7 +262,7 @@ generation land in the Captures tab.
|
||||
boundary. If you're scripting against a shared host, prefer
|
||||
`audio_base64` so you don't have to think about path sandboxing.
|
||||
- **Voice cloning consent applies.** See [Voice Cloning](/overview/voice-cloning#limitations)
|
||||
— an agent being able to call `voicebox.speak` in someone's voice
|
||||
— an agent being able to call `voicebox_speak` in someone's voice
|
||||
doesn't change the ethics of whose voices you clone.
|
||||
|
||||
## Implementation notes
|
||||
|
||||
@@ -428,12 +428,14 @@ bun run tauri build
|
||||
|
||||
```bash
|
||||
# macOS
|
||||
rm ~/Library/Application\ Support/sh.voicebox.app/data/voicebox.db
|
||||
rm ~/Library/Application\ Support/sh.voicebox.app/data/voicebox.db*
|
||||
|
||||
# Windows
|
||||
del %APPDATA%\sh.voicebox.app\data\voicebox.db
|
||||
del %APPDATA%\sh.voicebox.app\data\voicebox.db*
|
||||
```
|
||||
|
||||
The wildcard also removes the `voicebox.db-wal` and `voicebox.db-shm` sidecar files that SQLite keeps next to the database in WAL mode, so the reset starts completely clean.
|
||||
|
||||
Restart the app to create a fresh database.
|
||||
|
||||
## Model Issues
|
||||
|
||||
@@ -128,7 +128,7 @@ until you flip it back off.
|
||||
- **Agents that speak in a voice you own.** Combine the persona toggle with
|
||||
the built-in [MCP Server](/overview/mcp-server) so Claude Code, Cursor,
|
||||
Cline, or any MCP-aware agent can talk back through a profile with a
|
||||
personality. The agent calls `voicebox.speak({ text, profile, personality:
|
||||
personality. The agent calls `voicebox_speak({ text, profile, personality:
|
||||
true })` and Voicebox rewrites the text in character before speaking.
|
||||
- **Interactive characters.** Games, narrative tools, accessibility
|
||||
experiences. A character with a personality description plus a cloned
|
||||
@@ -151,7 +151,7 @@ Personalities are accessible via REST:
|
||||
| `POST` | `/generate` | Include `personality: true` to run input text through the personality LLM before TTS. Same for `POST /speak`. |
|
||||
|
||||
`POST /generate` with `personality: true` is the same primitive MCP's
|
||||
`voicebox.speak` tool uses when you pass `personality: true`. Scripts and
|
||||
`voicebox_speak` tool uses when you pass `personality: true`. Scripts and
|
||||
agents can use it directly.
|
||||
|
||||
## Limits and gotchas
|
||||
|
||||
Reference in New Issue
Block a user