mirror of
https://github.com/jamiepine/voicebox.git
synced 2026-10-03 17:15:19 -07:00
docs: refer to the MCP tools by their new underscore names
Update the README, MCP docs, in-app MCP page, i18n strings, landing copy, docstrings and comments to match the renamed tools. CHANGELOG entries and docs/plans are historical and left as they were.
This commit is contained in:
committed by
capy-ai-staging[bot]
parent
8990104081
commit
37af180887
@@ -80,7 +80,7 @@ The two cloud incumbents sit on opposite halves of the voice I/O loop — Eleven
|
||||
- **Unlimited length** — auto-chunking with crossfade for scripts, articles, and chapters
|
||||
- **Stories editor** — multi-track timeline for conversations, podcasts, and narratives
|
||||
- **Voice input** — global dictation hotkey with push-to-talk and toggle modes, accessibility-verified auto-paste on macOS, in-app mic on every text field, Whisper-based STT
|
||||
- **Agent voice output** — one tool call (`voicebox.speak`) and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you've cloned
|
||||
- **Agent voice output** — one tool call (`voicebox_speak`) and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you've cloned
|
||||
- **Voice personalities** — attach a free-form persona to any voice profile, then Compose, Rewrite, or Respond via a bundled local LLM — agents can invoke the same modes over MCP
|
||||
- **API-first** — REST API plus a built-in MCP server for integrating voice I/O into your own apps and agents
|
||||
- **Native performance** — built with Tauri (Rust), not Electron
|
||||
@@ -234,7 +234,7 @@ Every agent gets a voice. One tool call and any MCP-aware agent can speak to you
|
||||
|
||||
```ts
|
||||
// In any MCP-aware agent:
|
||||
await voicebox.speak({
|
||||
await voicebox_speak({
|
||||
text: "Deploy complete.",
|
||||
profile: "Morgan",
|
||||
});
|
||||
@@ -254,7 +254,7 @@ Attach a free-form personality to any voice profile — who this voice is, how t
|
||||
- **Compose** — a shuffle button that drops a fresh in-character line into the textarea; edit and speak, or click again for a different take
|
||||
- **Speak in character** — a toggle that routes your input text through the personality LLM to be rewritten in their voice before TTS
|
||||
|
||||
Agents can reach the same rewrite path over MCP by passing `personality: true` to `voicebox.speak`, turning the tool into a text-in → personality-LLM → TTS pipeline. The same LLM backs dictation's refinement step — one LLM in the app, one model cache, one GPU-memory footprint.
|
||||
Agents can reach the same rewrite path over MCP by passing `personality: true` to `voicebox_speak`, turning the tool into a text-in → personality-LLM → TTS pipeline. The same LLM backs dictation's refinement step — one LLM in the app, one model cache, one GPU-memory footprint.
|
||||
|
||||
**Local LLM options:** Qwen3 0.6B / 1.7B / 4B, sharing the TTS runtime (MLX on Apple Silicon, PyTorch elsewhere).
|
||||
|
||||
@@ -347,11 +347,11 @@ claude mcp add voicebox \
|
||||
}
|
||||
```
|
||||
|
||||
Four tools ship: `voicebox.speak`, `voicebox.transcribe`, `voicebox.list_captures`, `voicebox.list_profiles`. Per-client voice bindings are managed in **Voicebox → Settings → MCP**. See the [full MCP guide](docs/content/docs/overview/mcp-server.mdx) for tool signatures, resolution precedence, the speaking-pill contract, and security notes.
|
||||
Four tools ship: `voicebox_speak`, `voicebox_transcribe`, `voicebox_list_captures`, `voicebox_list_profiles`. Per-client voice bindings are managed in **Voicebox → Settings → MCP**. See the [full MCP guide](docs/content/docs/overview/mcp-server.mdx) for tool signatures, resolution precedence, the speaking-pill contract, and security notes.
|
||||
|
||||
```ts
|
||||
// In any MCP-aware agent:
|
||||
await voicebox.speak({
|
||||
await voicebox_speak({
|
||||
text: "Tests passing. Ready to merge.",
|
||||
profile: "Morgan", // optional — falls back to the per-client binding
|
||||
personality: true, // optional — rewrites text through the profile's personality LLM first
|
||||
|
||||
@@ -276,19 +276,19 @@ export function MCPPage() {
|
||||
<h3 className="text-sm font-semibold">{t('settings.mcp.sidebar.toolsTitle')}</h3>
|
||||
<ul className="text-sm text-muted-foreground space-y-1.5 leading-relaxed">
|
||||
<li>
|
||||
<code className="text-accent">voicebox.speak</code>
|
||||
<code className="text-accent">voicebox_speak</code>
|
||||
<div>{t('settings.mcp.sidebar.tools.speak')}</div>
|
||||
</li>
|
||||
<li>
|
||||
<code className="text-accent">voicebox.transcribe</code>
|
||||
<code className="text-accent">voicebox_transcribe</code>
|
||||
<div>{t('settings.mcp.sidebar.tools.transcribe')}</div>
|
||||
</li>
|
||||
<li>
|
||||
<code className="text-accent">voicebox.list_captures</code>
|
||||
<code className="text-accent">voicebox_list_captures</code>
|
||||
<div>{t('settings.mcp.sidebar.tools.listCaptures')}</div>
|
||||
</li>
|
||||
<li>
|
||||
<code className="text-accent">voicebox.list_profiles</code>
|
||||
<code className="text-accent">voicebox_list_profiles</code>
|
||||
<div>{t('settings.mcp.sidebar.tools.listProfiles')}</div>
|
||||
</li>
|
||||
</ul>
|
||||
|
||||
@@ -1055,7 +1055,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "Default voice",
|
||||
"description": "Used when an agent calls voicebox.speak without a specific profile and has no per-client binding.",
|
||||
"description": "Used when an agent calls voicebox_speak without a specific profile and has no per-client binding.",
|
||||
"label": "Default playback voice",
|
||||
"labelHint": "Shared with the Captures-tab 'Play as voice' dropdown — one default voice for passive playback.",
|
||||
"none": "(none)"
|
||||
|
||||
@@ -1055,7 +1055,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "Voz predeterminada",
|
||||
"description": "Se usa cuando un agente llama a voicebox.speak sin un perfil específico y no tiene una vinculación por cliente.",
|
||||
"description": "Se usa cuando un agente llama a voicebox_speak sin un perfil específico y no tiene una vinculación por cliente.",
|
||||
"label": "Voz de reproducción predeterminada",
|
||||
"labelHint": "Compartida con el desplegable 'Reproducir como voz' de la pestaña Capturas: una voz predeterminada para la reproducción pasiva.",
|
||||
"none": "(ninguna)"
|
||||
|
||||
@@ -1050,7 +1050,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "Voix par défaut",
|
||||
"description": "Utilisée quand un agent appelle voicebox.speak sans profil spécifique et sans liaison par client.",
|
||||
"description": "Utilisée quand un agent appelle voicebox_speak sans profil spécifique et sans liaison par client.",
|
||||
"label": "Voix de lecture par défaut",
|
||||
"labelHint": "Partagé avec la liste déroulante « Lire avec » de l'onglet Captures — une voix par défaut pour la lecture passive.",
|
||||
"none": "(aucune)"
|
||||
|
||||
@@ -1055,7 +1055,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "Voce predefinita",
|
||||
"description": "Utilizzata quando un agente chiama voicebox.speak senza specificare un profilo e non ha un'associazione per singolo client.",
|
||||
"description": "Utilizzata quando un agente chiama voicebox_speak senza specificare un profilo e non ha un'associazione per singolo client.",
|
||||
"label": "Voce di riproduzione predefinita",
|
||||
"labelHint": "Condivisa con il menu a discesa 'Riproduci come voce' della scheda Acquisizioni — una sola voce predefinita per la riproduzione passiva.",
|
||||
"none": "(nessuna)"
|
||||
|
||||
@@ -1050,7 +1050,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "デフォルトボイス",
|
||||
"description": "エージェントが特定のプロファイルを指定せず、クライアントごとのバインディングもない状態で voicebox.speak を呼び出したときに使われます。",
|
||||
"description": "エージェントが特定のプロファイルを指定せず、クライアントごとのバインディングもない状態で voicebox_speak を呼び出したときに使われます。",
|
||||
"label": "デフォルトの再生ボイス",
|
||||
"labelHint": "キャプチャタブの「ボイスで再生」ドロップダウンと共有 — パッシブ再生用に 1 つのデフォルトボイスを設定します。",
|
||||
"none": "(なし)"
|
||||
|
||||
@@ -1050,7 +1050,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "기본 음성",
|
||||
"description": "에이전트가 특정 프로필 없이 voicebox.speak를 호출하고 클라이언트별 바인딩도 없을 때 사용됩니다.",
|
||||
"description": "에이전트가 특정 프로필 없이 voicebox_speak를 호출하고 클라이언트별 바인딩도 없을 때 사용됩니다.",
|
||||
"label": "기본 재생 음성",
|
||||
"labelHint": "캡처 탭의 'Play as 음성' 드롭다운과 공유 — 수동 재생용 기본 음성입니다.",
|
||||
"none": "(없음)"
|
||||
|
||||
@@ -1050,7 +1050,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "Voz padrão",
|
||||
"description": "Usada quando um agente chama voicebox.speak sem um perfil específico e não tem vínculo por cliente.",
|
||||
"description": "Usada quando um agente chama voicebox_speak sem um perfil específico e não tem vínculo por cliente.",
|
||||
"label": "Voz de reprodução padrão",
|
||||
"labelHint": "Compartilhada com o menu 'Reproduzir como voz' da aba Capturas — uma voz padrão para reprodução passiva.",
|
||||
"none": "(nenhuma)"
|
||||
|
||||
@@ -1050,7 +1050,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "默认声音",
|
||||
"description": "当代理调用 voicebox.speak 但未指定具体档案、且没有按客户端绑定时使用。",
|
||||
"description": "当代理调用 voicebox_speak 但未指定具体档案、且没有按客户端绑定时使用。",
|
||||
"label": "默认播放声音",
|
||||
"labelHint": "与「捕获」标签页的「播放为」下拉菜单共享——被动播放的统一默认声音。",
|
||||
"none": "(无)"
|
||||
|
||||
@@ -1050,7 +1050,7 @@
|
||||
},
|
||||
"defaultVoice": {
|
||||
"title": "預設聲音",
|
||||
"description": "當代理呼叫 voicebox.speak 卻未指定聲音檔案,且沒有對應客戶端綁定時使用。",
|
||||
"description": "當代理呼叫 voicebox_speak 卻未指定聲音檔案,且沒有對應客戶端綁定時使用。",
|
||||
"label": "預設播放聲音",
|
||||
"labelHint": "與「擷取」分頁的「以聲音播放」下拉選單共用——一個用於被動播放的預設聲音。",
|
||||
"none": "(無)"
|
||||
|
||||
@@ -250,7 +250,7 @@ def _migrate_mcp_bindings(engine, inspector, tables: set[str]) -> None:
|
||||
"""Drop the legacy ``default_intent`` column and add ``default_personality``.
|
||||
|
||||
The intent tri-state (respond / rewrite / compose) has been collapsed
|
||||
to a boolean: when true, ``voicebox.speak`` rewrites input through the
|
||||
to a boolean: when true, ``voicebox_speak`` rewrites input through the
|
||||
profile's personality LLM before TTS.
|
||||
"""
|
||||
if "mcp_client_bindings" not in tables:
|
||||
|
||||
@@ -272,7 +272,7 @@ class MCPClientBinding(Base):
|
||||
label = Column(String, nullable=True) # display name
|
||||
profile_id = Column(String, ForeignKey("profiles.id"), nullable=True)
|
||||
default_engine = Column(String, nullable=True)
|
||||
# When true, voicebox.speak routes through the profile's personality LLM
|
||||
# When true, voicebox_speak routes through the profile's personality LLM
|
||||
# (rewrite) before TTS by default. Callers can still override per call.
|
||||
default_personality = Column(Boolean, nullable=False, default=False)
|
||||
last_seen_at = Column(DateTime, nullable=True)
|
||||
|
||||
@@ -49,10 +49,10 @@ claude mcp add voicebox \
|
||||
|
||||
| Name | Purpose |
|
||||
|---|---|
|
||||
| `voicebox.speak` | Speak text in a voice profile. Returns a generation id you can poll. |
|
||||
| `voicebox.transcribe` | Whisper transcription of a base64 blob or an absolute local path. |
|
||||
| `voicebox.list_captures` | Recent captures (dictation / recording / file) with transcripts. |
|
||||
| `voicebox.list_profiles` | Available voice profiles (cloned + preset). |
|
||||
| `voicebox_speak` | Speak text in a voice profile. Returns a generation id you can poll. |
|
||||
| `voicebox_transcribe` | Whisper transcription of a base64 blob or an absolute local path. |
|
||||
| `voicebox_list_captures` | Recent captures (dictation / recording / file) with transcripts. |
|
||||
| `voicebox_list_profiles` | Available voice profiles (cloned + preset). |
|
||||
|
||||
All tools resolve voice profiles in this precedence:
|
||||
|
||||
@@ -69,8 +69,8 @@ Settings → MCP.
|
||||
npx @modelcontextprotocol/inspector http://127.0.0.1:17493/mcp
|
||||
```
|
||||
|
||||
Point it at the URL, hit "List tools," call `voicebox.list_profiles`
|
||||
first to confirm wiring, then `voicebox.speak` for end-to-end.
|
||||
Point it at the URL, hit "List tools," call `voicebox_list_profiles`
|
||||
first to confirm wiring, then `voicebox_speak` for end-to-end.
|
||||
|
||||
## Non-MCP REST surface
|
||||
|
||||
|
||||
@@ -33,7 +33,7 @@ current_client_id: ContextVar[str | None] = ContextVar(
|
||||
)
|
||||
|
||||
# Remote address of the in-flight request. Used by tools that gate
|
||||
# host-filesystem access to loopback callers (see voicebox.transcribe).
|
||||
# host-filesystem access to loopback callers (see voicebox_transcribe).
|
||||
current_remote_addr: ContextVar[str | None] = ContextVar(
|
||||
"current_remote_addr", default=None
|
||||
)
|
||||
@@ -61,12 +61,12 @@ def request_is_loopback() -> bool:
|
||||
# ignored so the Settings UI's "last heard from" column only reflects
|
||||
# calls that actually acted on the client's bindings.
|
||||
#
|
||||
# - /mcp — FastMCP tool calls (voicebox.speak, voicebox.transcribe, …)
|
||||
# - /mcp — FastMCP tool calls (voicebox_speak, voicebox_transcribe, …)
|
||||
# and the /mcp/bindings admin surface. The admin surface is never
|
||||
# called with the header in practice (the frontend manages bindings
|
||||
# over plain REST), so the `startswith("/mcp")` match doesn't cause
|
||||
# false stamps.
|
||||
# - /speak — REST mirror of voicebox.speak for non-MCP agents (shell
|
||||
# - /speak — REST mirror of voicebox_speak for non-MCP agents (shell
|
||||
# scripts, ACP, A2A). Uses the same per-client binding lookup, so its
|
||||
# callers belong in the last-seen list too.
|
||||
_STAMPED_PATH_PREFIXES: tuple[str, ...] = ("/mcp", "/speak")
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
"""In-memory pub/sub for speaking-pill SSE broadcasts.
|
||||
|
||||
MCP ``voicebox.speak`` calls and the REST ``POST /speak`` route publish
|
||||
MCP ``voicebox_speak`` calls and the REST ``POST /speak`` route publish
|
||||
start/end events that DictateWindow subscribes to via /events/speak, so the
|
||||
floating pill surfaces whenever an agent is speaking.
|
||||
"""
|
||||
|
||||
@@ -27,8 +27,8 @@ def build_mcp_server() -> FastMCP:
|
||||
mcp = FastMCP(
|
||||
name="voicebox",
|
||||
instructions=(
|
||||
"Voicebox is a local voice I/O layer. Use `voicebox.speak` to "
|
||||
"play text in a voice profile, `voicebox.transcribe` for "
|
||||
"Voicebox is a local voice I/O layer. Use `voicebox_speak` to "
|
||||
"play text in a voice profile, `voicebox_transcribe` for "
|
||||
"audio→text, and the `list_*` tools to discover profiles and "
|
||||
"captures."
|
||||
),
|
||||
|
||||
@@ -1,8 +1,9 @@
|
||||
"""Voicebox MCP tool implementations.
|
||||
|
||||
Thin wrappers over existing services/routes. Tools are registered with dotted
|
||||
names (``voicebox.speak`` etc.) so they look natural in agent logs —
|
||||
the Python function name stays snake_case.
|
||||
Thin wrappers over existing services/routes. Tools are registered with
|
||||
underscore-separated names (``voicebox_speak`` etc.): MCP clients such as
|
||||
Claude Desktop validate tool names against ``^[a-zA-Z0-9_-]{1,64}$`` and
|
||||
reject the whole tool list if any name contains a dot (#790).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -202,7 +203,7 @@ def register_tools(mcp: FastMCP) -> None:
|
||||
name="voicebox_list_profiles",
|
||||
description=(
|
||||
"List available voice profiles (both cloned voices and presets). "
|
||||
"Use the returned `name` with voicebox.speak(profile=...)."
|
||||
"Use the returned `name` with voicebox_speak(profile=...)."
|
||||
),
|
||||
)
|
||||
async def voicebox_list_profiles() -> dict[str, Any]:
|
||||
|
||||
+2
-2
@@ -309,7 +309,7 @@ class GenerationSettingsUpdate(BaseModel):
|
||||
|
||||
class MCPClientBindingResponse(BaseModel):
|
||||
"""Per-MCP-client voice binding — what voice / engine the server should
|
||||
use when a given client_id calls voicebox.speak without args, plus an
|
||||
use when a given client_id calls voicebox_speak without args, plus an
|
||||
opt-in personality-rewrite default."""
|
||||
|
||||
client_id: str
|
||||
@@ -346,7 +346,7 @@ class MCPClientBindingListResponse(BaseModel):
|
||||
|
||||
|
||||
class SpeakRequest(BaseModel):
|
||||
"""Body for POST /speak — non-MCP REST surface that mirrors voicebox.speak."""
|
||||
"""Body for POST /speak — non-MCP REST surface that mirrors voicebox_speak."""
|
||||
|
||||
text: str = Field(..., min_length=1, max_length=10000)
|
||||
profile: Optional[str] = Field(
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
"""POST /speak — REST wrapper around voicebox.speak for non-MCP callers.
|
||||
"""POST /speak — REST wrapper around voicebox_speak for non-MCP callers.
|
||||
|
||||
Shell scripts, ACP, A2A, or any agent that doesn't speak MCP can hit this
|
||||
endpoint to play text through a cloned voice. Uses the same profile
|
||||
@@ -30,7 +30,7 @@ async def speak(
|
||||
request: Request,
|
||||
db: Session = Depends(get_db),
|
||||
):
|
||||
"""Speak text in a voice profile. Mirrors voicebox.speak (MCP).
|
||||
"""Speak text in a voice profile. Mirrors voicebox_speak (MCP).
|
||||
|
||||
Response shape matches POST /generate — a ``GenerationResponse`` with
|
||||
``status="generating"`` and an ``id`` the caller polls at
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
"""Tests for the voicebox.speak MCP tool's ``model_size`` plumbing (issue #884).
|
||||
"""Tests for the voicebox_speak MCP tool's ``model_size`` plumbing (issue #884).
|
||||
|
||||
The MCP speak path used to build its ``GenerationRequest`` without a
|
||||
``model_size``, so every agent-triggered generation silently fell back to the
|
||||
|
||||
@@ -121,7 +121,7 @@ POST /generate
|
||||
Shipped 2026-04-25 (PR #544). Voicebox went from a voice-cloning studio to a full voice studio — dictation in, agent speech out, a local LLM in the middle.
|
||||
|
||||
- **Dictation** — global hotkey capture (push-to-talk + toggle chords), on-screen pill with live state, auto-paste into the focused field with clipboard save/restore, chord-picker UI. Scoped Accessibility permission (transcripts still land if paste is denied).
|
||||
- **MCP server** at `http://127.0.0.1:17493/mcp` — `voicebox.speak` / `.transcribe` / `.list_captures` / `.list_profiles`. Streamable HTTP primary transport, stdio sidecar shim, per-client voice binding via `X-Voicebox-Client-Id`. Speaking pill always shows agent-initiated output.
|
||||
- **MCP server** at `http://127.0.0.1:17493/mcp` — `voicebox_speak` / `voicebox_transcribe` / `voicebox_list_captures` / `voicebox_list_profiles`. Streamable HTTP primary transport, stdio sidecar shim, per-client voice binding via `X-Voicebox-Client-Id`. Speaking pill always shows agent-initiated output.
|
||||
- **Personality** — voice profiles carry an optional ≤2000-char persona. Compose (shuffle an in-character line) and Speak-in-character (rewrite input before TTS), both on a local Qwen3 LLM that doubles as the refinement model.
|
||||
- **Refinement** — on-device Qwen3 strips fillers, fixes punctuation, optional self-correction rewrites; Whisper hallucination-loop stripping at a 6-token threshold; per-capture flag snapshots; model picker (0.6B / 1.7B / 4B).
|
||||
- **`POST /speak` REST wrapper** and **i18next foundation** (English + zh-CN) also landed.
|
||||
|
||||
@@ -105,15 +105,15 @@ for the shim to connect.
|
||||
|
||||
| Tool | Use |
|
||||
|---|---|
|
||||
| `voicebox.speak` | Speak text in a voice profile. Returns a `generation_id` to poll. |
|
||||
| `voicebox.transcribe` | Whisper transcription of base64 audio or an absolute local path. |
|
||||
| `voicebox.list_captures` | Recent captures with transcripts, paginated. |
|
||||
| `voicebox.list_profiles` | Available voice profiles (cloned + preset). |
|
||||
| `voicebox_speak` | Speak text in a voice profile. Returns a `generation_id` to poll. |
|
||||
| `voicebox_transcribe` | Whisper transcription of base64 audio or an absolute local path. |
|
||||
| `voicebox_list_captures` | Recent captures with transcripts, paginated. |
|
||||
| `voicebox_list_profiles` | Available voice profiles (cloned + preset). |
|
||||
|
||||
### `voicebox.speak`
|
||||
### `voicebox_speak`
|
||||
|
||||
```ts
|
||||
voicebox.speak({
|
||||
voicebox_speak({
|
||||
text: "Deploy complete.",
|
||||
profile?: "Morgan", // name or id; falls back to per-client binding, then default
|
||||
engine?: "qwen", // qwen | qwen_custom_voice | luxtts | chatterbox | chatterbox_turbo | tada | kokoro
|
||||
@@ -138,10 +138,10 @@ Returns:
|
||||
- **Persona mode** — `personality: true` and the profile must have a personality prompt set.
|
||||
The LLM rewrites the text in character before TTS. See [Voice Personalities](/overview/voice-personalities).
|
||||
|
||||
### `voicebox.transcribe`
|
||||
### `voicebox_transcribe`
|
||||
|
||||
```ts
|
||||
voicebox.transcribe({
|
||||
voicebox_transcribe({
|
||||
audio_base64?: "<base64>", // exactly one of these two
|
||||
audio_path?: "/absolute/path/to/file.wav",
|
||||
language?: "en",
|
||||
@@ -151,18 +151,18 @@ voicebox.transcribe({
|
||||
|
||||
Returns `{ text, duration, language, model }`. 200 MB ceiling on either path.
|
||||
|
||||
### `voicebox.list_captures`
|
||||
### `voicebox_list_captures`
|
||||
|
||||
`{ limit?: 20, offset?: 0 }` → `{ captures: [...], total }`. `limit` is
|
||||
clamped to `1..=200`.
|
||||
|
||||
### `voicebox.list_profiles`
|
||||
### `voicebox_list_profiles`
|
||||
|
||||
No args → `{ profiles: [{ id, name, voice_type, language, has_personality }] }`.
|
||||
|
||||
## Voice resolution
|
||||
|
||||
Every call to `voicebox.speak` (and `POST /speak`) resolves the voice profile
|
||||
Every call to `voicebox_speak` (and `POST /speak`) resolves the voice profile
|
||||
in this order:
|
||||
|
||||
<Steps>
|
||||
@@ -195,7 +195,7 @@ carries:
|
||||
| `label` | Display name in the Settings UI (e.g. "Claude Code"). |
|
||||
| `profile_id` | The voice this client uses when `profile` isn't passed. |
|
||||
| `default_engine` | Override the TTS engine for this client. |
|
||||
| `default_personality` | When true, `voicebox.speak` routes through the profile's personality LLM (rewrite) by default. |
|
||||
| `default_personality` | When true, `voicebox_speak` routes through the profile's personality LLM (rewrite) by default. |
|
||||
| `last_seen_at` | Last time the server saw a request from this client. |
|
||||
|
||||
`last_seen_at` is stamped automatically by middleware on every `/mcp/*`
|
||||
@@ -239,8 +239,8 @@ agent:
|
||||
npx @modelcontextprotocol/inspector http://127.0.0.1:17493/mcp
|
||||
```
|
||||
|
||||
Start with `voicebox.list_profiles` to confirm wiring, then
|
||||
`voicebox.speak` for end-to-end — you should hear audio and see the
|
||||
Start with `voicebox_list_profiles` to confirm wiring, then
|
||||
`voicebox_speak` for end-to-end — you should hear audio and see the
|
||||
generation land in the Captures tab.
|
||||
|
||||
<Callout type="info">
|
||||
@@ -262,7 +262,7 @@ generation land in the Captures tab.
|
||||
boundary. If you're scripting against a shared host, prefer
|
||||
`audio_base64` so you don't have to think about path sandboxing.
|
||||
- **Voice cloning consent applies.** See [Voice Cloning](/overview/voice-cloning#limitations)
|
||||
— an agent being able to call `voicebox.speak` in someone's voice
|
||||
— an agent being able to call `voicebox_speak` in someone's voice
|
||||
doesn't change the ethics of whose voices you clone.
|
||||
|
||||
## Implementation notes
|
||||
|
||||
@@ -128,7 +128,7 @@ until you flip it back off.
|
||||
- **Agents that speak in a voice you own.** Combine the persona toggle with
|
||||
the built-in [MCP Server](/overview/mcp-server) so Claude Code, Cursor,
|
||||
Cline, or any MCP-aware agent can talk back through a profile with a
|
||||
personality. The agent calls `voicebox.speak({ text, profile, personality:
|
||||
personality. The agent calls `voicebox_speak({ text, profile, personality:
|
||||
true })` and Voicebox rewrites the text in character before speaking.
|
||||
- **Interactive characters.** Games, narrative tools, accessibility
|
||||
experiences. A character with a personality description plus a cloned
|
||||
@@ -151,7 +151,7 @@ Personalities are accessible via REST:
|
||||
| `POST` | `/generate` | Include `personality: true` to run input text through the personality LLM before TTS. Same for `POST /speak`. |
|
||||
|
||||
`POST /generate` with `personality: true` is the same primitive MCP's
|
||||
`voicebox.speak` tool uses when you pass `personality: true`. Scripts and
|
||||
`voicebox_speak` tool uses when you pass `personality: true`. Scripts and
|
||||
agents can use it directly.
|
||||
|
||||
## Limits and gotchas
|
||||
|
||||
@@ -23,7 +23,7 @@ const SCENARIOS: Scenario[] = [
|
||||
{ prefix: '$', text: 'claude run', tone: 'accent' },
|
||||
{ prefix: '✓', text: 'Tests passing (42 files)', tone: 'success' },
|
||||
{ prefix: '✓', text: 'Build succeeded in 12.4s', tone: 'success' },
|
||||
{ prefix: '→', text: 'voicebox.speak({ profile: "Morgan" })', tone: 'dim' },
|
||||
{ prefix: '→', text: 'voicebox_speak({ profile: "Morgan" })', tone: 'dim' },
|
||||
],
|
||||
utterance: 'Tests passing. Ready to merge.',
|
||||
},
|
||||
@@ -35,7 +35,7 @@ const SCENARIOS: Scenario[] = [
|
||||
{ prefix: '$', text: 'cursor agent:deploy', tone: 'accent' },
|
||||
{ prefix: '✓', text: 'Migration applied (4 tables)', tone: 'success' },
|
||||
{ prefix: '✓', text: 'Deploy complete', tone: 'success' },
|
||||
{ prefix: '→', text: 'voicebox.speak({ profile: "Scarlett" })', tone: 'dim' },
|
||||
{ prefix: '→', text: 'voicebox_speak({ profile: "Scarlett" })', tone: 'dim' },
|
||||
],
|
||||
utterance: 'Deploy shipped. Prod is green.',
|
||||
},
|
||||
@@ -46,7 +46,7 @@ const SCENARIOS: Scenario[] = [
|
||||
log: [
|
||||
{ prefix: '$', text: 'cline task:review', tone: 'accent' },
|
||||
{ prefix: '!', text: '3 files need attention', tone: 'dim' },
|
||||
{ prefix: '→', text: 'voicebox.speak({ profile: "Jarvis" })', tone: 'dim' },
|
||||
{ prefix: '→', text: 'voicebox_speak({ profile: "Jarvis" })', tone: 'dim' },
|
||||
],
|
||||
utterance: 'Review ready. Three files to look at.',
|
||||
},
|
||||
@@ -202,7 +202,7 @@ const MCP_CONFIG = `{
|
||||
}`;
|
||||
|
||||
const SPEAK_EXAMPLE = `// In any MCP-aware agent:
|
||||
await voicebox.speak({
|
||||
await voicebox_speak({
|
||||
text: "Deploy complete.",
|
||||
profile: "Morgan",
|
||||
})`;
|
||||
@@ -300,7 +300,7 @@ export function AgentIntegration() {
|
||||
</h2>
|
||||
<p className="text-muted-foreground text-base md:text-lg leading-relaxed">
|
||||
One tool call —{' '}
|
||||
<code className="text-accent font-mono text-[0.9em]">voicebox.speak</code> —
|
||||
<code className="text-accent font-mono text-[0.9em]">voicebox_speak</code> —
|
||||
and any MCP-aware agent can talk to you in a voice you’ve cloned. Claude Code,
|
||||
Cursor, Cline, or anything that speaks MCP.
|
||||
</p>
|
||||
|
||||
@@ -152,7 +152,7 @@ const CAPTURES: Capture[] = [
|
||||
transcriptRaw:
|
||||
"draft an update for the blog about the agent voice feature the key point is one MCP tool call and any agent on your machine gets a voice claude code finishes a long task calls voicebox dot speak and you hear it in a voice you've cloned morgan scarlett whatever you set up same pill that shows when you're dictating also shows when an agent is speaking so you always know what's coming out of your machine closes the whole voice IO loop for agents",
|
||||
transcriptRefined:
|
||||
"Draft an update for the blog about the agent voice feature. The key point: one MCP tool call, and any agent on your machine gets a voice. Claude Code finishes a long task, calls voicebox.speak, and you hear it in a voice you've cloned — Morgan, Scarlett, whatever you've set up. The same pill that shows when you're dictating also shows when an agent is speaking, so you always know what's coming out of your machine. It closes the full voice I/O loop for agents.",
|
||||
"Draft an update for the blog about the agent voice feature. The key point: one MCP tool call, and any agent on your machine gets a voice. Claude Code finishes a long task, calls voicebox_speak, and you hear it in a voice you've cloned — Morgan, Scarlett, whatever you've set up. The same pill that shows when you're dictating also shows when an agent is speaking, so you always know what's coming out of your machine. It closes the full voice I/O loop for agents.",
|
||||
durationMs: 41000,
|
||||
ago: '22 min ago',
|
||||
createdAtLabel: 'Apr 22, 3:29 PM',
|
||||
|
||||
Reference in New Issue
Block a user