Merge remote-tracking branch 'origin/main' into prep/pr-659

# Conflicts:
#	backend/database/models.py
This commit is contained in:
jamiepine
2026-10-04 00:02:23 +00:00
49 changed files with 811 additions and 123 deletions
+5
View File
@@ -7,6 +7,11 @@
## [Unreleased] ## [Unreleased]
### Fixed
- Voice generation requests that omit `engine` now honor the selected profile's
configured engine instead of silently defaulting to Qwen.
### Linux ### Linux
- **ROCm setup works on Linux AMD systems.** Docker ROCm builds now keep PyTorch - **ROCm setup works on Linux AMD systems.** Docker ROCm builds now keep PyTorch
+8 -6
View File
@@ -80,7 +80,7 @@ The two cloud incumbents sit on opposite halves of the voice I/O loop — Eleven
- **Unlimited length** — auto-chunking with crossfade for scripts, articles, and chapters - **Unlimited length** — auto-chunking with crossfade for scripts, articles, and chapters
- **Stories editor** — multi-track timeline for conversations, podcasts, and narratives - **Stories editor** — multi-track timeline for conversations, podcasts, and narratives
- **Voice input** — global dictation hotkey with push-to-talk and toggle modes, accessibility-verified auto-paste on macOS, in-app mic on every text field, Whisper-based STT - **Voice input** — global dictation hotkey with push-to-talk and toggle modes, accessibility-verified auto-paste on macOS, in-app mic on every text field, Whisper-based STT
- **Agent voice output** — one tool call (`voicebox.speak`) and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you've cloned - **Agent voice output** — one tool call (`voicebox_speak`) and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you've cloned
- **Voice personalities** — attach a free-form persona to any voice profile, then Compose, Rewrite, or Respond via a bundled local LLM — agents can invoke the same modes over MCP - **Voice personalities** — attach a free-form persona to any voice profile, then Compose, Rewrite, or Respond via a bundled local LLM — agents can invoke the same modes over MCP
- **API-first** — REST API plus a built-in MCP server for integrating voice I/O into your own apps and agents - **API-first** — REST API plus a built-in MCP server for integrating voice I/O into your own apps and agents
- **Native performance** — built with Tauri (Rust), not Electron - **Native performance** — built with Tauri (Rust), not Electron
@@ -130,7 +130,9 @@ literally as text.
With **Chatterbox Turbo** selected, type `/` in the text input to open the tag With **Chatterbox Turbo** selected, type `/` in the text input to open the tag
inserter and add expressive tags inline with speech: inserter and add expressive tags inline with speech:
`[laugh]` `[chuckle]` `[gasp]` `[cough]` `[sigh]` `[groan]` `[sniff]` `[shush]` `[clear throat]` Sounds: `[laugh]` `[chuckle]` `[gasp]` `[cough]` `[sigh]` `[groan]` `[sniff]` `[shush]` `[clear throat]`
Delivery: `[angry]` `[crying]` `[dramatic]` `[fear]` `[happy]` `[narration]` `[sarcastic]` `[surprised]` `[whispering]` `[advertisement]`
### Post-Processing Effects ### Post-Processing Effects
@@ -232,7 +234,7 @@ Every agent gets a voice. One tool call and any MCP-aware agent can speak to you
```ts ```ts
// In any MCP-aware agent: // In any MCP-aware agent:
await voicebox.speak({ await voicebox_speak({
text: "Deploy complete.", text: "Deploy complete.",
profile: "Morgan", profile: "Morgan",
}); });
@@ -252,7 +254,7 @@ Attach a free-form personality to any voice profile — who this voice is, how t
- **Compose** — a shuffle button that drops a fresh in-character line into the textarea; edit and speak, or click again for a different take - **Compose** — a shuffle button that drops a fresh in-character line into the textarea; edit and speak, or click again for a different take
- **Speak in character** — a toggle that routes your input text through the personality LLM to be rewritten in their voice before TTS - **Speak in character** — a toggle that routes your input text through the personality LLM to be rewritten in their voice before TTS
Agents can reach the same rewrite path over MCP by passing `personality: true` to `voicebox.speak`, turning the tool into a text-in → personality-LLM → TTS pipeline. The same LLM backs dictation's refinement step — one LLM in the app, one model cache, one GPU-memory footprint. Agents can reach the same rewrite path over MCP by passing `personality: true` to `voicebox_speak`, turning the tool into a text-in → personality-LLM → TTS pipeline. The same LLM backs dictation's refinement step — one LLM in the app, one model cache, one GPU-memory footprint.
**Local LLM options:** Qwen3 0.6B / 1.7B / 4B, sharing the TTS runtime (MLX on Apple Silicon, PyTorch elsewhere). **Local LLM options:** Qwen3 0.6B / 1.7B / 4B, sharing the TTS runtime (MLX on Apple Silicon, PyTorch elsewhere).
@@ -345,11 +347,11 @@ claude mcp add voicebox \
} }
``` ```
Four tools ship: `voicebox.speak`, `voicebox.transcribe`, `voicebox.list_captures`, `voicebox.list_profiles`. Per-client voice bindings are managed in **Voicebox → Settings → MCP**. See the [full MCP guide](docs/content/docs/overview/mcp-server.mdx) for tool signatures, resolution precedence, the speaking-pill contract, and security notes. Four tools ship: `voicebox_speak`, `voicebox_transcribe`, `voicebox_list_captures`, `voicebox_list_profiles`. Per-client voice bindings are managed in **Voicebox → Settings → MCP**. See the [full MCP guide](docs/content/docs/overview/mcp-server.mdx) for tool signatures, resolution precedence, the speaking-pill contract, and security notes.
```ts ```ts
// In any MCP-aware agent: // In any MCP-aware agent:
await voicebox.speak({ await voicebox_speak({
text: "Tests passing. Ready to merge.", text: "Tests passing. Ready to merge.",
profile: "Morgan", // optional — falls back to the per-client binding profile: "Morgan", // optional — falls back to the per-client binding
personality: true, // optional — rewrites text through the profile's personality LLM first personality: true, // optional — rewrites text through the profile's personality LLM first
@@ -222,7 +222,7 @@ export function DictationReadinessChecklist({
/> />
)} )}
{readiness.llm && ( {readiness.autoRefine && readiness.llm && (
<ChecklistRow <ChecklistRow
icon={<Cpu className="h-3.5 w-3.5" />} icon={<Cpu className="h-3.5 w-3.5" />}
title={t('captures.readiness.llm.label', { name: readiness.llm.display_name })} title={t('captures.readiness.llm.label', { name: readiness.llm.display_name })}
@@ -23,9 +23,20 @@ const PARALINGUISTIC_TAGS = [
{ tag: '[sniff]', label: 'sniff', emoji: '\u{1F443}' }, { tag: '[sniff]', label: 'sniff', emoji: '\u{1F443}' },
{ tag: '[shush]', label: 'shush', emoji: '\u{1F92B}' }, { tag: '[shush]', label: 'shush', emoji: '\u{1F92B}' },
{ tag: '[clear throat]', label: 'clear throat', emoji: '\u{1F64A}' }, { tag: '[clear throat]', label: 'clear throat', emoji: '\u{1F64A}' },
{ tag: '[angry]', label: 'angry', emoji: '\u{1F621}' },
{ tag: '[crying]', label: 'crying', emoji: '\u{1F622}' },
{ tag: '[dramatic]', label: 'dramatic', emoji: '\u{1F3AD}' },
{ tag: '[fear]', label: 'fear', emoji: '\u{1F628}' },
{ tag: '[happy]', label: 'happy', emoji: '\u{1F60A}' },
{ tag: '[narration]', label: 'narration', emoji: '\u{1F4D6}' },
{ tag: '[sarcastic]', label: 'sarcastic', emoji: '\u{1F644}' },
{ tag: '[surprised]', label: 'surprised', emoji: '\u{1F632}' },
{ tag: '[whispering]', label: 'whispering', emoji: '\u{1F62F}' },
{ tag: '[advertisement]', label: 'advertisement', emoji: '\u{1F4E2}' },
] as const; ] as const;
const TAG_REGEX = /\[(laugh|chuckle|gasp|cough|sigh|groan|sniff|shush|clear throat)\]/gi; const TAG_REGEX =
/\[(laugh|chuckle|gasp|cough|sigh|groan|sniff|shush|clear throat|angry|crying|dramatic|fear|happy|narration|sarcastic|surprised|whispering|advertisement)\]/gi;
// Data attribute used to identify badge spans in the DOM // Data attribute used to identify badge spans in the DOM
const BADGE_ATTR = 'data-ptag'; const BADGE_ATTR = 'data-ptag';
@@ -136,6 +147,7 @@ export const ParalinguisticInput = forwardRef<ParalinguisticInputRef, Paralingui
left: 0, left: 0,
}); });
const triggerRangeRef = useRef<Range | null>(null); const triggerRangeRef = useRef<Range | null>(null);
const menuListRef = useRef<HTMLDivElement | null>(null);
const lastSerializedRef = useRef<string>(''); const lastSerializedRef = useRef<string>('');
const isComposingRef = useRef(false); const isComposingRef = useRef(false);
@@ -144,6 +156,14 @@ export const ParalinguisticInput = forwardRef<ParalinguisticInputRef, Paralingui
element: editorRef.current, element: editorRef.current,
})); }));
// The list is taller than the menu's max-height, so keep the keyboard
// highlight visible as ArrowUp/ArrowDown move it.
useEffect(() => {
if (!showMenu) return;
const item = menuListRef.current?.children[menuIndex] as HTMLElement | undefined;
item?.scrollIntoView({ block: 'nearest' });
}, [showMenu, menuIndex]);
// Filtered tag list for the autocomplete menu // Filtered tag list for the autocomplete menu
const filteredTags = PARALINGUISTIC_TAGS.filter((t) => const filteredTags = PARALINGUISTIC_TAGS.filter((t) =>
t.label.toLowerCase().includes(menuFilter.toLowerCase()), t.label.toLowerCase().includes(menuFilter.toLowerCase()),
@@ -381,6 +401,7 @@ export const ParalinguisticInput = forwardRef<ParalinguisticInputRef, Paralingui
createPortal( createPortal(
<AnimatePresence> <AnimatePresence>
<motion.div <motion.div
ref={menuListRef}
initial={{ opacity: 0, y: 4 }} initial={{ opacity: 0, y: 4 }}
animate={{ opacity: 1, y: 0 }} animate={{ opacity: 1, y: 0 }}
exit={{ opacity: 0, y: 4 }} exit={{ opacity: 0, y: 4 }}
+4 -4
View File
@@ -276,19 +276,19 @@ export function MCPPage() {
<h3 className="text-sm font-semibold">{t('settings.mcp.sidebar.toolsTitle')}</h3> <h3 className="text-sm font-semibold">{t('settings.mcp.sidebar.toolsTitle')}</h3>
<ul className="text-sm text-muted-foreground space-y-1.5 leading-relaxed"> <ul className="text-sm text-muted-foreground space-y-1.5 leading-relaxed">
<li> <li>
<code className="text-accent">voicebox.speak</code> <code className="text-accent">voicebox_speak</code>
<div>{t('settings.mcp.sidebar.tools.speak')}</div> <div>{t('settings.mcp.sidebar.tools.speak')}</div>
</li> </li>
<li> <li>
<code className="text-accent">voicebox.transcribe</code> <code className="text-accent">voicebox_transcribe</code>
<div>{t('settings.mcp.sidebar.tools.transcribe')}</div> <div>{t('settings.mcp.sidebar.tools.transcribe')}</div>
</li> </li>
<li> <li>
<code className="text-accent">voicebox.list_captures</code> <code className="text-accent">voicebox_list_captures</code>
<div>{t('settings.mcp.sidebar.tools.listCaptures')}</div> <div>{t('settings.mcp.sidebar.tools.listCaptures')}</div>
</li> </li>
<li> <li>
<code className="text-accent">voicebox.list_profiles</code> <code className="text-accent">voicebox_list_profiles</code>
<div>{t('settings.mcp.sidebar.tools.listProfiles')}</div> <div>{t('settings.mcp.sidebar.tools.listProfiles')}</div>
</li> </li>
</ul> </ul>
+12 -2
View File
@@ -22,7 +22,9 @@ const ToastViewport = React.forwardRef<
ToastViewport.displayName = ToastPrimitives.Viewport.displayName; ToastViewport.displayName = ToastPrimitives.Viewport.displayName;
const toastVariants = cva( const toastVariants = cva(
'group pointer-events-auto relative flex w-full items-center justify-between space-x-4 overflow-hidden rounded-md border p-6 pr-8 shadow-lg transition-all data-[swipe=cancel]:translate-x-0 data-[swipe=end]:translate-x-[var(--radix-toast-swipe-end-x)] data-[swipe=move]:translate-x-[var(--radix-toast-swipe-move-x)] data-[swipe=move]:transition-none data-[state=open]:animate-in data-[state=closed]:animate-out data-[swipe=end]:animate-out data-[state=closed]:fade-out-80 data-[state=closed]:slide-out-to-right-full data-[state=open]:slide-in-from-top-full data-[state=open]:sm:slide-in-from-bottom-full', // items-start rather than items-center: a long description scrolls inside a
// capped box, and centring it would push the title out of view.
'group pointer-events-auto relative flex w-full items-start justify-between space-x-4 overflow-hidden rounded-md border p-6 pr-8 shadow-lg transition-all data-[swipe=cancel]:translate-x-0 data-[swipe=end]:translate-x-[var(--radix-toast-swipe-end-x)] data-[swipe=move]:translate-x-[var(--radix-toast-swipe-move-x)] data-[swipe=move]:transition-none data-[state=open]:animate-in data-[state=closed]:animate-out data-[swipe=end]:animate-out data-[state=closed]:fade-out-80 data-[state=closed]:slide-out-to-right-full data-[state=open]:slide-in-from-top-full data-[state=open]:sm:slide-in-from-bottom-full',
{ {
variants: { variants: {
variant: { variant: {
@@ -98,7 +100,15 @@ const ToastDescription = React.forwardRef<
>(({ className, ...props }, ref) => ( >(({ className, ...props }, ref) => (
<ToastPrimitives.Description <ToastPrimitives.Description
ref={ref} ref={ref}
className={cn('text-sm opacity-90', className)} // A server traceback can run to thousands of characters. Unbounded, it
// overflowed the viewport and got clipped at both ends with the close
// button pushed off-screen, so the text was unreadable and the toast
// undismissable. Cap it and let it scroll; break-words keeps a long
// unspaced token (a path, a URL) from widening the toast.
className={cn(
'max-h-[40vh] overflow-y-auto overscroll-contain whitespace-pre-wrap break-words pr-1 text-sm opacity-90',
className,
)}
{...props} {...props}
/> />
)); ));
+1 -1
View File
@@ -1055,7 +1055,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "Default voice", "title": "Default voice",
"description": "Used when an agent calls voicebox.speak without a specific profile and has no per-client binding.", "description": "Used when an agent calls voicebox_speak without a specific profile and has no per-client binding.",
"label": "Default playback voice", "label": "Default playback voice",
"labelHint": "Shared with the Captures-tab 'Play as voice' dropdown — one default voice for passive playback.", "labelHint": "Shared with the Captures-tab 'Play as voice' dropdown — one default voice for passive playback.",
"none": "(none)" "none": "(none)"
+1 -1
View File
@@ -1055,7 +1055,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "Voz predeterminada", "title": "Voz predeterminada",
"description": "Se usa cuando un agente llama a voicebox.speak sin un perfil específico y no tiene una vinculación por cliente.", "description": "Se usa cuando un agente llama a voicebox_speak sin un perfil específico y no tiene una vinculación por cliente.",
"label": "Voz de reproducción predeterminada", "label": "Voz de reproducción predeterminada",
"labelHint": "Compartida con el desplegable 'Reproducir como voz' de la pestaña Capturas: una voz predeterminada para la reproducción pasiva.", "labelHint": "Compartida con el desplegable 'Reproducir como voz' de la pestaña Capturas: una voz predeterminada para la reproducción pasiva.",
"none": "(ninguna)" "none": "(ninguna)"
+1 -1
View File
@@ -1050,7 +1050,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "Voix par défaut", "title": "Voix par défaut",
"description": "Utilisée quand un agent appelle voicebox.speak sans profil spécifique et sans liaison par client.", "description": "Utilisée quand un agent appelle voicebox_speak sans profil spécifique et sans liaison par client.",
"label": "Voix de lecture par défaut", "label": "Voix de lecture par défaut",
"labelHint": "Partagé avec la liste déroulante « Lire avec » de l'onglet Captures — une voix par défaut pour la lecture passive.", "labelHint": "Partagé avec la liste déroulante « Lire avec » de l'onglet Captures — une voix par défaut pour la lecture passive.",
"none": "(aucune)" "none": "(aucune)"
+1 -1
View File
@@ -1055,7 +1055,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "Voce predefinita", "title": "Voce predefinita",
"description": "Utilizzata quando un agente chiama voicebox.speak senza specificare un profilo e non ha un'associazione per singolo client.", "description": "Utilizzata quando un agente chiama voicebox_speak senza specificare un profilo e non ha un'associazione per singolo client.",
"label": "Voce di riproduzione predefinita", "label": "Voce di riproduzione predefinita",
"labelHint": "Condivisa con il menu a discesa 'Riproduci come voce' della scheda Acquisizioni — una sola voce predefinita per la riproduzione passiva.", "labelHint": "Condivisa con il menu a discesa 'Riproduci come voce' della scheda Acquisizioni — una sola voce predefinita per la riproduzione passiva.",
"none": "(nessuna)" "none": "(nessuna)"
+1 -1
View File
@@ -1050,7 +1050,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "デフォルトボイス", "title": "デフォルトボイス",
"description": "エージェントが特定のプロファイルを指定せず、クライアントごとのバインディングもない状態で voicebox.speak を呼び出したときに使われます。", "description": "エージェントが特定のプロファイルを指定せず、クライアントごとのバインディングもない状態で voicebox_speak を呼び出したときに使われます。",
"label": "デフォルトの再生ボイス", "label": "デフォルトの再生ボイス",
"labelHint": "キャプチャタブの「ボイスで再生」ドロップダウンと共有 — パッシブ再生用に 1 つのデフォルトボイスを設定します。", "labelHint": "キャプチャタブの「ボイスで再生」ドロップダウンと共有 — パッシブ再生用に 1 つのデフォルトボイスを設定します。",
"none": "(なし)" "none": "(なし)"
+1 -1
View File
@@ -1050,7 +1050,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "기본 음성", "title": "기본 음성",
"description": "에이전트가 특정 프로필 없이 voicebox.speak를 호출하고 클라이언트별 바인딩도 없을 때 사용됩니다.", "description": "에이전트가 특정 프로필 없이 voicebox_speak를 호출하고 클라이언트별 바인딩도 없을 때 사용됩니다.",
"label": "기본 재생 음성", "label": "기본 재생 음성",
"labelHint": "캡처 탭의 'Play as 음성' 드롭다운과 공유 — 수동 재생용 기본 음성입니다.", "labelHint": "캡처 탭의 'Play as 음성' 드롭다운과 공유 — 수동 재생용 기본 음성입니다.",
"none": "(없음)" "none": "(없음)"
+1 -1
View File
@@ -1050,7 +1050,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "Voz padrão", "title": "Voz padrão",
"description": "Usada quando um agente chama voicebox.speak sem um perfil específico e não tem vínculo por cliente.", "description": "Usada quando um agente chama voicebox_speak sem um perfil específico e não tem vínculo por cliente.",
"label": "Voz de reprodução padrão", "label": "Voz de reprodução padrão",
"labelHint": "Compartilhada com o menu 'Reproduzir como voz' da aba Capturas — uma voz padrão para reprodução passiva.", "labelHint": "Compartilhada com o menu 'Reproduzir como voz' da aba Capturas — uma voz padrão para reprodução passiva.",
"none": "(nenhuma)" "none": "(nenhuma)"
+1 -1
View File
@@ -1050,7 +1050,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "默认声音", "title": "默认声音",
"description": "当代理调用 voicebox.speak 但未指定具体档案、且没有按客户端绑定时使用。", "description": "当代理调用 voicebox_speak 但未指定具体档案、且没有按客户端绑定时使用。",
"label": "默认播放声音", "label": "默认播放声音",
"labelHint": "与「捕获」标签页的「播放为」下拉菜单共享——被动播放的统一默认声音。", "labelHint": "与「捕获」标签页的「播放为」下拉菜单共享——被动播放的统一默认声音。",
"none": "(无)" "none": "(无)"
+1 -1
View File
@@ -1050,7 +1050,7 @@
}, },
"defaultVoice": { "defaultVoice": {
"title": "預設聲音", "title": "預設聲音",
"description": "當代理呼叫 voicebox.speak 卻未指定聲音檔案,且沒有對應客戶端綁定時使用。", "description": "當代理呼叫 voicebox_speak 卻未指定聲音檔案,且沒有對應客戶端綁定時使用。",
"label": "預設播放聲音", "label": "預設播放聲音",
"labelHint": "與「擷取」分頁的「以聲音播放」下拉選單共用——一個用於被動播放的預設聲音。", "labelHint": "與「擷取」分頁的「以聲音播放」下拉選單共用——一個用於被動播放的預設聲音。",
"none": "(無)" "none": "(無)"
+16 -8
View File
@@ -3,6 +3,7 @@ import { useAccessibilityPermission } from '@/components/AccessibilityGate/Acces
import { useInputMonitoringPermission } from '@/components/InputMonitoringGate/InputMonitoringGate'; import { useInputMonitoringPermission } from '@/components/InputMonitoringGate/InputMonitoringGate';
import { apiClient } from '@/lib/api/client'; import { apiClient } from '@/lib/api/client';
import type { ModelReadiness } from '@/lib/api/types'; import type { ModelReadiness } from '@/lib/api/types';
import { useCaptureSettings } from '@/lib/hooks/useSettings';
import { usePlatform } from '@/platform/PlatformContext'; import { usePlatform } from '@/platform/PlatformContext';
const READINESS_POLL_INTERVAL_MS = 5_000; const READINESS_POLL_INTERVAL_MS = 5_000;
@@ -17,6 +18,8 @@ export interface DictationReadiness {
missing: ReadinessGate[]; missing: ReadinessGate[];
stt: ModelReadiness | undefined; stt: ModelReadiness | undefined;
llm: ModelReadiness | undefined; llm: ModelReadiness | undefined;
/** Whether the user has auto-refine (LLM polish) enabled. */
autoRefine: boolean;
inputMonitoring: boolean; inputMonitoring: boolean;
accessibility: boolean; accessibility: boolean;
refetch: () => void; refetch: () => void;
@@ -46,6 +49,8 @@ export interface DictationReadiness {
export function useDictationReadiness(): DictationReadiness { export function useDictationReadiness(): DictationReadiness {
const platform = usePlatform(); const platform = usePlatform();
const isTauri = platform.metadata.isTauri; const isTauri = platform.metadata.isTauri;
const { settings } = useCaptureSettings();
const autoRefine = settings?.auto_refine ?? true;
const { const {
needsPermission: inputMonNeeds, needsPermission: inputMonNeeds,
@@ -61,17 +66,19 @@ export function useDictationReadiness(): DictationReadiness {
const { data, isLoading, refetch } = useQuery({ const { data, isLoading, refetch } = useQuery({
queryKey: ['capture-readiness'], queryKey: ['capture-readiness'],
queryFn: () => apiClient.getCaptureReadiness(), queryFn: () => apiClient.getCaptureReadiness(),
// Poll only while a model is still missing/downloading. Once both are // Poll only while a required model is still missing/downloading. Once
// green the endpoint's answer can't change until the user swaps models // all required models are green the endpoint's answer can't change
// in settings, and that path invalidates the query explicitly from // until the user swaps models in settings, and that path invalidates
// useSettings. refetchOnWindowFocus stays gated to the same condition. // the query explicitly from useSettings. refetchOnWindowFocus stays
// gated to the same condition.
refetchInterval: (query) => { refetchInterval: (query) => {
const d = query.state.data; const d = query.state.data;
return d && d.stt.ready && d.llm.ready ? false : READINESS_POLL_INTERVAL_MS; const allGreen = d && d.stt.ready && (!autoRefine || d.llm.ready);
return allGreen ? false : READINESS_POLL_INTERVAL_MS;
}, },
refetchOnWindowFocus: (query) => { refetchOnWindowFocus: (query) => {
const d = query.state.data; const d = query.state.data;
return !(d && d.stt.ready && d.llm.ready); return !(d && d.stt.ready && (!autoRefine || d.llm.ready));
}, },
}); });
@@ -84,10 +91,10 @@ export function useDictationReadiness(): DictationReadiness {
const missing: ReadinessGate[] = []; const missing: ReadinessGate[] = [];
if (!sttReady) missing.push('stt'); if (!sttReady) missing.push('stt');
if (!llmReady) missing.push('llm'); if (autoRefine && !llmReady) missing.push('llm');
if (!inputMonitoring) missing.push('input_monitoring'); if (!inputMonitoring) missing.push('input_monitoring');
if (!accessibility) missing.push('accessibility'); if (!accessibility) missing.push('accessibility');
const canRecord = sttReady && llmReady && inputMonitoring; const canRecord = sttReady && (!autoRefine || llmReady) && inputMonitoring;
return { return {
isLoading, isLoading,
@@ -96,6 +103,7 @@ export function useDictationReadiness(): DictationReadiness {
missing, missing,
stt: data?.stt, stt: data?.stt,
llm: data?.llm, llm: data?.llm,
autoRefine,
inputMonitoring, inputMonitoring,
accessibility, accessibility,
refetch: () => { refetch: () => {
@@ -1,8 +1,10 @@
import { useQueryClient } from '@tanstack/react-query'; import { useQueryClient } from '@tanstack/react-query';
import { useEffect, useRef } from 'react'; import { useEffect, useRef } from 'react';
import { ToastAction } from '@/components/ui/toast';
import { useToast } from '@/components/ui/use-toast'; import { useToast } from '@/components/ui/use-toast';
import { apiClient } from '@/lib/api/client'; import { apiClient } from '@/lib/api/client';
import { useGenerationSettings } from '@/lib/hooks/useSettings'; import { useGenerationSettings } from '@/lib/hooks/useSettings';
import { condenseError } from '@/lib/utils/errorText';
import { useGenerationStore } from '@/stores/generationStore'; import { useGenerationStore } from '@/stores/generationStore';
import { usePlayerStore } from '@/stores/playerStore'; import { usePlayerStore } from '@/stores/playerStore';
@@ -131,10 +133,41 @@ export function useGenerationProgress() {
queryClient.refetchQueries({ queryKey: ['history'] }); queryClient.refetchQueries({ queryKey: ['history'] });
const condensed = condenseError(data.error || 'An error occurred during generation');
toast({ toast({
title: data.status === 'not_found' ? 'Generation not found' : 'Generation failed', title: data.status === 'not_found' ? 'Generation not found' : 'Generation failed',
description: data.error || 'An error occurred during generation', description: condensed.truncated
? `${condensed.display}\n\n(${condensed.omitted} more characters — copy for the full error, or see Settings → Logs)`
: condensed.display,
variant: 'destructive', variant: 'destructive',
// Only offered when there is more to read than what is shown, so
// the common short error keeps a plain toast.
action: condensed.truncated ? (
<ToastAction
altText="Copy the full error text"
onClick={() => {
// Two ways this fails: the Clipboard API is absent outside
// a secure context (property access throws), or writeText
// rejects because permission was denied. The try/catch
// covers both, so the click never becomes an unhandled
// rejection and the user is told nothing was copied.
void (async () => {
try {
await navigator.clipboard.writeText(condensed.full);
} catch {
toast({
title: 'Could not copy',
description:
'Clipboard unavailable. The full error is in Settings → Logs.',
variant: 'destructive',
});
}
})();
}}
>
Copy
</ToastAction>
) : undefined,
}); });
} }
} catch { } catch {
+5 -1
View File
@@ -48,7 +48,11 @@ export function useCaptureSettings() {
// call, but its cached response keeps serving the previous // call, but its cached response keeps serving the previous
// model's state until the next 5 s poll. Invalidate on model // model's state until the next 5 s poll. Invalidate on model
// swaps so the readiness checklist re-checks immediately. // swaps so the readiness checklist re-checks immediately.
if (patch.stt_model !== undefined || patch.llm_model !== undefined) { if (
patch.stt_model !== undefined ||
patch.llm_model !== undefined ||
patch.auto_refine !== undefined
) {
queryClient.invalidateQueries({ queryKey: ['capture-readiness'] }); queryClient.invalidateQueries({ queryKey: ['capture-readiness'] });
} }
}, },
+97
View File
@@ -0,0 +1,97 @@
/**
* Condense a server error for display in a toast.
*
* Some backend errors are enormous and mostly noise. The transformers
* "Unrecognized model" error is ~4.8KB, of which the first sentence carries all
* the meaning and the remaining 4.7KB is an alphabetical list of every model
* architecture it knows about. Rendering that in a 420px toast clipped the text
* at both ends and pushed the close button off-screen.
*
* The rule is deliberately generic rather than pattern-matching any one
* library: keep the head, cut at the most natural boundary available inside the
* budget, and report how much was dropped so nobody assumes they read it all.
*/
/** Longest head of an oversized error kept for the toast, before the ellipsis.
*
* Roughly the first paragraph — enough for a sentence or two of real message.
* This bounds the *message* rather than the rendered string: `display` is this
* plus `TRUNCATION_SUFFIX` when something was cut. Spending two of these
* characters on the ellipsis instead would shorten the message to no purpose,
* since nothing downstream has a hard character limit — the description box
* scrolls.
*/
const HEAD_BUDGET = 400;
/** Marks a `display` value as incomplete. Appended after the head. */
const TRUNCATION_SUFFIX = ' …';
/** Below this, condensing is not worth it and the whole message is shown.
*
* Deliberately above the budget rather than equal to it. An error of 450
* characters would otherwise be cut to 400 to save 50 — a worse result than
* showing all of it, since the description scrolls anyway. The gap also gives
* the threshold hysteresis instead of flipping between full and truncated
* around a single character. No Copy action is offered in this range because
* nothing is being withheld: `display` already holds the entire message.
*/
const MIN_TO_CONDENSE = HEAD_BUDGET + 120;
export interface CondensedError {
/** What to show in the toast.
*
* Byte-identical to `full` whenever nothing is omitted, so a message that
* fits is never altered. Only an oversized one is rewritten, into a trimmed
* head followed by an ellipsis.
*/
display: string;
/** The original string exactly as received, for copying. Never modified. */
full: string;
/** Whether `display` omits part of the message. */
truncated: boolean;
/** How many characters `display` leaves out. */
omitted: number;
}
export function condenseError(raw: string | null | undefined): CondensedError {
// `full` is what the Copy action hands over, so it stays byte-for-byte what
// the server sent. The measuring and cutting below works on a trimmed copy
// instead — surrounding blank space should not count toward the budget or
// the omitted count — but `text` is never what gets returned as `display`
// unless the message is actually being shortened.
const full = raw ?? '';
const text = full.trim();
if (text.length <= MIN_TO_CONDENSE) {
return { display: full, full, truncated: false, omitted: 0 };
}
// A traceback's first line is nearly always the message; prefer it whenever
// it fits, since a newline is a stronger boundary than any punctuation.
const firstLine = text.split('\n', 1)[0].trim();
const useFirstLine = firstLine.length > 0 && firstLine.length <= HEAD_BUDGET;
let head = useFirstLine ? firstLine : text.slice(0, HEAD_BUDGET);
if (!useFirstLine) {
// A raw slice can stop mid-word, so back off to the last sentence end
// inside the budget. A first line that fit already ends at a newline, the
// stronger boundary, and is kept whole. Only accept the backoff if it
// keeps most of the budget — otherwise a stray early period would throw
// away usable context.
const lastStop = Math.max(head.lastIndexOf('. '), head.lastIndexOf('? '));
if (lastStop > HEAD_BUDGET * 0.4) {
head = head.slice(0, lastStop + 1);
}
}
head = head.trimEnd();
const omitted = text.length - head.length;
// Guard against the boundary search having produced nothing shorter. Nothing
// is omitted here either, so the message goes back exactly as it arrived.
if (omitted <= 0) {
return { display: full, full, truncated: false, omitted: 0 };
}
return { display: `${head}${TRUNCATION_SUFFIX}`, full, truncated: true, omitted };
}
+38 -4
View File
@@ -6,6 +6,7 @@ voice prompt combination, and model loading progress tracking.
""" """
import logging import logging
import os
import platform import platform
from contextlib import contextmanager from contextlib import contextmanager
from pathlib import Path from pathlib import Path
@@ -21,6 +22,24 @@ from ..utils.tasks import get_task_manager
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
def has_in_progress_download(blobs_dir: Path) -> bool:
"""
Whether a HuggingFace repo's ``blobs`` dir holds a genuinely in-progress download.
An ``.incomplete`` blob means a download is still in progress -- unless a
completed blob with the same hash already sits next to it, which happens
when a retried/concurrent download leaves a stale ``.incomplete`` behind
after the real transfer already finished. Only orphaned ``.incomplete``
files (no matching completed blob) count as "in progress".
"""
if not blobs_dir.exists():
return False
return any(
not incomplete.with_name(incomplete.name.removesuffix(".incomplete")).exists()
for incomplete in blobs_dir.glob("*.incomplete")
)
def is_model_cached( def is_model_cached(
hf_repo: str, hf_repo: str,
*, *,
@@ -47,10 +66,8 @@ def is_model_cached(
if not repo_cache.exists(): if not repo_cache.exists():
return False return False
# Incomplete blobs mean a download is still in progress if has_in_progress_download(repo_cache / "blobs"):
blobs_dir = repo_cache / "blobs" logger.debug(f"Found in-progress .incomplete file for {hf_repo}")
if blobs_dir.exists() and any(blobs_dir.glob("*.incomplete")):
logger.debug(f"Found .incomplete files for {hf_repo}")
return False return False
snapshots_dir = repo_cache / "snapshots" snapshots_dir = repo_cache / "snapshots"
@@ -77,6 +94,13 @@ def is_model_cached(
return False return False
# Documented escape hatch (docs/content/docs/overview/gpu-acceleration.mdx):
# users whose GPU has no compiled kernels in the bundled PyTorch set this to run
# on CPU instead of crashing at generation time.
FORCE_CPU_ENV_VAR = "VOICEBOX_FORCE_CPU"
FORCE_CPU_ENABLED_VALUE = "1"
def get_torch_device( def get_torch_device(
*, *,
allow_xpu: bool = False, allow_xpu: bool = False,
@@ -92,7 +116,17 @@ def get_torch_device(
allow_directml: Check for DirectML (Windows) support. allow_directml: Check for DirectML (Windows) support.
allow_mps: Allow MPS (Apple Silicon). If False, MPS falls back to CPU. allow_mps: Allow MPS (Apple Silicon). If False, MPS falls back to CPU.
force_cpu_on_mac: Force CPU on macOS regardless of GPU availability. force_cpu_on_mac: Force CPU on macOS regardless of GPU availability.
The VOICEBOX_FORCE_CPU override wins over every other candidate, and is
resolved before torch is imported so it still works when the installed
build is the reason CPU is wanted.
""" """
# Stripped: on Windows, where this override matters most, it is usually set
# through the GUI environment editor.
if os.environ.get(FORCE_CPU_ENV_VAR, "").strip() == FORCE_CPU_ENABLED_VALUE:
logger.info("%s=%s set, forcing CPU device", FORCE_CPU_ENV_VAR, FORCE_CPU_ENABLED_VALUE)
return "cpu"
if force_cpu_on_mac and platform.system() == "Darwin": if force_cpu_on_mac and platform.system() == "Darwin":
return "cpu" return "cpu"
+51 -1
View File
@@ -44,6 +44,7 @@ def run_migrations(engine) -> None:
_migrate_capture_settings(engine, inspector, tables) _migrate_capture_settings(engine, inspector, tables)
_migrate_mcp_bindings(engine, inspector, tables) _migrate_mcp_bindings(engine, inspector, tables)
_normalize_storage_paths(engine, tables) _normalize_storage_paths(engine, tables)
_migrate_add_indexes(engine, tables)
# -- helpers --------------------------------------------------------------- # -- helpers ---------------------------------------------------------------
@@ -249,7 +250,7 @@ def _migrate_mcp_bindings(engine, inspector, tables: set[str]) -> None:
"""Drop the legacy ``default_intent`` column and add ``default_personality``. """Drop the legacy ``default_intent`` column and add ``default_personality``.
The intent tri-state (respond / rewrite / compose) has been collapsed The intent tri-state (respond / rewrite / compose) has been collapsed
to a boolean: when true, ``voicebox.speak`` rewrites input through the to a boolean: when true, ``voicebox_speak`` rewrites input through the
profile's personality LLM before TTS. profile's personality LLM before TTS.
""" """
if "mcp_client_bindings" not in tables: if "mcp_client_bindings" not in tables:
@@ -292,6 +293,55 @@ def _supports_drop_column(engine) -> bool:
return tuple(int(p) for p in sqlite3.sqlite_version.split(".")[:3]) >= (3, 35, 0) return tuple(int(p) for p in sqlite3.sqlite_version.split(".")[:3]) >= (3, 35, 0)
def _migrate_add_indexes(engine, tables: set[str]) -> None:
"""Create missing indexes on high-traffic foreign keys and sort columns.
SQLite silently ignores ``CREATE INDEX IF NOT EXISTS``, so this is
safe to run on every startup regardless of whether the index already
exists. New installs get the indexes from ``Base.metadata.create_all``
(via the ``index=True`` column flags); this migration brings existing
databases into parity without dropping or recreating any data.
"""
indexes = [
# generations — filtered by profile, ordered/filtered by date, filtered by status
("ix_generations_profile_id", "generations", "profile_id"),
("ix_generations_created_at", "generations", "created_at"),
("ix_generations_status", "generations", "status"),
# story_items — every story lookup filters by story_id; join on generation_id
("ix_story_items_story_id", "story_items", "story_id"),
("ix_story_items_generation_id", "story_items", "generation_id"),
# generation_versions — always filtered/joined on generation_id
("ix_generation_versions_generation_id", "generation_versions", "generation_id"),
# profile_samples — loaded per-profile on every voice prompt build
("ix_profile_samples_profile_id", "profile_samples", "profile_id"),
# captures — ordered by date in list view
("ix_captures_created_at", "captures", "created_at"),
# channel_device_mappings — looked up per channel
("ix_channel_device_mappings_channel_id", "channel_device_mappings", "channel_id"),
]
with engine.connect() as conn:
existing = {
row[0]
for row in conn.execute(text("SELECT name FROM sqlite_master WHERE type = 'index'"))
}
created = []
for index_name, table, column in indexes:
if table not in tables or index_name in existing:
continue
conn.execute(
text(
f"CREATE INDEX IF NOT EXISTS {index_name}"
f" ON {table} ({column})"
)
)
created.append(index_name)
conn.commit()
if created:
logger.info("Created %d missing index(es): %s", len(created), ", ".join(created))
def _normalize_storage_paths(engine, tables: set[str]) -> None: def _normalize_storage_paths(engine, tables: set[str]) -> None:
"""Normalize stored file paths to be relative to the configured data dir.""" """Normalize stored file paths to be relative to the configured data dir."""
from pathlib import Path from pathlib import Path
+10 -10
View File
@@ -54,7 +54,7 @@ class ProfileSample(Base):
__tablename__ = "profile_samples" __tablename__ = "profile_samples"
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4())) id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
profile_id = Column(String, ForeignKey("profiles.id"), nullable=False) profile_id = Column(String, ForeignKey("profiles.id"), nullable=False, index=True)
audio_path = Column(String, nullable=False) audio_path = Column(String, nullable=False)
reference_text = Column(Text, nullable=False) reference_text = Column(Text, nullable=False)
@@ -65,7 +65,7 @@ class Generation(Base):
__tablename__ = "generations" __tablename__ = "generations"
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4())) id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
profile_id = Column(String, ForeignKey("profiles.id"), nullable=False) profile_id = Column(String, ForeignKey("profiles.id"), nullable=False, index=True)
text = Column(Text, nullable=False) text = Column(Text, nullable=False)
language = Column(String, default="en") language = Column(String, default="en")
audio_path = Column(String, nullable=True) audio_path = Column(String, nullable=True)
@@ -74,7 +74,7 @@ class Generation(Base):
instruct = Column(Text) instruct = Column(Text)
engine = Column(String, default="qwen") engine = Column(String, default="qwen")
model_size = Column(String, nullable=True) model_size = Column(String, nullable=True)
status = Column(String, default="completed") status = Column(String, default="completed", index=True)
error = Column(Text, nullable=True) error = Column(Text, nullable=True)
is_favorited = Column(Boolean, default=False) is_favorited = Column(Boolean, default=False)
# Origin of this generation — "manual" for plain /generate calls, # Origin of this generation — "manual" for plain /generate calls,
@@ -82,7 +82,7 @@ class Generation(Base):
# profile's personality LLM before TTS. Future sources (bulk import, # profile's personality LLM before TTS. Future sources (bulk import,
# agent replies, etc.) can extend this. # agent replies, etc.) can extend this.
source = Column(String, nullable=False, default="manual") source = Column(String, nullable=False, default="manual")
created_at = Column(DateTime, default=lambda: datetime.now(UTC)) created_at = Column(DateTime, default=lambda: datetime.now(UTC), index=True)
class Story(Base): class Story(Base):
@@ -103,8 +103,8 @@ class StoryItem(Base):
__tablename__ = "story_items" __tablename__ = "story_items"
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4())) id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
story_id = Column(String, ForeignKey("stories.id"), nullable=False) story_id = Column(String, ForeignKey("stories.id"), nullable=False, index=True)
generation_id = Column(String, ForeignKey("generations.id"), nullable=False) generation_id = Column(String, ForeignKey("generations.id"), nullable=False, index=True)
version_id = Column(String, ForeignKey("generation_versions.id"), nullable=True) version_id = Column(String, ForeignKey("generation_versions.id"), nullable=True)
start_time_ms = Column(Integer, nullable=False, default=0) start_time_ms = Column(Integer, nullable=False, default=0)
track = Column(Integer, nullable=False, default=0) track = Column(Integer, nullable=False, default=0)
@@ -132,7 +132,7 @@ class GenerationVersion(Base):
__tablename__ = "generation_versions" __tablename__ = "generation_versions"
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4())) id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
generation_id = Column(String, ForeignKey("generations.id"), nullable=False) generation_id = Column(String, ForeignKey("generations.id"), nullable=False, index=True)
label = Column(String, nullable=False) label = Column(String, nullable=False)
audio_path = Column(String, nullable=False) audio_path = Column(String, nullable=False)
effects_chain = Column(Text, nullable=True) effects_chain = Column(Text, nullable=True)
@@ -172,7 +172,7 @@ class ChannelDeviceMapping(Base):
__tablename__ = "channel_device_mappings" __tablename__ = "channel_device_mappings"
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4())) id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
channel_id = Column(String, ForeignKey("audio_channels.id"), nullable=False) channel_id = Column(String, ForeignKey("audio_channels.id"), nullable=False, index=True)
device_id = Column(String, nullable=False) device_id = Column(String, nullable=False)
@@ -272,7 +272,7 @@ class MCPClientBinding(Base):
label = Column(String, nullable=True) # display name label = Column(String, nullable=True) # display name
profile_id = Column(String, ForeignKey("profiles.id"), nullable=True) profile_id = Column(String, ForeignKey("profiles.id"), nullable=True)
default_engine = Column(String, nullable=True) default_engine = Column(String, nullable=True)
# When true, voicebox.speak routes through the profile's personality LLM # When true, voicebox_speak routes through the profile's personality LLM
# (rewrite) before TTS by default. Callers can still override per call. # (rewrite) before TTS by default. Callers can still override per call.
default_personality = Column(Boolean, nullable=False, default=False) default_personality = Column(Boolean, nullable=False, default=False)
last_seen_at = Column(DateTime, nullable=True) last_seen_at = Column(DateTime, nullable=True)
@@ -300,4 +300,4 @@ class Capture(Base):
stt_model = Column(String, nullable=True) stt_model = Column(String, nullable=True)
llm_model = Column(String, nullable=True) llm_model = Column(String, nullable=True)
refinement_flags = Column(Text, nullable=True) # JSON blob refinement_flags = Column(Text, nullable=True) # JSON blob
created_at = Column(DateTime, default=lambda: datetime.now(UTC)) created_at = Column(DateTime, default=lambda: datetime.now(UTC), index=True)
+16 -1
View File
@@ -3,7 +3,7 @@
import logging import logging
import uuid import uuid
from sqlalchemy import create_engine from sqlalchemy import create_engine, event
from sqlalchemy.orm import sessionmaker from sqlalchemy.orm import sessionmaker
from .. import config from .. import config
@@ -21,6 +21,7 @@ from .seed import backfill_generation_versions, seed_builtin_presets
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# Initialized by init_db() # Initialized by init_db()
engine = None engine = None
SessionLocal = None SessionLocal = None
@@ -39,6 +40,20 @@ def init_db() -> None:
connect_args={"check_same_thread": False}, connect_args={"check_same_thread": False},
) )
@event.listens_for(engine, "connect")
def _set_sqlite_pragmas(dbapi_connection, _record) -> None:
# Each pooled connection enables WAL journal mode and sets a 5-second
# busy timeout. WAL allows concurrent readers during a write (the
# default DELETE/ROLLBACK journal blocks all readers), which matters
# for voicebox because SSE status polls and history queries run
# concurrently with the generation worker writing to the same db.
# busy_timeout prevents "database is locked" errors when two
# connections briefly contend on the same write slot.
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA journal_mode=WAL")
cursor.execute("PRAGMA busy_timeout=5000")
cursor.close()
SessionLocal = sessionmaker(autocommit=False, autoflush=False, bind=engine) SessionLocal = sessionmaker(autocommit=False, autoflush=False, bind=engine)
run_migrations(engine) run_migrations(engine)
+6 -6
View File
@@ -49,10 +49,10 @@ claude mcp add voicebox \
| Name | Purpose | | Name | Purpose |
|---|---| |---|---|
| `voicebox.speak` | Speak text in a voice profile. Returns a generation id you can poll. | | `voicebox_speak` | Speak text in a voice profile. Returns a generation id you can poll. |
| `voicebox.transcribe` | Whisper transcription of a base64 blob or an absolute local path. | | `voicebox_transcribe` | Whisper transcription of a base64 blob or an absolute local path. |
| `voicebox.list_captures` | Recent captures (dictation / recording / file) with transcripts. | | `voicebox_list_captures` | Recent captures (dictation / recording / file) with transcripts. |
| `voicebox.list_profiles` | Available voice profiles (cloned + preset). | | `voicebox_list_profiles` | Available voice profiles (cloned + preset). |
All tools resolve voice profiles in this precedence: All tools resolve voice profiles in this precedence:
@@ -69,8 +69,8 @@ Settings → MCP.
npx @modelcontextprotocol/inspector http://127.0.0.1:17493/mcp npx @modelcontextprotocol/inspector http://127.0.0.1:17493/mcp
``` ```
Point it at the URL, hit "List tools," call `voicebox.list_profiles` Point it at the URL, hit "List tools," call `voicebox_list_profiles`
first to confirm wiring, then `voicebox.speak` for end-to-end. first to confirm wiring, then `voicebox_speak` for end-to-end.
## Non-MCP REST surface ## Non-MCP REST surface
+3 -3
View File
@@ -33,7 +33,7 @@ current_client_id: ContextVar[str | None] = ContextVar(
) )
# Remote address of the in-flight request. Used by tools that gate # Remote address of the in-flight request. Used by tools that gate
# host-filesystem access to loopback callers (see voicebox.transcribe). # host-filesystem access to loopback callers (see voicebox_transcribe).
current_remote_addr: ContextVar[str | None] = ContextVar( current_remote_addr: ContextVar[str | None] = ContextVar(
"current_remote_addr", default=None "current_remote_addr", default=None
) )
@@ -61,12 +61,12 @@ def request_is_loopback() -> bool:
# ignored so the Settings UI's "last heard from" column only reflects # ignored so the Settings UI's "last heard from" column only reflects
# calls that actually acted on the client's bindings. # calls that actually acted on the client's bindings.
# #
# - /mcp — FastMCP tool calls (voicebox.speak, voicebox.transcribe, …) # - /mcp — FastMCP tool calls (voicebox_speak, voicebox_transcribe, …)
# and the /mcp/bindings admin surface. The admin surface is never # and the /mcp/bindings admin surface. The admin surface is never
# called with the header in practice (the frontend manages bindings # called with the header in practice (the frontend manages bindings
# over plain REST), so the `startswith("/mcp")` match doesn't cause # over plain REST), so the `startswith("/mcp")` match doesn't cause
# false stamps. # false stamps.
# - /speak — REST mirror of voicebox.speak for non-MCP agents (shell # - /speak — REST mirror of voicebox_speak for non-MCP agents (shell
# scripts, ACP, A2A). Uses the same per-client binding lookup, so its # scripts, ACP, A2A). Uses the same per-client binding lookup, so its
# callers belong in the last-seen list too. # callers belong in the last-seen list too.
_STAMPED_PATH_PREFIXES: tuple[str, ...] = ("/mcp", "/speak") _STAMPED_PATH_PREFIXES: tuple[str, ...] = ("/mcp", "/speak")
+1 -1
View File
@@ -1,6 +1,6 @@
"""In-memory pub/sub for speaking-pill SSE broadcasts. """In-memory pub/sub for speaking-pill SSE broadcasts.
MCP ``voicebox.speak`` calls and the REST ``POST /speak`` route publish MCP ``voicebox_speak`` calls and the REST ``POST /speak`` route publish
start/end events that DictateWindow subscribes to via /events/speak, so the start/end events that DictateWindow subscribes to via /events/speak, so the
floating pill surfaces whenever an agent is speaking. floating pill surfaces whenever an agent is speaking.
""" """
+2 -2
View File
@@ -27,8 +27,8 @@ def build_mcp_server() -> FastMCP:
mcp = FastMCP( mcp = FastMCP(
name="voicebox", name="voicebox",
instructions=( instructions=(
"Voicebox is a local voice I/O layer. Use `voicebox.speak` to " "Voicebox is a local voice I/O layer. Use `voicebox_speak` to "
"play text in a voice profile, `voicebox.transcribe` for " "play text in a voice profile, `voicebox_transcribe` for "
"audio→text, and the `list_*` tools to discover profiles and " "audio→text, and the `list_*` tools to discover profiles and "
"captures." "captures."
), ),
+10 -9
View File
@@ -1,8 +1,9 @@
"""Voicebox MCP tool implementations. """Voicebox MCP tool implementations.
Thin wrappers over existing services/routes. Tools are registered with dotted Thin wrappers over existing services/routes. Tools are registered with
names (``voicebox.speak`` etc.) so they look natural in agent logs — underscore-separated names (``voicebox_speak`` etc.): MCP clients such as
the Python function name stays snake_case. Claude Desktop validate tool names against ``^[a-zA-Z0-9_-]{1,64}$`` and
reject the whole tool list if any name contains a dot (#790).
""" """
from __future__ import annotations from __future__ import annotations
@@ -36,7 +37,7 @@ def register_tools(mcp: FastMCP) -> None:
"""Attach all Voicebox tools to the given FastMCP instance.""" """Attach all Voicebox tools to the given FastMCP instance."""
@mcp.tool( @mcp.tool(
name="voicebox.speak", name="voicebox_speak",
description=( description=(
"Speak text in a Voicebox voice profile. Returns a generation id " "Speak text in a Voicebox voice profile. Returns a generation id "
"the caller can poll at /generate/{id}/status. Audio plays on the " "the caller can poll at /generate/{id}/status. Audio plays on the "
@@ -104,7 +105,7 @@ def register_tools(mcp: FastMCP) -> None:
profile_name=vp.name, profile_name=vp.name,
text=text, text=text,
engine=resolved_engine, engine=resolved_engine,
language=language, language=language or vp.language,
personality=use_persona, personality=use_persona,
model_size=model_size, model_size=model_size,
db=db, db=db,
@@ -113,7 +114,7 @@ def register_tools(mcp: FastMCP) -> None:
db.close() db.close()
@mcp.tool( @mcp.tool(
name="voicebox.transcribe", name="voicebox_transcribe",
description=( description=(
"Transcribe an audio clip to text using Voicebox's local Whisper. " "Transcribe an audio clip to text using Voicebox's local Whisper. "
"Pass exactly one of `audio_base64` (bytes as base64) or " "Pass exactly one of `audio_base64` (bytes as base64) or "
@@ -171,7 +172,7 @@ def register_tools(mcp: FastMCP) -> None:
tmp_path.unlink(missing_ok=True) tmp_path.unlink(missing_ok=True)
@mcp.tool( @mcp.tool(
name="voicebox.list_captures", name="voicebox_list_captures",
description=( description=(
"List recent voice captures (dictations, recordings, uploads) " "List recent voice captures (dictations, recordings, uploads) "
"with their transcripts. Most-recent first." "with their transcripts. Most-recent first."
@@ -199,10 +200,10 @@ def register_tools(mcp: FastMCP) -> None:
db.close() db.close()
@mcp.tool( @mcp.tool(
name="voicebox.list_profiles", name="voicebox_list_profiles",
description=( description=(
"List available voice profiles (both cloned voices and presets). " "List available voice profiles (both cloned voices and presets). "
"Use the returned `name` with voicebox.speak(profile=...)." "Use the returned `name` with voicebox_speak(profile=...)."
), ),
) )
async def voicebox_list_profiles() -> dict[str, Any]: async def voicebox_list_profiles() -> dict[str, Any]:
+3 -3
View File
@@ -85,7 +85,7 @@ class GenerationRequest(BaseModel):
seed: Optional[int] = Field(None, ge=0) seed: Optional[int] = Field(None, ge=0)
model_size: Optional[str] = Field(default="1.7B", pattern="^(1\\.7B|0\\.6B|1B|3B)$") model_size: Optional[str] = Field(default="1.7B", pattern="^(1\\.7B|0\\.6B|1B|3B)$")
instruct: Optional[str] = Field(None, max_length=500) instruct: Optional[str] = Field(None, max_length=500)
engine: Optional[str] = Field(default="qwen", pattern="^(qwen|qwen_custom_voice|luxtts|chatterbox|chatterbox_turbo|tada|kokoro)$") engine: Optional[str] = Field(default=None, pattern="^(qwen|qwen_custom_voice|luxtts|chatterbox|chatterbox_turbo|tada|kokoro)$")
personality: bool = Field( personality: bool = Field(
default=False, default=False,
description="When true and the profile has a personality prompt, the input text is rewritten in-character before TTS.", description="When true and the profile has a personality prompt, the input text is rewritten in-character before TTS.",
@@ -309,7 +309,7 @@ class GenerationSettingsUpdate(BaseModel):
class MCPClientBindingResponse(BaseModel): class MCPClientBindingResponse(BaseModel):
"""Per-MCP-client voice binding — what voice / engine the server should """Per-MCP-client voice binding — what voice / engine the server should
use when a given client_id calls voicebox.speak without args, plus an use when a given client_id calls voicebox_speak without args, plus an
opt-in personality-rewrite default.""" opt-in personality-rewrite default."""
client_id: str client_id: str
@@ -346,7 +346,7 @@ class MCPClientBindingListResponse(BaseModel):
class SpeakRequest(BaseModel): class SpeakRequest(BaseModel):
"""Body for POST /speak — non-MCP REST surface that mirrors voicebox.speak.""" """Body for POST /speak — non-MCP REST surface that mirrors voicebox_speak."""
text: str = Field(..., min_length=1, max_length=10000) text: str = Field(..., min_length=1, max_length=10000)
profile: Optional[str] = Field( profile: Optional[str] = Field(
+1 -1
View File
@@ -64,7 +64,7 @@ pedalboard>=0.9.0
httpx>=0.27.0 httpx>=0.27.0
# MCP server (Model Context Protocol) — lets local AI agents call # MCP server (Model Context Protocol) — lets local AI agents call
# voicebox.speak / .transcribe / .list_captures / .list_profiles # voicebox_speak / voicebox_transcribe / voicebox_list_captures / voicebox_list_profiles
fastmcp>=3.0,<4.0 fastmcp>=3.0,<4.0
sse-starlette>=2.0 sse-starlette>=2.0
+3 -3
View File
@@ -244,6 +244,7 @@ async def get_model_status():
use_scan_cache = False use_scan_cache = False
from ..backends import get_all_model_configs, check_model_loaded from ..backends import get_all_model_configs, check_model_loaded
from ..backends.base import has_in_progress_download
registry_configs = get_all_model_configs() registry_configs = get_all_model_configs()
model_configs = [ model_configs = [
@@ -293,8 +294,7 @@ async def get_model_status():
try: try:
cache_dir = hf_constants.HF_HUB_CACHE cache_dir = hf_constants.HF_HUB_CACHE
blobs_dir = Path(cache_dir) / ("models--" + repo_id.replace("/", "--")) / "blobs" blobs_dir = Path(cache_dir) / ("models--" + repo_id.replace("/", "--")) / "blobs"
if blobs_dir.exists(): has_incomplete = has_in_progress_download(blobs_dir)
has_incomplete = any(blobs_dir.glob("*.incomplete"))
except Exception: except Exception:
pass pass
@@ -314,7 +314,7 @@ async def get_model_status():
if repo_cache.exists(): if repo_cache.exists():
blobs_dir = repo_cache / "blobs" blobs_dir = repo_cache / "blobs"
has_incomplete = blobs_dir.exists() and any(blobs_dir.glob("*.incomplete")) has_incomplete = has_in_progress_download(blobs_dir)
if not has_incomplete: if not has_incomplete:
snapshots_dir = repo_cache / "snapshots" snapshots_dir = repo_cache / "snapshots"
+3 -3
View File
@@ -1,4 +1,4 @@
"""POST /speak — REST wrapper around voicebox.speak for non-MCP callers. """POST /speak — REST wrapper around voicebox_speak for non-MCP callers.
Shell scripts, ACP, A2A, or any agent that doesn't speak MCP can hit this Shell scripts, ACP, A2A, or any agent that doesn't speak MCP can hit this
endpoint to play text through a cloned voice. Uses the same profile endpoint to play text through a cloned voice. Uses the same profile
@@ -30,7 +30,7 @@ async def speak(
request: Request, request: Request,
db: Session = Depends(get_db), db: Session = Depends(get_db),
): ):
"""Speak text in a voice profile. Mirrors voicebox.speak (MCP). """Speak text in a voice profile. Mirrors voicebox_speak (MCP).
Response shape matches POST /generate — a ``GenerationResponse`` with Response shape matches POST /generate — a ``GenerationResponse`` with
``status="generating"`` and an ``id`` the caller polls at ``status="generating"`` and an ``id`` the caller polls at
@@ -75,7 +75,7 @@ async def speak(
models.GenerationRequest( models.GenerationRequest(
profile_id=profile.id, profile_id=profile.id,
text=data.text, text=data.text,
language=data.language or "en", language=data.language or profile.language or "en",
engine=engine, engine=engine,
personality=bool(personality_flag), personality=bool(personality_flag),
), ),
+37
View File
@@ -27,6 +27,43 @@ if not _is_writable(sys.stdout):
if not _is_writable(sys.stderr): if not _is_writable(sys.stderr):
sys.stderr = open(os.devnull, 'w') sys.stderr = open(os.devnull, 'w')
class _PipeSafeStream:
"""Falls back to devnull once the pipe to the Tauri app is gone.
The app reads our stdout/stderr through a pipe. When the server outlives
it (keep-running mode, or a sidecar the next launch reuses), every later
print()/tqdm write raises "[Errno 32] Broken pipe" and fails whatever
request triggered it, e.g. POST /captures.
"""
def __init__(self, stream):
self._stream = stream
def write(self, s):
try:
return self._stream.write(s)
except (OSError, ValueError):
self._stream = open(os.devnull, 'w')
return len(s)
def writelines(self, lines):
for line in lines:
self.write(line)
def flush(self):
try:
self._stream.flush()
except (OSError, ValueError):
self._stream = open(os.devnull, 'w')
def __getattr__(self, name):
return getattr(self._stream, name)
sys.stdout = _PipeSafeStream(sys.stdout)
sys.stderr = _PipeSafeStream(sys.stderr)
# PyInstaller + multiprocessing: child processes re-execute the frozen binary # PyInstaller + multiprocessing: child processes re-execute the frozen binary
# with internal arguments. freeze_support() handles this and exits early. # with internal arguments. freeze_support() handles this and exits early.
import multiprocessing import multiprocessing
+36 -10
View File
@@ -15,21 +15,30 @@ from ..database import Generation as DBGeneration, GenerationVersion as DBGenera
from .. import config from .. import config
def _get_versions_for_generation(generation_id: str, db: Session) -> tuple: def _get_versions_for_generations(generation_ids: list[str], db: Session) -> dict:
"""Get versions list and active version ID for a generation.""" """Fetch versions for many generations in a single query.
Returns a mapping of ``generation_id -> (versions, active_version_id)``
using the same shape as ``_get_versions_for_generation()``, so callers
can batch a whole page of generations without an N+1 query.
"""
import json import json
ids = list(dict.fromkeys(generation_ids))
if not ids:
return {}
versions_rows = ( versions_rows = (
db.query(DBGenerationVersion) db.query(DBGenerationVersion)
.filter_by(generation_id=generation_id) .filter(DBGenerationVersion.generation_id.in_(ids))
.order_by(DBGenerationVersion.created_at) .order_by(DBGenerationVersion.created_at)
.all() .all()
) )
if not versions_rows:
return None, None
versions = [] versions_by_generation: dict[str, list] = {}
active_version_id = None active_by_generation: dict[str, Optional[str]] = {}
for v in versions_rows: for v in versions_rows:
versions = versions_by_generation.setdefault(v.generation_id, [])
effects_chain = None effects_chain = None
if v.effects_chain: if v.effects_chain:
try: try:
@@ -47,9 +56,20 @@ def _get_versions_for_generation(generation_id: str, db: Session) -> tuple:
created_at=v.created_at, created_at=v.created_at,
)) ))
if v.is_default: if v.is_default:
active_version_id = v.id active_by_generation[v.generation_id] = v.id
return versions, active_version_id return {
generation_id: (
versions_by_generation.get(generation_id),
active_by_generation.get(generation_id),
)
for generation_id in ids
}
def _get_versions_for_generation(generation_id: str, db: Session) -> tuple:
"""Get versions list and active version ID for a single generation."""
return _get_versions_for_generations([generation_id], db)[generation_id]
async def create_generation( async def create_generation(
@@ -205,10 +225,16 @@ async def list_generations(
# Execute query # Execute query
results = q.all() results = q.all()
# Fetch versions for every generation on this page with a single
# query instead of one SELECT per generation (N+1).
versions_by_generation = _get_versions_for_generations(
[generation.id for generation, _ in results], db
)
# Convert to HistoryResponse with profile_name # Convert to HistoryResponse with profile_name
items = [] items = []
for generation, profile_name in results: for generation, profile_name in results:
versions, active_version_id = _get_versions_for_generation(generation.id, db) versions, active_version_id = versions_by_generation[generation.id]
items.append(HistoryResponse( items.append(HistoryResponse(
id=generation.id, id=generation.id,
profile_id=generation.profile_id, profile_id=generation.profile_id,
+80
View File
@@ -0,0 +1,80 @@
"""
Regression tests for the VOICEBOX_FORCE_CPU environment override.
The docs promise (docs/content/docs/developer/tts-generation.mdx) that
get_torch_device() layers "VOICEBOX_FORCE_CPU environment override" ahead of
CUDA/XPU/MPS detection, and gpu-acceleration.mdx tells users to set it to fall
back to CPU when the bundled PyTorch has no kernels for their GPU.
torch is stubbed through sys.modules so these run without a torch install and
without any GPU.
Usage:
python -m pytest backend/tests/test_force_cpu_env.py -v
"""
import sys
import pytest
from backend.backends.base import get_torch_device
# The documented public name and value of the override. Pinned here independently
# of the production constants so a rename of either fails these tests.
FORCE_CPU_ENV_VAR = "VOICEBOX_FORCE_CPU"
# Sentinel for "the variable is not set at all".
UNSET = None
class _FakeCuda:
@staticmethod
def is_available() -> bool:
return True
class _FakeTorch:
"""Minimal stand-in for a CUDA-enabled torch install."""
cuda = _FakeCuda
@pytest.fixture
def cuda_available(monkeypatch):
"""Make torch report a usable CUDA device without installing torch."""
monkeypatch.setitem(sys.modules, "torch", _FakeTorch)
def _set_override(monkeypatch, value):
if value is UNSET:
monkeypatch.delenv(FORCE_CPU_ENV_VAR, raising=False)
else:
monkeypatch.setenv(FORCE_CPU_ENV_VAR, value)
@pytest.mark.parametrize("value", ["1", " 1 "])
def test_force_cpu_wins_over_available_cuda(monkeypatch, cuda_available, value):
"""The documented value must beat an otherwise usable CUDA device.
Surrounding whitespace is tolerated: on Windows, where this override
matters most, it is typically set through the GUI environment editor."""
_set_override(monkeypatch, value)
assert get_torch_device(allow_xpu=True, allow_directml=True, allow_mps=True) == "cpu"
@pytest.mark.parametrize("value", [UNSET, "", "0"])
def test_without_override_cuda_is_still_selected(monkeypatch, cuda_available, value):
"""Unset or disabled must not disturb normal device detection."""
_set_override(monkeypatch, value)
assert get_torch_device() == "cuda"
def test_force_cpu_does_not_need_torch(monkeypatch):
"""The override is honoured before torch is imported, so it works on a
broken/incompatible torch install — which is the case it exists for."""
_set_override(monkeypatch, "1")
monkeypatch.setitem(sys.modules, "torch", None) # makes `import torch` raise
assert get_torch_device() == "cpu"
@@ -0,0 +1,27 @@
"""Tests for generation request engine selection."""
import pytest
from pydantic import ValidationError
from backend import models
def _request(**kwargs) -> models.GenerationRequest:
return models.GenerationRequest(profile_id="profile-1", text="hello", **kwargs)
def test_omitted_engine_does_not_override_profile_default():
request = _request()
assert request.engine is None
def test_explicit_engine_is_preserved():
request = _request(engine="chatterbox")
assert request.engine == "chatterbox"
def test_invalid_explicit_engine_is_rejected():
with pytest.raises(ValidationError):
_request(engine="invalid")
+74
View File
@@ -0,0 +1,74 @@
"""
Unit tests for ``is_model_cached``'s handling of stale ``.incomplete`` blobs.
A retried or concurrent download can leave an orphaned ``.incomplete`` file
next to its now-completed counterpart (same blob hash, no suffix). Only a
genuinely in-progress download -- an ``.incomplete`` with no completed blob
alongside it -- should mark the model as not cached.
``is_model_cached`` is extracted and exec'd standalone (instead of importing
``backend.backends.base``) so this test doesn't pull in the module's sibling
imports (audio/progress/hf_progress/tasks), which in turn require the full
ML stack (torch/transformers/librosa/fastapi/...) this pure filesystem check
never touches.
"""
import ast
import logging
from pathlib import Path
from typing import Optional
_SOURCE = (Path(__file__).parent.parent / "backends" / "base.py").read_text()
_MODULE = ast.parse(_SOURCE)
_FUNC_SRC = "\n\n".join(
ast.get_source_segment(_SOURCE, node)
for node in _MODULE.body
if isinstance(node, ast.FunctionDef) and node.name in ("has_in_progress_download", "is_model_cached")
)
_namespace = {
"Path": Path,
"Optional": Optional,
"logger": logging.getLogger("test_is_model_cached"),
}
exec(_FUNC_SRC, _namespace) # noqa: S102
is_model_cached = _namespace["is_model_cached"]
def _make_repo_cache(tmp_path, repo="org/model"):
repo_dir = tmp_path / ("models--" + repo.replace("/", "--"))
blobs_dir = repo_dir / "blobs"
snapshots_dir = repo_dir / "snapshots" / "abc123"
blobs_dir.mkdir(parents=True)
snapshots_dir.mkdir(parents=True)
return repo_dir, blobs_dir, snapshots_dir
def test_orphaned_incomplete_blob_does_not_block_cache_hit(tmp_path, monkeypatch):
import huggingface_hub.constants as hf_constants
monkeypatch.setattr(hf_constants, "HF_HUB_CACHE", str(tmp_path))
repo = "org/model"
_, blobs_dir, snapshots_dir = _make_repo_cache(tmp_path, repo)
completed_blob = blobs_dir / "deadbeef"
completed_blob.write_bytes(b"weights")
(blobs_dir / "deadbeef.incomplete").write_bytes(b"stale partial")
(snapshots_dir / "model.safetensors").symlink_to(completed_blob)
assert is_model_cached(repo) is True
def test_genuinely_in_progress_download_is_not_cached(tmp_path, monkeypatch):
import huggingface_hub.constants as hf_constants
monkeypatch.setattr(hf_constants, "HF_HUB_CACHE", str(tmp_path))
repo = "org/model"
_, blobs_dir, snapshots_dir = _make_repo_cache(tmp_path, repo)
(blobs_dir / "feedface.incomplete").write_bytes(b"partial")
(snapshots_dir / "config.json").write_text("{}")
assert is_model_cached(repo) is False
+1 -1
View File
@@ -1,4 +1,4 @@
"""Tests for the voicebox.speak MCP tool's ``model_size`` plumbing (issue #884). """Tests for the voicebox_speak MCP tool's ``model_size`` plumbing (issue #884).
The MCP speak path used to build its ``GenerationRequest`` without a The MCP speak path used to build its ``GenerationRequest`` without a
``model_size``, so every agent-triggered generation silently fell back to the ``model_size``, so every agent-triggered generation silently fell back to the
+158
View File
@@ -0,0 +1,158 @@
"""Tests for voice-profile language fallback on the two speak surfaces.
Both speak paths built their ``GenerationRequest`` with a hardcoded ``"en"``
fallback and never consulted the resolved profile, so a profile created with
``language="fr"`` was still synthesised as English unless every caller passed
``language=`` explicitly. Agents going through MCP had no way to know the
profile's language, so they couldn't pass it either.
These tests pin the fix: the fallback chain is now explicit argument →
resolved profile's language → ``"en"``, matching how ``engine`` and
``personality`` already consult the resolved binding.
"""
import pytest
import backend.routes.generations as generations
import backend.routes.speak as speak_route
from backend import models
from backend.mcp_server import tools
class _FakeGeneration:
"""Minimal stand-in for GenerationResponse consumed by the speak paths."""
id = "gen-test"
status = "generating"
def model_dump(self, mode="json"):
return {"id": self.id, "status": self.status}
class _FakeProfile:
def __init__(self, language):
self.id = "p1"
self.name = "Siwis"
self.language = language
self.personality = None
class _FakeQuery:
def filter(self, *args, **kwargs):
return self
def first(self):
# No per-client binding — engine/personality fall through to their
# own defaults, leaving language as the only variable under test.
return None
class _FakeDB:
def query(self, *args, **kwargs):
return _FakeQuery()
def close(self):
pass
class _FakeRequest:
"""Stands in for starlette's Request — only headers are read."""
def __init__(self, client_id=None):
self.headers = {"X-Voicebox-Client-Id": client_id} if client_id else {}
@pytest.fixture
def captured_request(monkeypatch):
"""Capture the GenerationRequest instead of running a real generation.
Both speak paths import ``generate_speech`` lazily from
``routes.generations``, so patching the attribute on that module
intercepts the call on either surface.
"""
captured = {}
async def fake_generate_speech(req, db):
captured["req"] = req
return _FakeGeneration()
monkeypatch.setattr(generations, "generate_speech", fake_generate_speech)
monkeypatch.setattr(speak_route.mcp_events, "publish", lambda *a, **k: None)
monkeypatch.setattr(tools.mcp_events, "publish", lambda *a, **k: None)
return captured
# ─── REST: POST /speak ────────────────────────────────────────────────────
async def _call_rest(monkeypatch, profile_language, requested_language=None):
monkeypatch.setattr(
speak_route,
"resolve_profile",
lambda profile, client_id, db: _FakeProfile(profile_language),
)
await speak_route.speak(
models.SpeakRequest(text="Bonjour", language=requested_language),
_FakeRequest(client_id="claude-code"),
_FakeDB(),
)
async def test_rest_speak_falls_back_to_profile_language(captured_request, monkeypatch):
await _call_rest(monkeypatch, profile_language="fr")
assert captured_request["req"].language == "fr"
async def test_rest_speak_explicit_language_wins(captured_request, monkeypatch):
# An explicit argument still overrides the profile — a French profile can
# be asked to read an English string.
await _call_rest(monkeypatch, profile_language="fr", requested_language="en")
assert captured_request["req"].language == "en"
async def test_rest_speak_defaults_to_en_without_profile_language(captured_request, monkeypatch):
# Profiles predating the language column resolve to None; the "en"
# backstop keeps their behaviour unchanged.
await _call_rest(monkeypatch, profile_language=None)
assert captured_request["req"].language == "en"
# ─── MCP: voicebox.speak ──────────────────────────────────────────────────
async def _call_mcp(monkeypatch, profile_language, requested_language=None):
# Build the server from the same ``fastmcp`` package production imports so
# the registered ``voicebox.speak`` wrapper — where the profile fallback
# lives — is the code under test.
from fastmcp import FastMCP
monkeypatch.setattr(
tools,
"resolve_profile",
lambda profile, client_id, db: _FakeProfile(profile_language),
)
monkeypatch.setattr(tools, "get_db", lambda: iter([_FakeDB()]))
mcp = FastMCP("test")
tools.register_tools(mcp)
args = {"text": "Bonjour"}
if requested_language is not None:
args["language"] = requested_language
await mcp.call_tool("voicebox.speak", args)
async def test_mcp_speak_falls_back_to_profile_language(captured_request, monkeypatch):
# The agent-facing path matters most: an MCP client can't know the
# profile's language, so omitting it must not silently mean English.
await _call_mcp(monkeypatch, profile_language="fr")
assert captured_request["req"].language == "fr"
async def test_mcp_speak_explicit_language_wins(captured_request, monkeypatch):
await _call_mcp(monkeypatch, profile_language="fr", requested_language="en")
assert captured_request["req"].language == "en"
async def test_mcp_speak_defaults_to_en_without_profile_language(captured_request, monkeypatch):
await _call_mcp(monkeypatch, profile_language=None)
assert captured_request["req"].language == "en"
+4 -3
View File
@@ -152,6 +152,9 @@
}, },
}, },
}, },
"overrides": {
"seroval": "1.5.3",
},
"packages": { "packages": {
"@alloc/quick-lru": ["@alloc/[email protected]", "", {}, "sha512-UrcABB+4bUrFABwbluTIBErXwvbsU/V7TZWfmbgJfbkwiBuziS9gxdODUyuiecfdGQ85jglMW6juS3+z5TsKLw=="], "@alloc/quick-lru": ["@alloc/[email protected]", "", {}, "sha512-UrcABB+4bUrFABwbluTIBErXwvbsU/V7TZWfmbgJfbkwiBuziS9gxdODUyuiecfdGQ85jglMW6juS3+z5TsKLw=="],
@@ -1051,7 +1054,7 @@
"semver": ["[email protected]", "", { "bin": { "semver": "bin/semver.js" } }, "sha512-BR7VvDCVHO+q2xBEWskxS6DJE1qRnb7DxzUrogb71CWoSficBxYsiAGd+Kl0mmq/MprG9yArRkyrQxTO6XjMzA=="], "semver": ["[email protected]", "", { "bin": { "semver": "bin/semver.js" } }, "sha512-BR7VvDCVHO+q2xBEWskxS6DJE1qRnb7DxzUrogb71CWoSficBxYsiAGd+Kl0mmq/MprG9yArRkyrQxTO6XjMzA=="],
"seroval": ["[email protected].0", "", {}, "sha512-OE4cvmJ1uSPrKorFIH9/w/Qwuvi/IMcGbv5RKgcJ/zjA/IohDLU6SVaxFN9FwajbP7nsX0dQqMDes1whk3y+yw=="], "seroval": ["[email protected].3", "", {}, "sha512-BXe0x4buEeYiIKaRUnth1WqCILQ3k4O67KP/B4pC3pVz0Mv2c96ngA9QDREUYxWY1sb2RZVRqwI9RcpVMyHCVw=="],
"seroval-plugins": ["[email protected]", "", { "peerDependencies": { "seroval": "^1.0" } }, "sha512-EAHqADIQondwRZIdeW2I636zgsODzoBDwb3PT/+7TLDWyw1Dy/Xv7iGUIEXXav7usHDE9HVhOU61irI3EnyyHA=="], "seroval-plugins": ["[email protected]", "", { "peerDependencies": { "seroval": "^1.0" } }, "sha512-EAHqADIQondwRZIdeW2I636zgsODzoBDwb3PT/+7TLDWyw1Dy/Xv7iGUIEXXav7usHDE9HVhOU61irI3EnyyHA=="],
@@ -1203,8 +1206,6 @@
"@voicebox/landing/tailwindcss": ["[email protected]", "", { "dependencies": { "@alloc/quick-lru": "^5.2.0", "arg": "^5.0.2", "chokidar": "^3.6.0", "didyoumean": "^1.2.2", "dlv": "^1.1.3", "fast-glob": "^3.3.2", "glob-parent": "^6.0.2", "is-glob": "^4.0.3", "jiti": "^1.21.7", "lilconfig": "^3.1.3", "micromatch": "^4.0.8", "normalize-path": "^3.0.0", "object-hash": "^3.0.0", "picocolors": "^1.1.1", "postcss": "^8.4.47", "postcss-import": "^15.1.0", "postcss-js": "^4.0.1", "postcss-load-config": "^4.0.2 || ^5.0 || ^6.0", "postcss-nested": "^6.2.0", "postcss-selector-parser": "^6.1.2", "resolve": "^1.22.8", "sucrase": "^3.35.0" }, "bin": { "tailwind": "lib/cli.js", "tailwindcss": "lib/cli.js" } }, "sha512-3ofp+LL8E+pK/JuPLPggVAIaEuhvIz4qNcf3nA1Xn2o/7fb7s/TYpHhwGDv1ZU3PkBluUVaF8PyCHcm48cKLWQ=="], "@voicebox/landing/tailwindcss": ["[email protected]", "", { "dependencies": { "@alloc/quick-lru": "^5.2.0", "arg": "^5.0.2", "chokidar": "^3.6.0", "didyoumean": "^1.2.2", "dlv": "^1.1.3", "fast-glob": "^3.3.2", "glob-parent": "^6.0.2", "is-glob": "^4.0.3", "jiti": "^1.21.7", "lilconfig": "^3.1.3", "micromatch": "^4.0.8", "normalize-path": "^3.0.0", "object-hash": "^3.0.0", "picocolors": "^1.1.1", "postcss": "^8.4.47", "postcss-import": "^15.1.0", "postcss-js": "^4.0.1", "postcss-load-config": "^4.0.2 || ^5.0 || ^6.0", "postcss-nested": "^6.2.0", "postcss-selector-parser": "^6.1.2", "resolve": "^1.22.8", "sucrase": "^3.35.0" }, "bin": { "tailwind": "lib/cli.js", "tailwindcss": "lib/cli.js" } }, "sha512-3ofp+LL8E+pK/JuPLPggVAIaEuhvIz4qNcf3nA1Xn2o/7fb7s/TYpHhwGDv1ZU3PkBluUVaF8PyCHcm48cKLWQ=="],
"@voicebox/web/wavesurfer.js": ["[email protected]", "", {}, "sha512-NswPjVHxk0Q1F/VMRemCPUzSojjuHHisQrBqQiRXg7MVbe3f5vQ6r0rTTXA/a/neC/4hnOEC4YpXca4LpH0SUg=="],
"chokidar/glob-parent": ["[email protected]", "", { "dependencies": { "is-glob": "^4.0.1" } }, "sha512-AOIgSQCepiJYwP3ARnGx+5VnTu2HBYdzbGP45eLw1vr3zB3vZLeyed1sC9hnbcOc9/SrMyM5RPQrkGz4aS9Zow=="], "chokidar/glob-parent": ["[email protected]", "", { "dependencies": { "is-glob": "^4.0.1" } }, "sha512-AOIgSQCepiJYwP3ARnGx+5VnTu2HBYdzbGP45eLw1vr3zB3vZLeyed1sC9hnbcOc9/SrMyM5RPQrkGz4aS9Zow=="],
"eslint/js-yaml": ["[email protected]", "", { "dependencies": { "argparse": "^2.0.1" }, "bin": { "js-yaml": "bin/js-yaml.js" } }, "sha512-qQKT4zQxXl8lLwBtHMWwaTcGfFOZviOJet3Oy/xmGk2gZH677CJM9EvtfdSkgWcATZhj/55JZ0rmy3myCT5lsA=="], "eslint/js-yaml": ["[email protected]", "", { "dependencies": { "argparse": "^2.0.1" }, "bin": { "js-yaml": "bin/js-yaml.js" } }, "sha512-qQKT4zQxXl8lLwBtHMWwaTcGfFOZviOJet3Oy/xmGk2gZH677CJM9EvtfdSkgWcATZhj/55JZ0rmy3myCT5lsA=="],
+1 -1
View File
@@ -121,7 +121,7 @@ POST /generate
Shipped 2026-04-25 (PR #544). Voicebox went from a voice-cloning studio to a full voice studio — dictation in, agent speech out, a local LLM in the middle. Shipped 2026-04-25 (PR #544). Voicebox went from a voice-cloning studio to a full voice studio — dictation in, agent speech out, a local LLM in the middle.
- **Dictation** — global hotkey capture (push-to-talk + toggle chords), on-screen pill with live state, auto-paste into the focused field with clipboard save/restore, chord-picker UI. Scoped Accessibility permission (transcripts still land if paste is denied). - **Dictation** — global hotkey capture (push-to-talk + toggle chords), on-screen pill with live state, auto-paste into the focused field with clipboard save/restore, chord-picker UI. Scoped Accessibility permission (transcripts still land if paste is denied).
- **MCP server** at `http://127.0.0.1:17493/mcp` — `voicebox.speak` / `.transcribe` / `.list_captures` / `.list_profiles`. Streamable HTTP primary transport, stdio sidecar shim, per-client voice binding via `X-Voicebox-Client-Id`. Speaking pill always shows agent-initiated output. - **MCP server** at `http://127.0.0.1:17493/mcp` — `voicebox_speak` / `voicebox_transcribe` / `voicebox_list_captures` / `voicebox_list_profiles`. Streamable HTTP primary transport, stdio sidecar shim, per-client voice binding via `X-Voicebox-Client-Id`. Speaking pill always shows agent-initiated output.
- **Personality** — voice profiles carry an optional ≤2000-char persona. Compose (shuffle an in-character line) and Speak-in-character (rewrite input before TTS), both on a local Qwen3 LLM that doubles as the refinement model. - **Personality** — voice profiles carry an optional ≤2000-char persona. Compose (shuffle an in-character line) and Speak-in-character (rewrite input before TTS), both on a local Qwen3 LLM that doubles as the refinement model.
- **Refinement** — on-device Qwen3 strips fillers, fixes punctuation, optional self-correction rewrites; Whisper hallucination-loop stripping at a 6-token threshold; per-capture flag snapshots; model picker (0.6B / 1.7B / 4B). - **Refinement** — on-device Qwen3 strips fillers, fixes punctuation, optional self-correction rewrites; Whisper hallucination-loop stripping at a 6-token threshold; per-capture flag snapshots; model picker (0.6B / 1.7B / 4B).
- **`POST /speak` REST wrapper** and **i18next foundation** (English + zh-CN) also landed. - **`POST /speak` REST wrapper** and **i18next foundation** (English + zh-CN) also landed.
+15 -15
View File
@@ -105,15 +105,15 @@ for the shim to connect.
| Tool | Use | | Tool | Use |
|---|---| |---|---|
| `voicebox.speak` | Speak text in a voice profile. Returns a `generation_id` to poll. | | `voicebox_speak` | Speak text in a voice profile. Returns a `generation_id` to poll. |
| `voicebox.transcribe` | Whisper transcription of base64 audio or an absolute local path. | | `voicebox_transcribe` | Whisper transcription of base64 audio or an absolute local path. |
| `voicebox.list_captures` | Recent captures with transcripts, paginated. | | `voicebox_list_captures` | Recent captures with transcripts, paginated. |
| `voicebox.list_profiles` | Available voice profiles (cloned + preset). | | `voicebox_list_profiles` | Available voice profiles (cloned + preset). |
### `voicebox.speak` ### `voicebox_speak`
```ts ```ts
voicebox.speak({ voicebox_speak({
text: "Deploy complete.", text: "Deploy complete.",
profile?: "Morgan", // name or id; falls back to per-client binding, then default profile?: "Morgan", // name or id; falls back to per-client binding, then default
engine?: "qwen", // qwen | qwen_custom_voice | luxtts | chatterbox | chatterbox_turbo | tada | kokoro engine?: "qwen", // qwen | qwen_custom_voice | luxtts | chatterbox | chatterbox_turbo | tada | kokoro
@@ -138,10 +138,10 @@ Returns:
- **Persona mode** — `personality: true` and the profile must have a personality prompt set. - **Persona mode** — `personality: true` and the profile must have a personality prompt set.
The LLM rewrites the text in character before TTS. See [Voice Personalities](/overview/voice-personalities). The LLM rewrites the text in character before TTS. See [Voice Personalities](/overview/voice-personalities).
### `voicebox.transcribe` ### `voicebox_transcribe`
```ts ```ts
voicebox.transcribe({ voicebox_transcribe({
audio_base64?: "<base64>", // exactly one of these two audio_base64?: "<base64>", // exactly one of these two
audio_path?: "/absolute/path/to/file.wav", audio_path?: "/absolute/path/to/file.wav",
language?: "en", language?: "en",
@@ -151,18 +151,18 @@ voicebox.transcribe({
Returns `{ text, duration, language, model }`. 200 MB ceiling on either path. Returns `{ text, duration, language, model }`. 200 MB ceiling on either path.
### `voicebox.list_captures` ### `voicebox_list_captures`
`{ limit?: 20, offset?: 0 }` → `{ captures: [...], total }`. `limit` is `{ limit?: 20, offset?: 0 }` → `{ captures: [...], total }`. `limit` is
clamped to `1..=200`. clamped to `1..=200`.
### `voicebox.list_profiles` ### `voicebox_list_profiles`
No args → `{ profiles: [{ id, name, voice_type, language, has_personality }] }`. No args → `{ profiles: [{ id, name, voice_type, language, has_personality }] }`.
## Voice resolution ## Voice resolution
Every call to `voicebox.speak` (and `POST /speak`) resolves the voice profile Every call to `voicebox_speak` (and `POST /speak`) resolves the voice profile
in this order: in this order:
<Steps> <Steps>
@@ -195,7 +195,7 @@ carries:
| `label` | Display name in the Settings UI (e.g. "Claude Code"). | | `label` | Display name in the Settings UI (e.g. "Claude Code"). |
| `profile_id` | The voice this client uses when `profile` isn't passed. | | `profile_id` | The voice this client uses when `profile` isn't passed. |
| `default_engine` | Override the TTS engine for this client. | | `default_engine` | Override the TTS engine for this client. |
| `default_personality` | When true, `voicebox.speak` routes through the profile's personality LLM (rewrite) by default. | | `default_personality` | When true, `voicebox_speak` routes through the profile's personality LLM (rewrite) by default. |
| `last_seen_at` | Last time the server saw a request from this client. | | `last_seen_at` | Last time the server saw a request from this client. |
`last_seen_at` is stamped automatically by middleware on every `/mcp/*` `last_seen_at` is stamped automatically by middleware on every `/mcp/*`
@@ -239,8 +239,8 @@ agent:
npx @modelcontextprotocol/inspector http://127.0.0.1:17493/mcp npx @modelcontextprotocol/inspector http://127.0.0.1:17493/mcp
``` ```
Start with `voicebox.list_profiles` to confirm wiring, then Start with `voicebox_list_profiles` to confirm wiring, then
`voicebox.speak` for end-to-end — you should hear audio and see the `voicebox_speak` for end-to-end — you should hear audio and see the
generation land in the Captures tab. generation land in the Captures tab.
<Callout type="info"> <Callout type="info">
@@ -262,7 +262,7 @@ generation land in the Captures tab.
boundary. If you're scripting against a shared host, prefer boundary. If you're scripting against a shared host, prefer
`audio_base64` so you don't have to think about path sandboxing. `audio_base64` so you don't have to think about path sandboxing.
- **Voice cloning consent applies.** See [Voice Cloning](/overview/voice-cloning#limitations) - **Voice cloning consent applies.** See [Voice Cloning](/overview/voice-cloning#limitations)
— an agent being able to call `voicebox.speak` in someone's voice — an agent being able to call `voicebox_speak` in someone's voice
doesn't change the ethics of whose voices you clone. doesn't change the ethics of whose voices you clone.
## Implementation notes ## Implementation notes
@@ -419,12 +419,14 @@ bun run tauri build
```bash ```bash
# macOS # macOS
rm ~/Library/Application\ Support/sh.voicebox.app/data/voicebox.db rm ~/Library/Application\ Support/sh.voicebox.app/data/voicebox.db*
# Windows # Windows
del %APPDATA%\sh.voicebox.app\data\voicebox.db del %APPDATA%\sh.voicebox.app\data\voicebox.db*
``` ```
The wildcard also removes the `voicebox.db-wal` and `voicebox.db-shm` sidecar files that SQLite keeps next to the database in WAL mode, so the reset starts completely clean.
Restart the app to create a fresh database. Restart the app to create a fresh database.
## Model Issues ## Model Issues
@@ -128,7 +128,7 @@ until you flip it back off.
- **Agents that speak in a voice you own.** Combine the persona toggle with - **Agents that speak in a voice you own.** Combine the persona toggle with
the built-in [MCP Server](/overview/mcp-server) so Claude Code, Cursor, the built-in [MCP Server](/overview/mcp-server) so Claude Code, Cursor,
Cline, or any MCP-aware agent can talk back through a profile with a Cline, or any MCP-aware agent can talk back through a profile with a
personality. The agent calls `voicebox.speak({ text, profile, personality: personality. The agent calls `voicebox_speak({ text, profile, personality:
true })` and Voicebox rewrites the text in character before speaking. true })` and Voicebox rewrites the text in character before speaking.
- **Interactive characters.** Games, narrative tools, accessibility - **Interactive characters.** Games, narrative tools, accessibility
experiences. A character with a personality description plus a cloned experiences. A character with a personality description plus a cloned
@@ -151,7 +151,7 @@ Personalities are accessible via REST:
| `POST` | `/generate` | Include `personality: true` to run input text through the personality LLM before TTS. Same for `POST /speak`. | | `POST` | `/generate` | Include `personality: true` to run input text through the personality LLM before TTS. Same for `POST /speak`. |
`POST /generate` with `personality: true` is the same primitive MCP's `POST /generate` with `personality: true` is the same primitive MCP's
`voicebox.speak` tool uses when you pass `personality: true`. Scripts and `voicebox_speak` tool uses when you pass `personality: true`. Scripts and
agents can use it directly. agents can use it directly.
## Limits and gotchas ## Limits and gotchas
+2 -2
View File
@@ -348,12 +348,12 @@ db-init: _ensure-venv
# Reset database (delete + reinit) # Reset database (delete + reinit)
[unix] [unix]
db-reset: db-reset:
rm -f {{ backend_dir }}/data/voicebox.db rm -f {{ backend_dir }}/data/voicebox.db {{ backend_dir }}/data/voicebox.db-wal {{ backend_dir }}/data/voicebox.db-shm
just db-init just db-init
[windows] [windows]
db-reset: db-reset:
if (Test-Path "{{ backend_dir }}/data/voicebox.db") { Remove-Item -Force "{{ backend_dir }}/data/voicebox.db" } Remove-Item -Force -ErrorAction SilentlyContinue "{{ backend_dir }}/data/voicebox.db", "{{ backend_dir }}/data/voicebox.db-wal", "{{ backend_dir }}/data/voicebox.db-shm"
just db-init just db-init
# ─── Utilities ──────────────────────────────────────────────────────── # ─── Utilities ────────────────────────────────────────────────────────
+5 -5
View File
@@ -23,7 +23,7 @@ const SCENARIOS: Scenario[] = [
{ prefix: '$', text: 'claude run', tone: 'accent' }, { prefix: '$', text: 'claude run', tone: 'accent' },
{ prefix: '✓', text: 'Tests passing (42 files)', tone: 'success' }, { prefix: '✓', text: 'Tests passing (42 files)', tone: 'success' },
{ prefix: '✓', text: 'Build succeeded in 12.4s', tone: 'success' }, { prefix: '✓', text: 'Build succeeded in 12.4s', tone: 'success' },
{ prefix: '→', text: 'voicebox.speak({ profile: "Morgan" })', tone: 'dim' }, { prefix: '→', text: 'voicebox_speak({ profile: "Morgan" })', tone: 'dim' },
], ],
utterance: 'Tests passing. Ready to merge.', utterance: 'Tests passing. Ready to merge.',
}, },
@@ -35,7 +35,7 @@ const SCENARIOS: Scenario[] = [
{ prefix: '$', text: 'cursor agent:deploy', tone: 'accent' }, { prefix: '$', text: 'cursor agent:deploy', tone: 'accent' },
{ prefix: '✓', text: 'Migration applied (4 tables)', tone: 'success' }, { prefix: '✓', text: 'Migration applied (4 tables)', tone: 'success' },
{ prefix: '✓', text: 'Deploy complete', tone: 'success' }, { prefix: '✓', text: 'Deploy complete', tone: 'success' },
{ prefix: '→', text: 'voicebox.speak({ profile: "Scarlett" })', tone: 'dim' }, { prefix: '→', text: 'voicebox_speak({ profile: "Scarlett" })', tone: 'dim' },
], ],
utterance: 'Deploy shipped. Prod is green.', utterance: 'Deploy shipped. Prod is green.',
}, },
@@ -46,7 +46,7 @@ const SCENARIOS: Scenario[] = [
log: [ log: [
{ prefix: '$', text: 'cline task:review', tone: 'accent' }, { prefix: '$', text: 'cline task:review', tone: 'accent' },
{ prefix: '!', text: '3 files need attention', tone: 'dim' }, { prefix: '!', text: '3 files need attention', tone: 'dim' },
{ prefix: '→', text: 'voicebox.speak({ profile: "Jarvis" })', tone: 'dim' }, { prefix: '→', text: 'voicebox_speak({ profile: "Jarvis" })', tone: 'dim' },
], ],
utterance: 'Review ready. Three files to look at.', utterance: 'Review ready. Three files to look at.',
}, },
@@ -202,7 +202,7 @@ const MCP_CONFIG = `{
}`; }`;
const SPEAK_EXAMPLE = `// In any MCP-aware agent: const SPEAK_EXAMPLE = `// In any MCP-aware agent:
await voicebox.speak({ await voicebox_speak({
text: "Deploy complete.", text: "Deploy complete.",
profile: "Morgan", profile: "Morgan",
})`; })`;
@@ -300,7 +300,7 @@ export function AgentIntegration() {
</h2> </h2>
<p className="text-muted-foreground text-base md:text-lg leading-relaxed"> <p className="text-muted-foreground text-base md:text-lg leading-relaxed">
One tool call —{' '} One tool call —{' '}
<code className="text-accent font-mono text-[0.9em]">voicebox.speak</code> — <code className="text-accent font-mono text-[0.9em]">voicebox_speak</code> —
and any MCP-aware agent can talk to you in a voice you&rsquo;ve cloned. Claude Code, and any MCP-aware agent can talk to you in a voice you&rsquo;ve cloned. Claude Code,
Cursor, Cline, or anything that speaks MCP. Cursor, Cline, or anything that speaks MCP.
</p> </p>
+1 -1
View File
@@ -152,7 +152,7 @@ const CAPTURES: Capture[] = [
transcriptRaw: transcriptRaw:
"draft an update for the blog about the agent voice feature the key point is one MCP tool call and any agent on your machine gets a voice claude code finishes a long task calls voicebox dot speak and you hear it in a voice you've cloned morgan scarlett whatever you set up same pill that shows when you're dictating also shows when an agent is speaking so you always know what's coming out of your machine closes the whole voice IO loop for agents", "draft an update for the blog about the agent voice feature the key point is one MCP tool call and any agent on your machine gets a voice claude code finishes a long task calls voicebox dot speak and you hear it in a voice you've cloned morgan scarlett whatever you set up same pill that shows when you're dictating also shows when an agent is speaking so you always know what's coming out of your machine closes the whole voice IO loop for agents",
transcriptRefined: transcriptRefined:
"Draft an update for the blog about the agent voice feature. The key point: one MCP tool call, and any agent on your machine gets a voice. Claude Code finishes a long task, calls voicebox.speak, and you hear it in a voice you've cloned — Morgan, Scarlett, whatever you've set up. The same pill that shows when you're dictating also shows when an agent is speaking, so you always know what's coming out of your machine. It closes the full voice I/O loop for agents.", "Draft an update for the blog about the agent voice feature. The key point: one MCP tool call, and any agent on your machine gets a voice. Claude Code finishes a long task, calls voicebox_speak, and you hear it in a voice you've cloned — Morgan, Scarlett, whatever you've set up. The same pill that shows when you're dictating also shows when an agent is speaking, so you always know what's coming out of your machine. It closes the full voice I/O loop for agents.",
durationMs: 41000, durationMs: 41000,
ago: '22 min ago', ago: '22 min ago',
createdAtLabel: 'Apr 22, 3:29 PM', createdAtLabel: 'Apr 22, 3:29 PM',
+3
View File
@@ -45,5 +45,8 @@
"dependencies": { "dependencies": {
"loaders.css": "^0.1.2", "loaders.css": "^0.1.2",
"react-loaders": "^3.0.1" "react-loaders": "^3.0.1"
},
"overrides": {
"seroval": "1.5.3"
} }
} }
+1 -1
View File
@@ -1434,7 +1434,7 @@ pub fn run() {
} }
}); });
// Agent-initiated speech (voicebox.speak over MCP or POST /speak) // Agent-initiated speech (voicebox_speak over MCP or POST /speak)
// pops the pill up so the user can see what's coming out of their // pops the pill up so the user can see what's coming out of their
// machine. The `dictate:show` listener is kept for any frontend // machine. The `dictate:show` listener is kept for any frontend
// caller that wants to force-surface the pill directly, but the // caller that wants to force-surface the pill directly, but the