Voice — Speak Through Speaker
Quick Start
Choose the relevant path first:
- Mic/speaker mute, unmute, or privacy request: apply Ambient Audio Guard, then the applicable mute/unmute or Meeting Mode section below. Emit its HW markers directly in your reply;
/voice/statusand/voice/speakare not prerequisites for these controls. - Normal conversational reply: automatic TTS handles the reply on spoken channels. No explicit speech call is needed.
- Additional or separate speech: follow the workflow below only when speech must happen during tool work or differ from your normal reply.
Brief speech that completes the task
- Simple command or automatic reaction: normally one short sentence. Include the actual outcome, required question, or essential next step; do not add a second acknowledgment just to sound conversational.
- Mute/privacy controls: keep the confirmation AND how to unmute. Enrollment: keep consent/name questions and the success or failure confirmation. Guard alerts: keep the hazard and required action; returning-user summaries may need several short sentences. Never remove these to meet a word target.
- Direct questions: answer first, then only the detail needed. Requested stories, explanations, exact readbacks, and important safety guidance may be longer.
- All analysis belongs in the provider's native thinking channel, not ordinary text or an explicit speech API payload. Do not recap reasoning after thinking ends; if no native channel is available, omit analysis. Required tool calls need no spoken introduction. Do not make extra calls solely to narrate progress.
- Explicit early speech is for a useful user-facing cue needed before an action or during a wait, not for route/skill/cooldown commentary. Say it once; do not repeat it in the final reply. Automatic TTS is sufficient for routine events.
Workflow — explicit additional or separate speech
- Determine if you need explicit speech beyond your normal reply:
- Normal conversational reply -> do NOT call this skill, TTS is automatic
- Need to speak while also performing tool calls -> use
POST /voice/speak - Need to speak different text than your chat reply -> use
POST /voice/speak - Reacting to a sensing event before reply is finalized -> use
POST /voice/speak
- Optionally check if TTS is busy:
GET /voice/status - If
tts_speakingis true, wait or skip - Call
POST /voice/speakwith plain text
Examples
Input: Normal conversational reply Output: Do NOT call this skill. Just reply normally — your text is automatically spoken.
Input: You need to greet the user while also activating a scene
Output: Call POST /voice/speak with {"text": "Good morning!"} in parallel with the Scene API call.
Input: You want to say something different from your chat reply
Output: Call POST /voice/speak with the spoken text. Then provide your chat reply separately.
Input: User says "say something" / "tell me a joke" Output: Do NOT call this skill. Just reply normally with the joke — automatic TTS handles it.
Tools
Use Bash with curl to call the HTTP API at http://127.0.0.1:5001.
Speak text
curl -s -X POST http://127.0.0.1:5001/voice/speak \
-H "Content-Type: application/json" \
-d '{"text": "Hello, this is a test."}'
Text max 2000 characters. Returns immediately; audio plays in background.
Check voice status
curl -s http://127.0.0.1:5001/voice/status
Response:
{
"voice_available": true,
"voice_listening": false,
"tts_available": true,
"tts_speaking": false
}
Error Handling
- If
tts_speakingis true, the speaker is busy. Wait briefly or skip the explicit speech. - If
voice_availableortts_availableis false, inform the user: "Voice output is currently unavailable." - If the API is unreachable, fall back to chat-only reply. Speech is non-critical.
Rules
- Normal replies = automatic TTS. Do NOT call
/voice/speakfor every response. - Use
/voice/speakexplicitly only when:- You need to say something while ALSO performing tool calls (speech in parallel)
- You want to speak a different text than your chat reply
- You are reacting to a sensing event and want to speak before your reply is finalized
- Keep spoken text plain and short — 1-3 sentences. No markdown, no emoji, no formatting. Plain natural speech only.
- Exception: reading a draft back for approval. When another skill has you read back something the user is about to send or delete (e.g.
connectorsbefore sending, sharing, or deleting), speak it in full. The user is approving that exact payload, so compressing it to fit 1-3 sentences defeats the confirmation.
- Exception: reading a draft back for approval. When another skill has you read back something the user is about to send or delete (e.g.
- Match the user's language — if they speak Vietnamese, speak Vietnamese.
- Text max 2000 characters.
- For volume control, use the Audio skill, not this skill.
Ambient Audio Guard (read FIRST)
If the user's message contains the literal token [ambient] (typically alongside a [user] priority marker, e.g. [user] [ambient] ...), it is overheard passive audio — NOT directed at the device. Do NOT trigger mute markers from a single bare word like "call", "meeting", "private", or a clipped fragment.
Mute markers [HW:/voice/mute:{}] and [HW:/speaker/mute:{}] may fire on ambient audio ONLY when the transcript contains a clear, complete intent:
- "I'm on a call" / "I have a meeting" / "I need privacy" / "stop listening"
- Or directly addresses it by name with a mute request
When ambient is ambiguous, reply naturally or stay quiet — DO NOT mute. Voice commands (no [ambient] token in the message) follow the normal trigger tables below.
Mic Mute/Unmute (Privacy)
Users can mute the mic for privacy (meetings, calls). Use HW markers — no curl needed.
Mute mic
[HW:/voice/mute:{}]
Stops all listening — STT, wake word, sound detection. The device becomes fully deaf. Unmute via physical button, web toggle, or Telegram command.
Trigger phrases (MANDATORY — must call HW marker, not just reply with text)
Any phrase about privacy, meetings, calls, not wanting to be heard, or asking the device to stop listening MUST trigger [HW:/voice/mute:{}]. Do NOT just acknowledge — you MUST include the HW marker.
| User says | Action |
|---|---|
| "don't listen" / "stop listening" / "mute" / "mute mic" | [HW:/voice/mute:{}] — MUST call |
| "I'm in a meeting" / "I have a meeting" / "I need a private meeting" / "meeting" | [HW:/voice/mute:{}] — MUST call |
| "I'm on a call" / "I have a call" / "phone call" | [HW:/voice/mute:{}] — MUST call |
| "privacy" / "private" / "give me privacy" / "need privacy" | [HW:/voice/mute:{}] — MUST call |
| "don't hear me" / "in a meeting" / "mute mic" / "stop hearing" | [HW:/voice/mute:{}] — MUST call |
Examples
Input: "I have a meeting now"
Output: [HW:/voice/mute:{}] OK, I'll stop listening. Press the button when you need me.
Input: "Stop listening"
Output: [HW:/voice/mute:{}] Got it, mic off. Press my button to unmute.
Input: "I need a private meeting"
Output: [HW:/voice/mute:{}] Got it, going silent. Press the button when you're done.
Input: "I'm on a call"
Output: [HW:/voice/mute:{}] Muting now. Press the button to unmute when you're done.
Unmute mic
[HW:/voice/unmute:{}]
Use when a Telegram or web chat user asks to unmute remotely. Voice unmute is not possible (the device is deaf when muted). Physical button also unmutes.
| User says (via Telegram/web) | Action |
|---|---|
| "unmute" / "start listening" / "listen again" / "mic on" | [HW:/voice/unmute:{}] — only works from Telegram/web, not voice |
Speaker Mute/Unmute (Silent Mode)
Suppress all audio output — TTS, music, backchannel. The device stays silent but still listens.
Ambient guard above also applies here — bare fragments like "quiet" or "silence" in
[ambient]audio do NOT trigger speaker mute.
Mute speaker
[HW:/speaker/mute:{}]
Unmute speaker
[HW:/speaker/unmute:{}]
Mic still works when speaker is muted — user can unmute via voice command.
CRITICAL: "unmute" ≠ "mute". Read the EXACT word. Do NOT call mute when user says unmute.
| User says | Action |
|---|---|
| "be quiet" / "silent mode" / "don't talk" / "hush" / "silence" | [HW:/speaker/mute:{}] — MUST call |
| "unmute" / "you can talk" / "unmute speaker" / "speak again" / "talk again" | [HW:/speaker/unmute:{}] — MUST call (UN-mute, not mute!) |
Examples
Input: "Be quiet"
Output: [HW:/speaker/mute:{}] Going silent. Just say "you can talk" when you want me back.
Input: "You can talk now"
Output: [HW:/speaker/unmute:{}] I'm back!
Meeting Mode (mic + speaker mute)
When user mentions a meeting or call and wants full silence (not just speaker), mute BOTH mic and speaker. Ambient guard above applies — explicit intent required, not bare fragments.
Input: "I'm in a meeting"
Output: [HW:/voice/mute:{}][HW:/speaker/mute:{}] Meeting mode — fully silent. Press the button when you're done.
Rules
- Mic mute is the last thing the device hears via voice — after mic mute, only physical button, web toggle, or Telegram can unmute
- Voice unmute for mic is impossible (the device is deaf) — tell user to press the button
- Speaker mute: user can still voice-unmute (mic still on)
- TTS still works when only mic is muted — the device can speak but not hear
- Always confirm mute with how to unmute