Vox — Claude Code Skill
Project Overview
Vox is a lightweight voice MCP server (~1,500 lines of Rust) providing local text-to-speech (Kokoro) and speech-to-text (Moonshine Base) via the MCP protocol. It runs as a stdio subprocess per MCP client, or as a shared HTTP daemon.
Build & Test
cargo check # type-check only
cargo test # run all unit tests (84 tests across 10 modules)
cargo clippy -- -D warnings # lint — must pass with zero warnings
cargo build --release # optimized build (LTO + single codegen unit)
cargo bench -- resample # benchmark resampling
All of cargo test, cargo clippy -- -D warnings, and cargo fmt --check must pass before submitting changes.
Architecture
| Module |
Purpose |
main.rs |
Entry point, config loading, model download, stdio/daemon startup |
cli.rs |
Clap CLI parser: daemon, config, download-models subcommands |
server.rs |
MCP tool handlers (say, listen, converse), streaming TTS pipeline |
tts.rs |
Kokoro TTS engine wrapper, voice name → speaker ID resolution, sentence splitting |
audio.rs |
cpal-based mic capture and speaker playback, Lanczos-3 sinc resampling |
stt.rs |
Moonshine Base STT engine wrapper |
vad.rs |
Voice activity detection (Silero ONNX) |
config.rs |
TOML config loading, env var overrides (VOX_* prefix), path resolution |
daemon.rs |
HTTP daemon lifecycle: daemonize, PID file, start/stop/status/log |
models.rs |
Model readiness checks and download/extraction |
lib.rs |
Public re-exports for benchmarks (audio, config, error, tts) |
error.rs |
VoiceError enum with thiserror derives |
Transport Modes
- Stdio (default):
rmcp::transport::stdio(). One process per MCP client.
- Daemon (
vox daemon start [--port PORT]): StreamableHttpService via rmcp. Single process, models loaded once, multiple clients connect over HTTP/SSE. Factory closure creates a VoiceMcpServer per session with shared Arc<Mutex<TtsEngine>> and Arc<Mutex<SttEngine>>.
Config System
Precedence (highest wins):
- Environment variables:
VOX_SPEED, VOX_VOICE, VOX_MODEL_DIR, VOX_LOG_LEVEL, VOX_PORT
- TOML file:
$XDG_CONFIG_HOME/vox/config.toml
- Compiled defaults (
Config::default())
CLI management: vox config get [key], vox config set <key> <value>, vox config path
MCP Tools
| Tool |
Description |
say |
Speak text aloud through speakers (TTS only) |
listen |
Record from microphone and transcribe (STT only) |
converse |
Speak text then listen for response (TTS + STT round-trip) |
Available Voices
American female (af_*): heart, alloy, aoede, bella, jessica, kore, nicole, nova, river, sarah, sky
American male (am_*): adam, echo, eric, liam, michael, onyx, puck, santa
British female (bf_*): alice, emma, lily
British male (bm_*): daniel, fable, george, lewis
Default: af_heart (ID 0). Voices can be specified by name or numeric ID.
Code Conventions
- Edition 2024 — uses
let chains (if let Ok(x) = ... && let Ok(y) = ...)
- Visibility:
pub(crate) for test-only exposure, not fully pub
- Clippy: treat all warnings as errors (
-D warnings)
- Tests: inline
#[cfg(test)] mod tests per module, tempfile for filesystem tests
unsafe impl Send: TtsEngine and CaptureHandle have manual Send impls due to non-Send cpal/sherpa internals confined to dedicated threads
Dev Workflow
- Make changes
cargo test — verify all 84 tests pass
cargo clippy -- -D warnings — zero warnings
cargo fmt --check — formatting
- If touching
audio.rs resampling: cargo bench -- resample
1---2name: vox3description: Runs a local voice MCP server in Rust for text-to-speech and speech-to-text, with build, test, and configuration guidance.4---56# Vox — Claude Code Skill78## Project Overview910Vox is a lightweight voice MCP server (~1,500 lines of Rust) providing local text-to-speech (Kokoro) and speech-to-text (Moonshine Base) via the MCP protocol. It runs as a stdio subprocess per MCP client, or as a shared HTTP daemon.1112## Build & Test1314```bash15cargo check # type-check only16cargo test # run all unit tests (84 tests across 10 modules)17cargo clippy -- -D warnings # lint — must pass with zero warnings18cargo build --release # optimized build (LTO + single codegen unit)19cargo bench -- resample # benchmark resampling20```2122All of `cargo test`, `cargo clippy -- -D warnings`, and `cargo fmt --check` must pass before submitting changes.2324## Architecture2526| Module | Purpose |27|--------|---------|28| `main.rs` | Entry point, config loading, model download, stdio/daemon startup |29| `cli.rs` | Clap CLI parser: daemon, config, download-models subcommands |30| `server.rs` | MCP tool handlers (`say`, `listen`, `converse`), streaming TTS pipeline |31| `tts.rs` | Kokoro TTS engine wrapper, voice name → speaker ID resolution, sentence splitting |32| `audio.rs` | cpal-based mic capture and speaker playback, Lanczos-3 sinc resampling |33| `stt.rs` | Moonshine Base STT engine wrapper |34| `vad.rs` | Voice activity detection (Silero ONNX) |35| `config.rs` | TOML config loading, env var overrides (`VOX_*` prefix), path resolution |36| `daemon.rs` | HTTP daemon lifecycle: daemonize, PID file, start/stop/status/log |37| `models.rs` | Model readiness checks and download/extraction |38| `lib.rs` | Public re-exports for benchmarks (`audio`, `config`, `error`, `tts`) |39| `error.rs` | `VoiceError` enum with `thiserror` derives |4041## Transport Modes4243- **Stdio** (default): `rmcp::transport::stdio()`. One process per MCP client.44- **Daemon** (`vox daemon start [--port PORT]`): `StreamableHttpService` via rmcp. Single process, models loaded once, multiple clients connect over HTTP/SSE. Factory closure creates a `VoiceMcpServer` per session with shared `Arc<Mutex<TtsEngine>>` and `Arc<Mutex<SttEngine>>`.4546## Config System4748Precedence (highest wins):491. Environment variables: `VOX_SPEED`, `VOX_VOICE`, `VOX_MODEL_DIR`, `VOX_LOG_LEVEL`, `VOX_PORT`502. TOML file: `$XDG_CONFIG_HOME/vox/config.toml`513. Compiled defaults (`Config::default()`)5253CLI management: `vox config get [key]`, `vox config set <key> <value>`, `vox config path`5455## MCP Tools5657| Tool | Description |58|------|-------------|59| `say` | Speak text aloud through speakers (TTS only) |60| `listen` | Record from microphone and transcribe (STT only) |61| `converse` | Speak text then listen for response (TTS + STT round-trip) |6263## Available Voices6465American female (`af_*`): heart, alloy, aoede, bella, jessica, kore, nicole, nova, river, sarah, sky66American male (`am_*`): adam, echo, eric, liam, michael, onyx, puck, santa67British female (`bf_*`): alice, emma, lily68British male (`bm_*`): daniel, fable, george, lewis6970Default: `af_heart` (ID 0). Voices can be specified by name or numeric ID.7172## Code Conventions7374- **Edition 2024** — uses `let` chains (`if let Ok(x) = ... && let Ok(y) = ...`)75- **Visibility**: `pub(crate)` for test-only exposure, not fully `pub`76- **Clippy**: treat all warnings as errors (`-D warnings`)77- **Tests**: inline `#[cfg(test)] mod tests` per module, `tempfile` for filesystem tests78- **`unsafe impl Send`**: `TtsEngine` and `CaptureHandle` have manual `Send` impls due to non-Send cpal/sherpa internals confined to dedicated threads7980## Dev Workflow81821. Make changes832. `cargo test` — verify all 84 tests pass843. `cargo clippy -- -D warnings` — zero warnings854. `cargo fmt --check` — formatting865. If touching `audio.rs` resampling: `cargo bench -- resample`