Audio Processing Skill
A comprehensive toolset for audio manipulation and analysis with security validations.
Security
- File paths are validated to prevent path traversal attacks
- Access to system directories (/etc, /proc, /sys, /root) is blocked
- TTS text input is limited to 10,000 characters
- All file operations use resolved absolute paths
Tool API
audio_tool
Perform audio operations like transcription, text-to-speech, and feature extraction.
- Parameters:
action (string, required): One of transcribe, tts, extract_features, vad_segments, transform.
file_path (string, optional): Path to input audio file.
text (string, optional): Text for TTS (max 10,000 chars).
output_path (string, optional): Path for output file (default: auto-generated).
model (string, optional): Whisper model size (tiny, base, small, medium, large). Default: base.
ops (string, optional): JSON string of operations for transform action.
Usage:
# Transcribe audio file
uv run --with "openai-whisper" --with "pydub" --with "numpy" skills/audio-processing/tool.py transcribe --file_path input.wav
# Transcribe with specific model
uv run --with "openai-whisper" skills/audio-processing/tool.py transcribe --file_path input.wav --model small
# Text-to-speech
uv run --with "gTTS" skills/audio-processing/tool.py tts --text "Hello world" --output_path hello.mp3
# Extract audio features
uv run --with "librosa" --with "numpy" --with "soundfile" skills/audio-processing/tool.py extract_features --file_path input.wav
# Voice activity detection (find speech segments)
uv run --with "pydub" skills/audio-processing/tool.py vad_segments --file_path input.wav
# Transform audio (trim, resample, normalize)
uv run --with "pydub" skills/audio-processing/tool.py transform --file_path input.wav --ops '[{"op": "trim", "start": 10, "end": 30}, {"op": "normalize"}]'
Actions
transcribe
Convert speech to text using OpenAI Whisper.
- Returns:
{ "text": "...", "segments": [...] }
- Models: tiny, base, small, medium, large (larger = more accurate, slower)
tts
Generate speech from text using Google TTS.
- Returns:
{ "file_path": "output.mp3", "status": "created" }
- Language: English (default)
extract_features
Extract audio features for analysis.
- Returns: duration, sample_rate, mfcc_mean, rms_mean
- Useful for audio classification, quality analysis
vad_segments
Detect speech segments using silence detection.
- Returns:
{ "segments": [{ "start": 0.5, "end": 3.2 }, ...] }
- Uses FFmpeg silencedetect filter
- Aggressiveness: 1-3 (default: 2)
transform
Apply transformations to audio files.
- Operations: trim, resample, normalize
- Returns:
{ "file_path": "output.wav" }
Requirements
- ffmpeg: Required for VAD and transform operations
- Python 3.8+: All operations
- Disk Space: Whisper models range from 100MB (tiny) to 3GB (large)
Error Handling
- Returns JSON error object on failure
- Validates all file paths before processing
- Gracefully handles missing dependencies
1---2name: audio-processing3description: Audio ingestion, analysis, transformation, and generation (Transcribe, TTS, VAD, Features).4---5
6# Audio Processing Skill
7
8A comprehensive toolset for audio manipulation and analysis with security validations.
9
10## Security
11
12- File paths are validated to prevent path traversal attacks
13- Access to system directories (/etc, /proc, /sys, /root) is blocked
14- TTS text input is limited to 10,000 characters
15- All file operations use resolved absolute paths
16
17## Tool API
18
19### audio_tool
20Perform audio operations like transcription, text-to-speech, and feature extraction.
21
22- **Parameters:**
23 - `action` (string, required): One of `transcribe`, `tts`, `extract_features`, `vad_segments`, `transform`.
24 - `file_path` (string, optional): Path to input audio file.
25 - `text` (string, optional): Text for TTS (max 10,000 chars).
26 - `output_path` (string, optional): Path for output file (default: auto-generated).
27 - `model` (string, optional): Whisper model size (tiny, base, small, medium, large). Default: `base`.
28 - `ops` (string, optional): JSON string of operations for transform action.
29
30**Usage:**
31
32```bash
33# Transcribe audio file
34uv run --with "openai-whisper" --with "pydub" --with "numpy" skills/audio-processing/tool.py transcribe --file_path input.wav
35
36# Transcribe with specific model
37uv run --with "openai-whisper" skills/audio-processing/tool.py transcribe --file_path input.wav --model small
38
39# Text-to-speech
40uv run --with "gTTS" skills/audio-processing/tool.py tts --text "Hello world" --output_path hello.mp3
41
42# Extract audio features
43uv run --with "librosa" --with "numpy" --with "soundfile" skills/audio-processing/tool.py extract_features --file_path input.wav
44
45# Voice activity detection (find speech segments)
46uv run --with "pydub" skills/audio-processing/tool.py vad_segments --file_path input.wav
47
48# Transform audio (trim, resample, normalize)
49uv run --with "pydub" skills/audio-processing/tool.py transform --file_path input.wav --ops '[{"op": "trim", "start": 10, "end": 30}, {"op": "normalize"}]'
50```
51
52## Actions
53
54### transcribe
55Convert speech to text using OpenAI Whisper.
56
57- Returns: `{ "text": "...", "segments": [...] }`
58- Models: tiny, base, small, medium, large (larger = more accurate, slower)
59
60### tts
61Generate speech from text using Google TTS.
62
63- Returns: `{ "file_path": "output.mp3", "status": "created" }`
64- Language: English (default)
65
66### extract_features
67Extract audio features for analysis.
68
69- Returns: duration, sample_rate, mfcc_mean, rms_mean
70- Useful for audio classification, quality analysis
71
72### vad_segments
73Detect speech segments using silence detection.
74
75- Returns: `{ "segments": [{ "start": 0.5, "end": 3.2 }, ...] }`
76- Uses FFmpeg silencedetect filter
77- Aggressiveness: 1-3 (default: 2)
78
79### transform
80Apply transformations to audio files.
81
82- Operations: trim, resample, normalize
83- Returns: `{ "file_path": "output.wav" }`
84
85## Requirements
86
87- **ffmpeg:** Required for VAD and transform operations
88- **Python 3.8+:** All operations
89- **Disk Space:** Whisper models range from 100MB (tiny) to 3GB (large)
90
91## Error Handling
92
93- Returns JSON error object on failure
94- Validates all file paths before processing
95- Gracefully handles missing dependencies