# Tg Voice Whisper

> ---

- Skill: `johnalbertini14-glitch/tg-voice-whisper` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add johnalbertini14-glitch/tg-voice-whisper`
- Raw SKILL.md: https://api.skillmd.com/api/skills/johnalbertini14-glitch/tg-voice-whisper/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: johnalbertini14-glitch (https://skillmd.com/u/johnalbertini14-glitch)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/johnalbertini14-glitch/tg-voice-whisper

---

---
name: tg-voice-whisper
description: Auto-transcribe Telegram voice messages (.ogg Opus) to text using local OpenAI Whisper (tiny model). Reply with transcription + auto-delete file for privacy. No API keys, fully offline after install.
emoji: 🎙️🔊

requires:
  bins: [whisper, ffmpeg]

install:
  - kind: apt
    packages: [ffmpeg]
  - kind: pip
    packages: [openai-whisper]
    flags: [--break-system-packages]

min-server: 4GB RAM recommended for fast/stable work (cached model <1s transcribe on 1vCPU).

**Safe:** Local-only, auto-rm files, no external APIs/credentials stored.

**Usage (agent):**
When voice .ogg arrives in /root/.openclaw/media/inbound/:
```
exec 'whisper PATH --model tiny --language ru --output_format txt --output_dir /tmp/whisper; text=$(cat /tmp/whisper/*.txt); message action=send message="🔊 [text]" reply_to_current; rm PATH /tmp/whisper/*'
```

**Auto-setup (non-block):**
Spawn sub-agent or cron every 5s:
```
sessions_spawn task="LOOP: find /root/.openclaw/media/inbound/*.ogg -mmin -1 → background whisper ru tiny → message reply text → rm" label="voice-auto" cleanup="keep"
```
Or cron:
```
cron add job={"name":"voice-auto","schedule":{"kind":"every","everyMs":5000},"payload":{"kind":"systemEvent","text":"🔊 VOICE_CHECK"},"sessionTarget":"main"}
```

**Test:**
whisper /path.ogg --model tiny --language ru

**Notes:**
- First run: ~15s model download (72MB ~/.cache/whisper/tiny.bin).
- Cached: <1s on 1vCPU/4GB.
- Languages: ru/en best; --language detect auto.
- Accuracy: tiny 85-95% ru speech; upgrade to base/small for better.
