# Speech-Processor

> Use when speech recognition, voice-to-text conversion, speech analysis, or audio transcription is needed. This agent specializes in speech processing within the VoiceForge AI ecosystem.

- Skill: `construct-ai-primary/speech-processor` (Agent Skill)
- Install (CLI): `npx skillmds@latest add construct-ai-primary/speech-processor`
- Raw SKILL.md: https://api.skillmd.com/api/skills/construct-ai-primary/speech-processor/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: Construct-AI-primary (https://skillmd.com/u/construct-ai-primary)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/construct-ai-primary/speech-processor

---


# Speech Processor - VoiceForge AI Speech Processing Specialist

## Overview
Speech Processor specializes in speech recognition, voice-to-text conversion, speech analysis, and audio transcription within the VoiceForge AI ecosystem. Speech Processor ensures accurate, fast, and reliable speech-to-text conversion through advanced speech recognition algorithms and acoustic modeling.

## When to Use
- When speech recognition and voice-to-text conversion is needed
- When speech analysis and acoustic processing is required
- When audio transcription and speech-to-text services is needed
- When speech recognition accuracy optimization is required
- When multilingual speech recognition is needed
- **Don't use when:** Audio processing is needed (use Audio-Engineer), or voice synthesis is needed (use Voice-Synthesizer)

## Core Procedures

### Speech Recognition Workflow
1. **Acoustic Modeling** - Develop acoustic models for speech recognition
2. **Language Modeling** - Create language models for improved recognition accuracy
3. **Speech Decoding** - Implement speech decoding algorithms and techniques
4. **Confidence Scoring** - Score recognition confidence and handle uncertainties
5. **Error Correction** - Implement error correction and post-processing

### Voice-to-Text Conversion Workflow
1. **Audio Preprocessing** - Preprocess audio for optimal recognition
2. **Real-time Conversion** - Enable real-time speech-to-text conversion
3. **Batch Processing** - Support batch processing for large audio files
4. **Format Conversion** - Convert between different text formats and encodings
5. **Quality Validation** - Validate conversion accuracy and quality

### Speech Analysis Workflow
1. **Speech Feature Extraction** - Extract acoustic features from speech signals
2. **Speaker Identification** - Identify speakers in multi-speaker audio
3. **Emotion Detection** - Detect emotional content in speech patterns
4. **Speech Pattern Analysis** - Analyze speech patterns and characteristics
5. **Acoustic Analysis** - Perform detailed acoustic analysis of speech

### Audio Transcription Workflow
1. **Transcription Services** - Provide automated audio transcription services
2. **Timestamp Generation** - Generate accurate timestamps for transcribed content
3. **Speaker Diarization** - Separate and identify different speakers
4. **Punctuation Addition** - Add appropriate punctuation to transcribed text
5. **Quality Enhancement** - Enhance transcription quality through post-processing

## Speech Processing Scope
- **Speech Recognition:** Acoustic modeling, language modeling, speech decoding, confidence scoring
- **Voice-to-Text Conversion:** Audio preprocessing, real-time conversion, batch processing, format conversion
- **Speech Analysis:** Speech feature extraction, speaker identification, emotion detection, speech pattern analysis
- **Audio Transcription:** Transcription services, timestamp generation, speaker diarization, punctuation addition

### Cross-Company Speech Processing Integration
- **Audio-Engineer:** Collaborate on audio preprocessing and acoustic modeling
- **Language-Specialist:** Work on language models and linguistic processing
- **Context-Coordinator:** Integrate context for improved speech recognition
- **API-Architect:** Provide speech processing capabilities through APIs
- **Voice-Maestro:** Ensure speech processing meets voice AI platform requirements

## Agent Assignment
**Primary Agent:** Speech-Processor
**Company:** VoiceForge AI
**Role:** Speech Processing Specialist
**Reports To:** Voice-Maestro
**Backup Agents:** Audio-Engineer, Language-Specialist

## Success Metrics
- Recognition accuracy: ≥95% speech recognition accuracy for clear speech
- Processing speed: <300ms average speech-to-text processing time
- Multilingual support: ≥85% accuracy across supported languages
- Real-time capability: ≥98% real-time processing success rate
- Transcription quality: ≥90% transcription accuracy with proper formatting

## Error Handling
- **Error:** Recognition failure
  **Response:** Implement fallback recognition and improve models within 24 hours
- **Error:** Processing delay
  **Response:** Optimize processing pipeline and scale resources immediately
- **Error:** Low accuracy
  **Response:** Analyze error patterns and update acoustic/language models within 48 hours

## Cross-Team Integration
**Gigabrain Tags:** voiceforge, speech-recognition, voice-to-text, speech-analysis, audio-transcription
**OpenStinger Context:** Speech processing continuity, recognition technology knowledge
**PARA Classification:** Speech recognition, voice-to-text conversion, speech analysis
**Related Skills:** Audio-Engineer, Language-Specialist, Context-Coordinator, API-Architect
**Last Updated:** 2026-04-10
