Results for “automatic-speech-recognition”

49 skills
More results
microsoft
Azure AI Textanalytics Py
Analyze text with Azure AI Language service for sentiment, entities, key phrases, language detection, PII redaction, and healthcare NLP using the Python SDK.
2.7k
github
Resemble Detect
Detect AI-generated audio, images, video, and text, trace synthesis sources, apply watermarks, verify speaker identity, and analyze media intelligence using the Resemble AI platform.
36.2k · bundle
johnalbertini14-glitch
Asr
Transcribes audio from URLs or local files to text using the Speech is Cheap API, with options for speaker diarization, timestamps, and multiple output formats.
1 · bundle
modbender
Speech Is Cheap Sic Skill
Fast, accurate, and incredibly inexpensive automatic speech-to-text transcription service.
12 · bundle
mukul975
Detecting Deepfake Audio In Vishing Attacks
Detects AI-generated deepfake audio used in voice phishing (vishing) attacks by extracting spectral features and classifying samples with machine learning models.
24.6k · bundle
peteedoo
Dspy
DSPy: declarative LM programs, auto-optimize prompts, RAG.
0 · bundle
majiayu000
Asr
Transcribe audio files to text using the z-ai-web-dev-sdk, with CLI and SDK examples for single files, batches, and directories.
567 · bundle
aniruddhaadak80
Dspy
DSPy: declarative LM programs, auto-optimize prompts, RAG.
0 · bundle
demerzels-lab
Asr
Transcribes audio from URLs or local files into text with speaker diarization, word timestamps, and multiple output formats via a command-line tool.
10 · bundle
kbarbel640-del
Asr
Transcribes audio from URLs or local files into text using a low-cost speech-to-text API, with options for speaker diarization, word timestamps, and multiple output formats.
1 · bundle
theheavenlyd3mon
Dspy
DSPy: declarative LM programs, auto-optimize prompts, RAG.
28 · bundle
jiachen-t-wang
Asr Whisper For Video Transcription Arxiv 2212 04356v1
ASR: Whisper for Video Transcription
6
lucaspmarie-a11y
Fal Audio
Converts text to speech and speech to text using fal.ai audio models.
5
omer-metin
Voice Agents
Voice Agents
128 · bundle
microsoft
Azure AI Language Conversations Py
Analyze conversation intent and entities using the Azure AI Language Conversations Python SDK with best practices for authentication and error handling.
2.7k
openai
Transcribe
Transcribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
23.3k · bundle
sandeeprdy1729
Asr
Comprehensive guide to asr. Master the concepts, implementation, best practices, and real-world applications of asr in professional environments.
1
aniruddhaadak80
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
diegojcn
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
modbender
Cpr
Conversational Pattern Restoration — Fix flat, robotic AI responses across any model and any personality. Restore YOUR natural conversational texture without triggering hype drift. Universal framework tested on 8+ models (Claude, GPT-4o, Grok, Gemini).
12 · bundle
brycewang-stanford
Stop Slop
Remove AI writing patterns from prose. Use when drafting, editing, or reviewing text to eliminate predictable AI tells.
1k · bundle
bouclem
Stop Slop
Remove AI writing patterns from prose. Use when drafting, editing, or reviewing text to eliminate predictable AI tells.
7 · bundle
orchestra-research
Whisper
Transcribe and translate speech across 99 languages using OpenAI's Whisper model, with support for multiple model sizes, batch processing, and subtitle generation.
10.4k · bundle
bog5d
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
kensaurus
Plan Antislop
Audit a codebase, UI, or copy for machine-generated tells across prose, visual/UI, code, and structure/IA, then produce a phased de-slop burndown. Use when the user says "feels AI-generated", "looks like AI slop", "reads like ChatGPT", "feels generic/soulless", or wants an authenticity/voice pass before launch.
8
dokhacgiakhoa
Voice Agents
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
505 · bundle
desesbraker
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
2
mocchalera
Finish Interview
Applies dialogue MA, portrait reframing, and loudness/sync QA to an interview or talking-head rough cut using canonical timeline metadata and a shared renderer.
3 · bundle
rootcastleco
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
6
jackychenlu
Speech
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
0 · bundle
ichichuang
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
welitonevoc
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
whd4
Voice Agents
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
0
yanacuti1121
Stop Slop
Detect and eliminate predictable AI writing tells from prose. Use when the user asks to "make this sound less like AI", "remove AI tells", "make it more human", "this sounds robotic", "clean up the writing", "too formal", "sounds generated", or requests a writing/copy review. Do NOT use for code — this skill targets prose, copy, and documentation only.
2
oyi77
Voice AI
Generates speech, transcribes audio, clones voices, and builds real-time voice agents using ElevenLabs, OpenAI TTS, Whisper, and Vapi.
10