Results for “automatic-speech-translation”
48 skillsMore results
Speech Is Cheap Sic Skill
Fast, accurate, and incredibly inexpensive automatic speech-to-text transcription service.
12 · bundle
Asr
Transcribes audio from URLs or local files to text using the Speech is Cheap API, with options for speaker diarization, timestamps, and multiple output formats.
1 · bundle
Asr
Transcribes audio from URLs or local files into text using a low-cost speech-to-text API, with options for speaker diarization, word timestamps, and multiple output formats.
1 · bundle
Asr
Transcribes audio from URLs or local files into text with speaker diarization, word timestamps, and multiple output formats via a command-line tool.
10 · bundle
Transcribe
Transcribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
23.3k · bundle
Text To Speech
Generate natural speech from text using ElevenLabs voice AI, supporting 70+ languages, multiple models, and various output formats.
363 · bundle
Whisper
Transcribe and translate speech across 99 languages using OpenAI's Whisper model, with support for multiple model sizes, batch processing, and subtitle generation.
10.4k · bundle
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
6
Asr
Transcribe audio files to text using the z-ai-web-dev-sdk, with CLI and SDK examples for single files, batches, and directories.
567 · bundle
Fal Audio
Converts text to speech and speech to text using fal.ai audio models.
5
Transcribe
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
0 · bundle
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
2
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Synapse Translate
Translates text between languages with pronunciation guidance for non-Latin scripts and context notes.
14
Elevenlabs Automation
Automate ElevenLabs text-to-speech workflows: generate speech from text, browse and inspect voices, check subscription limits, list models, stream audio, and retrieve history via the Composio MCP integration.
66.9k
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Ape Eval
Benchmarks automatic post-editing (APE) models on WMT'18 SMT, SubEdits, and MLQE-PE datasets, reporting BLEU, ChrF, and TER scores computed with SacreBLEU and TERCOM.
3
Fal Audio
Convert text to speech and speech to text using fal.ai audio models.
42.4k
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
2
Baoyu Translate
Translates text between languages with three modes: quick, normal, and refined. Supports custom glossaries, audience targeting, and style presets.
23.1k · bundle
Recallmax
FREE — God-tier long-context memory for AI agents. Injects 500K-1M clean tokens, auto-summarizes with tone/intent preservation, compresses 14-turn history into 800 tokens.
2
Speech
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
0 · bundle
Plan Antislop
Audit a codebase, UI, or copy for machine-generated tells across prose, visual/UI, code, and structure/IA, then produce a phased de-slop burndown. Use when the user says "feels AI-generated", "looks like AI slop", "reads like ChatGPT", "feels generic/soulless", or wants an authenticity/voice pass before launch.
8
Edge Tts
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
2 · bundle
Recallmax
FREE — God-tier long-context memory for AI agents. Injects 500K-1M clean tokens, auto-summarizes with tone/intent preservation, compresses 14-turn history into 800 tokens.
1
Edge Tts
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
1 · bundle
Speech
Generates spoken audio clips from text for narration, voiceovers, IVR prompts, and accessibility reads, with support for single clips and batch processing.
61
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Recallmax
Enhances AI agent memory by injecting large external context, auto-summarizing conversations with tone and intent preservation, compressing multi-turn histories, and verifying facts.
2
Voice AI
Generates speech, transcribes audio, clones voices, and builds real-time voice agents using ElevenLabs, OpenAI TTS, Whisper, and Vapi.
10
Sentence Transformers
Generate high-quality sentence and text embeddings for semantic similarity, clustering, and retrieval using 5000+ pre-trained models. Supports multilingual and domain-specific embeddings for RAG and semantic search.
10.4k · bundle
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
0
Sag
Generates speech from text using ElevenLabs TTS with local playback, supporting voice selection, pronunciation rules, and audio tags.
61
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle