Packs

2 packs

Results for “audio”

26 skills
openai
speech
Generate spoken audio for narration, voiceovers, IVR prompts, and accessibility reads using the OpenAI Audio API with bundled CLI and built-in voices.
23.3k · bundle
remotion-dev
mediabunny
Extract audio and video metadata such as duration and dimensions using the Mediabunny library in the browser.
3.9k · bundle
ecnu-icalk
remotion-best-practices
Provides domain-specific guidance for building videos with Remotion, covering captions, FFmpeg operations, audio visualization, and a wide range of composition techniques.
559 · bundle
alirezarezvani
notebooklm
Browser-automates Google's NotebookLM to read notebooks, add sources, generate Studio outputs (Audio/Video Overviews, Mind Maps, Reports), and create new notebooks.
20.4k · bundle
sdiamante13
remotion-best-practices
Loads Remotion best practices for React video creation, covering animations, composition setup, assets, audio, captions, sequencing, transitions, media metadata, Three.js, Tailwind, and rendering patterns.
7 · bundle
brycewang-stanford
markitdown
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. Use when converting documents to markdown, extracting text from PDFs/Office files, transcribing audio, performing OCR on images, extracting YouTube transcripts, or processing batches of files. Supports 20+ formats including DOCX, XLSX, PPTX, PDF, HTML, EPUB, CSV, JSON, images with OCR, and audio with transcription.
1k · bundle
More results
microsoft
podcast-generation
Generate AI-powered podcast-style audio narratives from text using Azure OpenAI's GPT Realtime Mini model via WebSocket, with full-stack implementation from React frontend to Python FastAPI backend.
2.7k · bundle
k-dense-ai
markitdown
Convert files and office documents to Markdown using Microsoft's MarkItDown tool. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.
30.2k · bundle
inference-sh
ai-voice-cloning
Generate natural AI voices, text-to-speech, and voice synthesis using the inference.sh CLI with models like Inworld TTS, ElevenLabs, and Kokoro TTS for voiceovers, audiobooks, podcasts, and more.
584
jackychenlu
speech
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
0 · bundle
jarbitechture
speech
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
0 · bundle
seaworld008
speech
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
65 · bundle
metinduraktr-44
speech
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
0 · bundle
om-scogo
edge-tts
Text-to-speech conversion using `uvx edge-tts` for generating audio from text. Use when (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
0 · bundle
om-scogo
zai-tts
Text-to-speech conversion using GLM-TTS service via the `uvx zai-tts` command for generating audio from text. Use when (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, podcast, driving, cooking). (3) Using pre-cloned voices for speech.
0 · bundle
pymodel
opentui
Build terminal UIs with OpenTUI. Covers core, components, audio, keymaps, React, Solid, plugins, testing, standalone executables, QR encoding, SSH, and Three.js WebGPU.
14 · bundle
composiohq
youtube-downloader
Download YouTube videos with customizable quality and format options, including audio-only MP3 extraction.
66.9k · bundle
yanacuti1121
markitdown
Convert mọi file (PDF, Word, Excel, PowerPoint, Image, Audio, HTML, CSV, YouTube URL) thành Markdown cho LLM. Tích hợp vào YAMTAM vault để import tài liệu.
2
timlai666
edge-tts
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
1 · bundle
danstrem2
edge-tts
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
2 · bundle
chen-yu-hao
markitdown
Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.
5 · bundle
auto-skiller
remotion-video-creation
Provides domain-specific best practices for creating videos with Remotion in React, covering 3D, animations, audio, captions, charts, transitions, and more.
1 · bundle
inference-sh
elevenlabs-tts
Generate high-quality speech from text using ElevenLabs' premium voices, with support for 32 languages, multiple models, and voice tuning parameters.
584
remotion-dev
remotion-markup
Guidance for writing Remotion React markup with animation patterns, media handling, sequences, effects, and advanced features like 3D, audio visualization, and parameterized videos.
3.9k · bundle
omer-metin
voiceover
World-class voiceover expertise combining the narrative craft of documentary producers, the commercial precision of advertising agencies, and the accessibility of modern AI voice technology. Voiceover is the invisible art that makes or breaks video content. Great voiceover isn't just speaking clearly—it's performing the script in a way that creates the intended emotional response. The best voiceover work understands pacing, tone, emphasis, and the subtle art of making scripted words sound natural. It knows when to use human talent versus AI, and how to get the best from both. Use when "voiceover, voice over, VO, narration, narrator, voice recording, voice talent, voice actor, AI voice, text to speech, audio narration, voice direction, voiceover, audio, narration, voice, recording, AI-voice, talent, direction" mentioned.
128 · bundle
akillness
jeo-skill
Browse, group, relate, and selectively install the jeo-skills catalog through the lightweight `jeo-skill` CLI. Use when the user wants skills organized by web, infrastructure, game, creative media, CLI tools, AI/agents, engineering, research, business, or utilities; needs a frontend/backend/game-audio/game-VFX subcategory; wants overlapping skills connected instead of duplicated; or wants a category, bundle, or named skills installed without copying the full repository.
42 · bundle