Packs
2 packscurated
Text-to-Speech Podcast
Convert a script into a multi-speaker podcast audio with TTS, dialogue, and optional music.
4 skills · pack
curated
Build Gemini Live API App
Build real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket-based audio/video/text streaming and function calling.
4 skills · pack
Results for “audio”
8 skillscaa-eval
Benchmarks large audio-language models against adversarial audio attacks using the CAA dataset, computing WER, ROUGE-L, cosine similarity, and coherence scores to assess robustness in conversational settings.
3
notebooklm
Browser-automates Google's NotebookLM to read notebooks, add sources, generate Studio outputs (Audio/Video Overviews, Mind Maps, Reports), and create new notebooks.
20.4k · bundle
summarize
Summarize URLs or files with the summarize CLI (web, PDFs, images, audio, YouTube).
0 · bundle
More results
summarize
Summarize URLs or files with the summarize CLI (web, PDFs, images, audio, YouTube).
228
competition-stego-media
Inspects metadata, hidden channels, and appended payloads in media files to recover concealed data in steganography challenges.
12.8k · bundle
notebooklm-integration
Wraps the notebooklm-py CLI to ingest sources and generate synthesized artifacts like podcasts, slide decks, and quizzes from Google NotebookLM.
2
gemini-interactions-api
Call the Gemini API for text generation, chat, multimodal understanding, image/video/audio generation, streaming, function calling, structured output, and managed agents using the Interactions API in Python and TypeScript.
3.8k · bundle
podcast-transcript-txt
Finds and exports podcast episode transcripts as cleaned TXT files from YouTube URLs, episode webpages, Apple Podcasts, X/Twitter links, direct audio URLs, or plain titles, with a deterministic decision tree and optional local ASR fallback.
37 · bundle