Results for “speech-recognition”
22 skillsNemotron Speech
Routes NVIDIA Nemotron Speech (Riva) NIM tasks for ASR, TTS, and NMT, covering cloud-hosted inference, self-hosted Docker deployment, and custom model builds.
2.2k · bundle
Asr
Transcribe audio files to text using local speech recognition. Triggers on: "转录", "transcribe", "语音转文字", "ASR", "识别音频", "把这段音频转成文字".
0 · bundle
More results
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Resemble Detect
Detect AI-generated audio, images, video, and text, trace synthesis sources, apply watermarks, verify speaker identity, and analyze media intelligence using the Resemble AI platform.
36.2k · bundle
Transcribe
Transcribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
23.3k · bundle
Talking Head Recut
Packages an existing talking-head, interview, or podcast video with timed, designed graphic overlay cards—kinetic titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture—synced to the transcript, on a 16:9, 9:16, or 4:5 canvas.
· bundle
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
2
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Detecting Deepfake Audio In Vishing Attacks
Detects AI-generated deepfake audio used in voice phishing (vishing) attacks by extracting spectral features and classifying samples with machine learning models.
24.6k · bundle
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
6
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Open Vocabulary Object Detection Using Captions Arxiv 2011 1
Open-Vocabulary Object Detection Using Captions
6
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
0
Brand Voice
从真实的帖子、文章、发布说明、文档或网站文案中构建基于源材料的写作风格档案,然后在内容、外展和社交工作流中重复使用该档案。当用户希望保持声音一致性而不使用通用的AI写作套路时使用。
0 · bundle
Asr Whisper For Video Transcription Arxiv 2212 04356v1
ASR: Whisper for Video Transcription
6
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
2
Transcribe
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
0 · bundle
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
Voice Agents
Voice Agents
128 · bundle
Transcribe
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
65 · bundle
Transformers JS
Run state-of-the-art machine learning models directly in JavaScript/TypeScript across browsers and server-side runtimes using Transformers.js.
10.8k · bundle
Digital Health Clinical Asr Setup
Bootstraps a clinical ASR evaluation environment by verifying NVIDIA_API_KEY, installing Python dependencies, and running a smoke test against hosted TTS/ASR services.
2.2k · bundle