Apify Youtube Transcripts LLM Training Data

Build LLM training data from YouTube transcripts in bulk with the Apify Actor johnvc/YoutubeTranscripts. Feed it an array of video URLs and get one clean row per video, with non_timestamped transcript text ready for tokenization, timestamped snippets, language_code, and provenance metadata (title, channel_name, upload_date, view_count) for dataset documentation. Filter by language or translate every transcript into one target language for a consistent corpus. Use when the user wants LLM training data from videos, a fine-tuning or RAG corpus from YouTube, bulk YouTube transcripts for machine learning, or a text dataset from spoken video content without running speech-to-text. Pay-per-video billing, MCP-ready for Claude and other AI agents.

johnisanerd f653c69 3 files · 11.1 KB Updated

File contents

johnisanerd/claude-skill-youtube-transcripts-llm-training-data/tree/main/apify-youtube-transcripts-llm-training-data commit f653c6913b

Frequently asked questions

npx skillmds@latest add johnisanerd/apify-youtube-transcripts-llm-training-data