Results for “full-text-extraction”
4 skillsMore results
Extract Article Text
Extract clean article content — title, author, date, and body text — from PDFs, Word docs, and web pages.
2
Ocr And Documents
Extract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
0 · bundle
Youtube Transcript
Extracts a YouTube transcript with yt-dlp, cleans it while preserving all substantive speech, and saves a structured markdown file with metadata, overview, and optional timestamps.
1