Trafilatura Web Text Extraction and Crawling Toolkit

Trafilatura is a Python package and CLI tool for gathering text from the web. It handles crawling, downloading, and extracting main text content, metadata, and comments from raw HTML, outputting clean structured data in CSV, JSON, Markdown, XML, and TXT formats.

agentskillexchange Updated 28 repo stars

File contents

agentskillexchange/skills/tree/main/skills/trafilatura-web-text-extraction-crawling commit aa1bdbedf5

Frequently asked questions

npx skillmds@latest add agentskillexchange/trafilatura-web-text-extraction-and-crawling-toolkit