ETL & Pipelines Agent Skills
ETL & Pipelines
153 skillssurreal-sync
Migrates data from MongoDB, PostgreSQL, MySQL, Neo4j, Kafka, and JSONL into SurrealDB with full and incremental CDC synchronization.
34
csv
Generates Python code to batch-translate the English column of tab-separated CSV files while preserving the original Chinese column and output format.
559
apify-actor-runner
Runs Apify cloud actors for structured web scraping and exports datasets to S3, with input schema validation and webhook notifications.
28
connect
Build streaming data pipelines with Redpanda Connect using declarative YAML configs, Bloblang mappings, and component discovery. Covers running, linting, and dry-running pipelines.
6 · bundle
connect-cdc-oracle
Streams change data capture from Oracle Database into Redpanda or Kafka using the oracledb_cdc input in Redpanda Connect, which reads redo logs via LogMiner. Covers configuration, Oracle setup, checkpointing, and enterprise features.
6 · bundle
connect-cdc-mongodb
Streams change data capture from MongoDB into Redpanda or Kafka using Redpanda Connect's mongodb_cdc input, covering Change Streams, snapshots, document modes, and resume-token checkpointing.
6 · bundle
connect-cdc-spanner
Streams change data capture from Google Cloud Spanner into Redpanda or Kafka using Redpanda Connect's gcp_spanner_cdc input, with partition-aware watermarked delivery and support for Redpanda Enterprise features.
6 · bundle
connect-cdc-dynamodb
Guides setup and operation of the aws_dynamodb_cdc input in Redpanda Connect, which streams change data capture from AWS DynamoDB into Redpanda or Kafka using DynamoDB Streams. Covers enabling streams, IAM policies, checkpoint tables, snapshot modes, table discovery, and operational constraints.
6 · bundle
connect-cdc-postgres
Streams change data capture from PostgreSQL into Redpanda or Kafka using Redpanda Connect's postgres_cdc input, covering setup, snapshotting, and troubleshooting.
6 · bundle
connect-cdc-sqlserver
Streams change data capture from Microsoft SQL Server into Redpanda or Kafka using Redpanda Connect's microsoft_sql_server_cdc input, covering setup, configuration, and troubleshooting.
6 · bundle
connect-cdc-tigerbeetle
Streams change data capture events from a TigerBeetle financial transactions database into Redpanda or Kafka using the tigerbeetle_cdc input, with checkpointing, filtering, and routing guidance.
6 · bundle
pysam
Read, write, and analyze genomic datasets including SAM/BAM/CRAM alignments, VCF/BCF variants, and FASTA/FASTQ sequences using a Pythonic interface to htslib.
253 · bundle
scanpy
Runs standard single-cell RNA-seq analysis with Scanpy, covering QC, normalization, dimensionality reduction, clustering, marker identification, visualization, and conversion of R single-cell formats to h5ad.
253 · bundle
anndata
Manages annotated data matrices for single-cell genomics, covering creation, I/O, concatenation, and manipulation of AnnData objects in h5ad and zarr formats.
253 · bundle
lamindb
Manages biological datasets and models with LaminDB, covering setup, artifact registration, querying, lineage tracking, validation, ontology annotation, collections, branches, storage, and workflow integrations.
253 · bundle
export-premiere
Exports a timeline.json as a Premiere Pro editor packet, including FCP7 XML, captions, and media, with verification steps.
3
data-analyst
Guides data analysis, EDA, and ML tasks by teaching, diagnosing MCPs, and deciding with the user, offering multiple options and documenting decisions.
0
etl-pipelines
Construye pipelines ETL/ELT con Pandas: extracción, transformación y carga de datos con logging, manejo de errores, idempotencia y opciones de orquestación.
0 · bundle
polars
Process tabular data with Polars' expression API, lazy evaluation, and parallel execution for faster pandas-style workflows.
2
interactive-page-repost-publisher
Ingests JavaScript-heavy external pages into StaticFlow as standalone local interactive mirrors backed by LanceDB, with bilingual article write-back and localized interactive locales.
0
audio-scrape
Discovers podcasts via the iTunes Search API, parses RSS feeds, downloads audio, transcribes with OpenAI Whisper, chunks transcripts, and upserts results into a database table.
1
book-citations
Restores and normalizes footnotes/endnotes in books by extracting from PDF/EPUB, matching anchors, staging markers, and compiling to HTML with assertions.
1 · bundle
youtube-scrape
Scrapes YouTube channel or video metadata via the YouTube Data API, downloads transcripts with yt-dlp, chunks them for search, and upserts results into a database.
1
book-chunk
Chunks a book into canonical retrieval units with heading-aware structure splitting, recursive token targets, and contextual prefixes for downstream RAG ingestion.
1
book-ingest
Upserts a validated MDX book corpus into Supabase via Drizzle, hydrating books, chapters, sections, and chunks tables while preserving stable bookmark anchors and only re-embedding changed content.
1
soul2dna
Compile SOUL.md character profiles into synthetic diploid genomes (.genome.json) via trait-to-allele mapping.
17 · bundle
dbt
Provides guidance and best practices for working with dbt, covering core concepts, patterns, troubleshooting, and deployment strategies.
1
big-data
Designs and implements big data architectures, processes large-scale datasets with distributed systems, and optimizes data pipelines for throughput using Hadoop, Spark, and cloud platforms.
1
data-scraper-agent
Builds a scheduled, AI-powered data collection agent that scrapes public sources, enriches results with Gemini Flash, and stores them in Notion, Sheets, or Supabase.
1 · bundle
029-how-eebc0870
Guides when to use Salesforce Bulk API 2.0 and provides commands for importing, updating, upserting, deleting, and exporting records, along with CSV format requirements, limits, and error handling.
7 · bundle
polars
Process in-memory tabular data with a fast, expression-based DataFrame library that supports lazy evaluation, parallel execution, and Apache Arrow semantics.
3
database
Guides database design, implementation, optimization, migration, pipeline development, quality, and operations across SQL and NoSQL platforms.
3
database
Guides database design, implementation, optimization, migrations, data pipelines, and operations across SQL and NoSQL platforms.
5
dbt-etl
Guides the setup and use of dbt for extract-transform-load workflows.
2 · bundle
scrape-webpage
Extract content, metadata, and images from a webpage for import or migration to AEM Edge Delivery Services.
142 · bundle
issue-fields-migration
Bulk-migrate repo labels and Project V2 fields into GitHub org-level issue fields (single select, text, number, date).
36.2k · bundle