God Data Cleaning

God-level data cleaning, preprocessing, and feature engineering skill. Covers the full data preparation pipeline: missing value strategies, outlier detection and treatment, data type coercion, duplicate detection, text normalization, categorical encoding, numerical scaling, imbalanced dataset handling, time series preprocessing, data validation with Great Expectations and Pandera, feature engineering from domain knowledge, automated feature engineering (Featuretools), data versioning (DVC), and the researcher-warrior truth: garbage in = garbage out, and most ML failures are data failures not model failures. Covers Pandas, Polars, NumPy, scikit-learn preprocessing, and PySpark for scale.

ArdurAI 41231a6 3 files · 32.3 KB Updated

File contents

ArdurAI/god-skill-suite/tree/main/skills/god-data-cleaning commit 41231a6862

Frequently asked questions

npx skillmds@latest add ardurai/god-data-cleaning