Splitting Datasets

Split datasets into training, validation, and test partitions with the right stratification and temporal rules. Use as a narrow preprocessing helper once the broader ML workflow is already chosen, not as the main route owner for an end-to-end ML task.

gabrielmoreira Updated 17 repo stars

File contents

Dataset Splitter

Positioning

Treat this skill as a narrow helper for partition strategy.

When to Use

Use this skill when:

  • Prepare a dataset for machine learning model training.
  • Create training, validation, and testing sets.
  • Partition data to evaluate model performance.

Not For / Boundaries

  • Full preprocessing-pipeline ownership: use preprocessing-data-with-automated-pipelines
  • Leakage audits and prediction-time checks: use ml-data-leakage-guard
  • Model training and tuning after the split: use scikit-learn

Typical Outputs

  • Partition strategy with ratios, random seeds, and stratification rules
  • Notes on temporal or grouped split constraints
  • Handoff guidance for leakage review and downstream training

Related Skills

  • preprocessing-data-with-automated-pipelines for the broader preprocessing sequence
  • ml-data-leakage-guard to verify the split does not leak future or test information

gabrielmoreira/agent-skills-mirror/tree/main/mirrors/repos/foryourhealth111-pixel@Vibe-Skills/bundled/skills/splitting-datasets commit f793d7d92c

Frequently asked questions

npx skillmds@latest add gabrielmoreira/splitting-datasets