Skill: shared-git-data
Git for data science. Different from code — data needs its own rules.
Trigger
- Starting a data project and creating the repo
- Committing a notebook and you don't want to push megabytes of outputs
- Should the dataset be in the repo or not?
- You need to version pipelines without versioning heavy data
Workflow LEND
ANALYZE ├── Project: analysis, pipeline, dashboard? ├── Data: how big? can it go in the repo? (< 100 MB) ├── Notebooks: outputs included or stripped? └── Collaboration: solo or team?
OFFER (Senior Menu) ├── A) Standard git — code + clean notebooks + small data (< 100 MB) ├── B) Git + DVC — large data versioned outside the repo └── C) Git + nbstripout — notebooks without outputs, clean diffs
CHOOSE → user confirms
EXECUTE ├── .gitignore: .venv/, pycache/, *.pyc, .env, data/raw/ ├── nbstripout:
git config filter.nbstripout.extrakeys 'metadata.kernel'├── DVC:dvc init && dvc add data/raw/dataset.csv && git add data/raw/dataset.csv.dvc├── README.md with reproduction instructions ├── requirements.txt frozen ├── Commits in English US, semantic (feat, fix, chore) └── Branch + PR workflow for data pipelinesVERIFY ├── .gitignore covers what shouldn't be in the repo ├── Notebooks can be diffed (no outputs) └── A fresh
git clone && pip installis enough to start
- Save decisions and changes to Engram (mem_save)
- Check if README, AGENTS.md, or ARCHITECTURE.md need updating
- If docs changed → update them in the same PR/commit