# Shared Git Data

> Sets up Git-based version control for data science projects, handling notebooks, datasets, and pipelines with DVC and nbstripout.

- Skill: `leandrobenjaminl/shared-git-data` (Agent Skill)
- Install (CLI): `npx skillmds@latest add leandrobenjaminl/shared-git-data`
- Raw SKILL.md: https://api.skillmd.com/api/skills/leandrobenjaminl/shared-git-data/raw
- Safety review: PASS (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools, Data & Analytics, Dev Tooling, ETL & Pipelines
- Tags: Data Science, Data Versioning, Datasets, Dvc, Git, Gitignore, Nbstripout, Notebooks
- License: MIT
- Author: LeandroBenjaminL (https://skillmd.com/u/leandrobenjaminl)
- Updated: 2026-08-22
- Page: https://skillmd.com/skills/leandrobenjaminl/shared-git-data

---


# Skill: shared-git-data

Git for data science. Different from code — data needs its own rules.

## Trigger

- Starting a data project and creating the repo
- Committing a notebook and you don't want to push megabytes of outputs
- Should the dataset be in the repo or not?
- You need to version pipelines without versioning heavy data

## Workflow LEND

1. ANALYZE
   ├── Project: analysis, pipeline, dashboard?
   ├── Data: how big? can it go in the repo? (< 100 MB)
   ├── Notebooks: outputs included or stripped?
   └── Collaboration: solo or team?

2. OFFER (Senior Menu)
   ├── A) Standard git — code + clean notebooks + small data (< 100 MB)
   ├── B) Git + DVC — large data versioned outside the repo
   └── C) Git + nbstripout — notebooks without outputs, clean diffs

3. CHOOSE → user confirms

4. EXECUTE
   ├── .gitignore: .venv/, __pycache__/, *.pyc, .env, data/raw/
   ├── nbstripout: `git config filter.nbstripout.extrakeys 'metadata.kernel'`
   ├── DVC: `dvc init && dvc add data/raw/dataset.csv && git add data/raw/dataset.csv.dvc`
   ├── README.md with reproduction instructions
   ├── requirements.txt frozen
   ├── Commits in English US, semantic (feat, fix, chore)
   └── Branch + PR workflow for data pipelines

5. VERIFY
   ├── .gitignore covers what shouldn't be in the repo
   ├── Notebooks can be diffed (no outputs)
   └── A fresh `git clone && pip install` is enough to start

- [ ] Save decisions and changes to Engram (mem_save)
- [ ] Check if README, AGENTS.md, or ARCHITECTURE.md need updating
- [ ] If docs changed → update them in the same PR/commit

