Data Engineering Discipline

Discipline guardrails for data-engineering work with downstream consumers — activate at the START of the task, before writing code, because silent semantic drift is the dominant risk. Activate on: migrating or porting a pipeline, refactoring a transform, backfilling or replaying history, evolving a schema (add / rename / retype / drop a column), creating a new dataset — or a metadata / catalog / lineage emitter whose output a separate tool loads — that has consumers, designing or reviewing a data contract, reshaping a tool / API response payload a client depends on, writing or changing the tests, fixtures, or expected values that gate a pipeline, generating pipeline code with an LLM, and investigating a consumed dataset that misbehaves — "the numbers changed / look different", or a table/extract that "ran but didn't update / is stale / isn't refreshing / the watermark didn't advance". A hand-authored schema-as-data document counts when code is generated from or validated against it; not when its only consumer

grimaldost 8308333 20 files · 230.5 KB Updated

File contents

grimaldost/craft-collection/tree/main/plugins/engineering-discipline/skills/data-engineering-discipline commit 8308333d00

Frequently asked questions

npx skillmds@latest add grimaldost/data-engineering-discipline