Data Pipeline Architecture
Overview
This public intake copy packages plugins/antigravity-awesome-skills/skills/data-engineering-data-pipeline from https://github.com/sickn33/antigravity-awesome-skills into the native Omni Skills editorial shape without hiding its origin.
Use it when the operator needs the upstream workflow, support files, and repository context to stay intact while the public validator and private enhancer continue their normal downstream flow.
This intake keeps the copied upstream files intact and uses the external_source block in metadata.json plus ORIGIN.md as the provenance anchor for review.
Data Pipeline Architecture You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.
Imported source sections that did not map cleanly to the public headings are still preserved below or in the support files. Notable imported sections: Requirements, Core Capabilities, Output Deliverables, Success Criteria, Limitations.
When to Use This Skill
Use this section as the trigger filter. It should make the activation boundary explicit before the operator loads files, runs commands, or opens a pull request.
- Working on data pipeline architecture tasks or workflows
- Needing guidance, best practices, or checklists for data pipeline architecture
- The task is unrelated to data pipeline architecture
- You need a different domain or tool outside this scope
- Use when the request clearly matches the imported source intent: You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.
- Use when the operator should preserve upstream workflow detail instead of rewriting the process from scratch.
Operating Table
| Situation |
Start here |
Why it matters |
| First-time use |
metadata.json |
Confirms repository, branch, commit, and imported path through the external_source block before touching the copied workflow |
| Provenance review |
ORIGIN.md |
Gives reviewers a plain-language audit trail for the imported source |
| Workflow execution |
SKILL.md |
Starts with the smallest copied file that materially changes execution |
| Supporting context |
SKILL.md |
Adds the next most relevant copied source file without loading the entire package |
| Handoff decision |
## Related Skills |
Helps the operator switch to a stronger native skill when the task drifts |
Workflow
This workflow is intentionally editorial and operational at the same time. It keeps the imported source useful to the operator while still satisfying the public intake standards that feed the downstream enhancer flow.
- Assess: sources, volume, latency requirements, targets
- Select pattern: ETL (transform before load), ELT (load then transform), Lambda (batch + speed layers), Kappa (stream-only), Lakehouse (unified)
- Design flow: sources → ingestion → processing → storage → serving
- Add observability touchpoints
- Incremental loading with watermark columns
- Retry logic with exponential backoff
- Schema validation and dead letter queue for invalid records
Imported Workflow Notes
Imported: Instructions
1. Architecture Design
- Assess: sources, volume, latency requirements, targets
- Select pattern: ETL (transform before load), ELT (load then transform), Lambda (batch + speed layers), Kappa (stream-only), Lakehouse (unified)
- Design flow: sources → ingestion → processing → storage → serving
- Add observability touchpoints
2. Ingestion Implementation
Batch
- Incremental loading with watermark columns
- Retry logic with exponential backoff
- Schema validation and dead letter queue for invalid records
- Metadata tracking (_extracted_at, _source)
Streaming
- Kafka consumers with exactly-once semantics
- Manual offset commits within transactions
- Windowing for time-based aggregations
- Error handling and replay capability
3. Orchestration
Airflow
- Task groups for logical organization
- XCom for inter-task communication
- SLA monitoring and email alerts
- Incremental execution with execution_date
- Retry with exponential backoff
Prefect
- Task caching for idempotency
- Parallel execution with .submit()
- Artifacts for visibility
- Automatic retries with configurable delays
4. Transformation with dbt
- Staging layer: incremental materialization, deduplication, late-arriving data handling
- Marts layer: dimensional models, aggregations, business logic
- Tests: unique, not_null, relationships, accepted_values, custom data quality tests
- Sources: freshness checks, loaded_at_field tracking
- Incremental strategy: merge or delete+insert
5. Data Quality Framework
Great Expectations
- Table-level: row count, column count
- Column-level: uniqueness, nullability, type validation, value sets, ranges
- Checkpoints for validation execution
- Data docs for documentation
- Failure notifications
dbt Tests
- Schema tests in YAML
- Custom data quality tests with dbt-expectations
- Test results tracked in metadata
6. Storage Strategy
Delta Lake
- ACID transactions with append/overwrite/merge modes
- Upsert with predicate-based matching
- Time travel for historical queries
- Optimize: compact small files, Z-order clustering
- Vacuum to remove old files
Apache Iceberg
- Partitioning and sort order optimization
- MERGE INTO for upserts
- Snapshot isolation and time travel
- File compaction with binpack strategy
- Snapshot expiration for cleanup
7. Monitoring & Cost Optimization
Monitoring
- Track: records processed/failed, data size, execution time, success/failure rates
- CloudWatch metrics and custom namespaces
- SNS alerts for critical/warning/info events
- Data freshness checks
- Performance trend analysis
Cost Optimization
- Partitioning: date/entity-based, avoid over-partitioning (keep >1GB)
- File sizes: 512MB-1GB for Parquet
- Lifecycle policies: hot (Standard) → warm (IA) → cold (Glacier)
- Compute: spot instances for batch, on-demand for streaming, serverless for adhoc
- Query optimization: partition pruning, clustering, predicate pushdown
Imported: Requirements
$ARGUMENTS
Examples
Example 1: Ask for the upstream workflow directly
Use @data-engineering-data-pipeline-v2 to handle <task>. Start from the copied upstream workflow, load only the files that change the outcome, and keep provenance visible in the answer.
Explanation: This is the safest starting point when the operator needs the imported workflow, but not the entire repository.
Example 2: Ask for a provenance-grounded review
Review @data-engineering-data-pipeline-v2 against metadata.json and ORIGIN.md, then explain which copied upstream files you would load first and why.
Explanation: Use this before review or troubleshooting when you need a precise, auditable explanation of origin and file selection.
Example 3: Narrow the copied support files before execution
Use @data-engineering-data-pipeline-v2 for <task>. Load only the copied references, examples, or scripts that change the outcome, and name the files explicitly before proceeding.
Explanation: This keeps the skill aligned with progressive disclosure instead of loading the whole copied package by default.
Example 4: Build a reviewer packet
Review @data-engineering-data-pipeline-v2 using the copied upstream files plus provenance, then summarize any gaps before merge.
Explanation: This is useful when the PR is waiting for human review and you want a repeatable audit packet.
Imported Usage Notes
Imported: Example: Minimal Batch Pipeline
# Batch ingestion with validation
from batch_ingestion import BatchDataIngester
from storage.delta_lake_manager import DeltaLakeManager
from data_quality.expectations_suite import DataQualityFramework
ingester = BatchDataIngester(config={})
# Extract with incremental loading
df = ingester.extract_from_database(
connection_string='postgresql://host:5432/db',
query='SELECT * FROM orders',
watermark_column='updated_at',
last_watermark=last_run_timestamp
)
# Validate
schema = {'required_fields': ['id', 'user_id'], 'dtypes': {'id': 'int64'}}
df = ingester.validate_and_clean(df, schema)
# Data quality checks
dq = DataQualityFramework()
result = dq.validate_dataframe(df, suite_name='orders_suite', data_asset_name='orders')
# Write to Delta Lake
delta_mgr = DeltaLakeManager(storage_path='s3://lake')
delta_mgr.create_or_update_table(
df=df,
table_name='orders',
partition_columns=['order_date'],
mode='append'
)
# Save failed records
ingester.save_dead_letter_queue('s3://lake/dlq/orders')
Best Practices
Treat the generated public skill as a reviewable packaging layer around the upstream repository. The goal is to keep provenance explicit and load only the copied source material that materially improves execution.
- Keep the imported skill grounded in the upstream repository; do not invent steps that the source material cannot support.
- Prefer the smallest useful set of support files so the workflow stays auditable and fast to review.
- Keep provenance, source commit, and imported file paths visible in notes and PR descriptions.
- Point directly at the copied upstream files that justify the workflow instead of relying on generic review boilerplate.
- Treat generated examples as scaffolding; adapt them to the concrete task before execution.
- Route to a stronger native skill when architecture, debugging, design, or security concerns become dominant.
Troubleshooting
Problem: The operator skipped the imported context and answered too generically
Symptoms: The result ignores the upstream workflow in plugins/antigravity-awesome-skills/skills/data-engineering-data-pipeline, fails to mention provenance, or does not use any copied source files at all.
Solution: Re-open metadata.json, ORIGIN.md, and the most relevant copied upstream files. Check the external_source block first, then restate the provenance before continuing.
Problem: The imported workflow feels incomplete during review
Symptoms: Reviewers can see the generated SKILL.md, but they cannot quickly tell which references, examples, or scripts matter for the current task.
Solution: Point at the exact copied references, examples, scripts, or assets that justify the path you took. If the gap is still real, record it in the PR instead of hiding it.
Problem: The task drifted into a different specialization
Symptoms: The imported skill starts in the right place, but the work turns into debugging, architecture, design, security, or release orchestration that a native skill handles better.
Solution: Use the related skills section to hand off deliberately. Keep the imported provenance visible so the next skill inherits the right context instead of starting blind.
Related Skills
@00-andruia-consultant - Use when the work is better handled by that native specialization after this imported skill establishes context.
@00-andruia-consultant-v2 - Use when the work is better handled by that native specialization after this imported skill establishes context.
@10-andruia-skill-smith - Use when the work is better handled by that native specialization after this imported skill establishes context.
@10-andruia-skill-smith-v2 - Use when the work is better handled by that native specialization after this imported skill establishes context.
Additional Resources
Use this support matrix and the linked files below as the operator packet for this imported skill. They should reflect real copied source material, not generic scaffolding.
| Resource family |
What it gives the reviewer |
Example path |
references |
copied reference notes, guides, or background material from upstream |
references/n/a |
examples |
worked examples or reusable prompts copied from upstream |
examples/n/a |
scripts |
upstream helper scripts that change execution or validation |
scripts/n/a |
agents |
routing or delegation notes that are genuinely part of the imported package |
agents/n/a |
assets |
supporting assets or schemas copied from the source package |
assets/n/a |
Imported Reference Notes
Imported: Core Capabilities
- Design ETL/ELT, Lambda, Kappa, and Lakehouse architectures
- Implement batch and streaming data ingestion
- Build workflow orchestration with Airflow/Prefect
- Transform data using dbt and Spark
- Manage Delta Lake/Iceberg storage with ACID transactions
- Implement data quality frameworks (Great Expectations, dbt tests)
- Monitor pipelines with CloudWatch/Prometheus/Grafana
- Optimize costs through partitioning, lifecycle policies, and compute optimization
Imported: Output Deliverables
1. Architecture Documentation
- Architecture diagram with data flow
- Technology stack with justification
- Scalability analysis and growth patterns
- Failure modes and recovery strategies
2. Implementation Code
- Ingestion: batch/streaming with error handling
- Transformation: dbt models (staging → marts) or Spark jobs
- Orchestration: Airflow/Prefect DAGs with dependencies
- Storage: Delta/Iceberg table management
- Data quality: Great Expectations suites and dbt tests
3. Configuration Files
- Orchestration: DAG definitions, schedules, retry policies
- dbt: models, sources, tests, project config
- Infrastructure: Docker Compose, K8s manifests, Terraform
- Environment: dev/staging/prod configs
4. Monitoring & Observability
- Metrics: execution time, records processed, quality scores
- Alerts: failures, performance degradation, data freshness
- Dashboards: Grafana/CloudWatch for pipeline health
- Logging: structured logs with correlation IDs
5. Operations Guide
- Deployment procedures and rollback strategy
- Troubleshooting guide for common issues
- Scaling guide for increased volume
- Cost optimization strategies and savings
- Disaster recovery and backup procedures
Imported: Success Criteria
- Pipeline meets defined SLA (latency, throughput)
- Data quality checks pass with >99% success rate
- Automatic retry and alerting on failures
- Comprehensive monitoring shows health and performance
- Documentation enables team maintenance
- Cost optimization reduces infrastructure costs by 30-50%
- Schema evolution without downtime
- End-to-end data lineage tracked
Imported: Limitations
- Use this skill only when the task clearly matches the scope described above.
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
1---2name: data-engineering-data-pipeline-v23description: Data Pipeline Architecture workflow skill. Use this skill when the user needs You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.4---56# Data Pipeline Architecture78## Overview910This public intake copy packages `plugins/antigravity-awesome-skills/skills/data-engineering-data-pipeline` from `https://github.com/sickn33/antigravity-awesome-skills` into the native Omni Skills editorial shape without hiding its origin.1112Use it when the operator needs the upstream workflow, support files, and repository context to stay intact while the public validator and private enhancer continue their normal downstream flow.1314This intake keeps the copied upstream files intact and uses the `external_source` block in `metadata.json` plus `ORIGIN.md` as the provenance anchor for review.1516# Data Pipeline Architecture You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.1718Imported source sections that did not map cleanly to the public headings are still preserved below or in the support files. Notable imported sections: Requirements, Core Capabilities, Output Deliverables, Success Criteria, Limitations.1920## When to Use This Skill2122Use this section as the trigger filter. It should make the activation boundary explicit before the operator loads files, runs commands, or opens a pull request.2324- Working on data pipeline architecture tasks or workflows25- Needing guidance, best practices, or checklists for data pipeline architecture26- The task is unrelated to data pipeline architecture27- You need a different domain or tool outside this scope28- Use when the request clearly matches the imported source intent: You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.29- Use when the operator should preserve upstream workflow detail instead of rewriting the process from scratch.3031## Operating Table3233| Situation | Start here | Why it matters |34| --- | --- | --- |35| First-time use | `metadata.json` | Confirms repository, branch, commit, and imported path through the `external_source` block before touching the copied workflow |36| Provenance review | `ORIGIN.md` | Gives reviewers a plain-language audit trail for the imported source |37| Workflow execution | `SKILL.md` | Starts with the smallest copied file that materially changes execution |38| Supporting context | `SKILL.md` | Adds the next most relevant copied source file without loading the entire package |39| Handoff decision | `## Related Skills` | Helps the operator switch to a stronger native skill when the task drifts |4041## Workflow4243This workflow is intentionally editorial and operational at the same time. It keeps the imported source useful to the operator while still satisfying the public intake standards that feed the downstream enhancer flow.44451. Assess: sources, volume, latency requirements, targets462. Select pattern: ETL (transform before load), ELT (load then transform), Lambda (batch + speed layers), Kappa (stream-only), Lakehouse (unified)473. Design flow: sources → ingestion → processing → storage → serving484. Add observability touchpoints495. Incremental loading with watermark columns506. Retry logic with exponential backoff517. Schema validation and dead letter queue for invalid records5253### Imported Workflow Notes5455#### Imported: Instructions5657### 1. Architecture Design58- Assess: sources, volume, latency requirements, targets59- Select pattern: ETL (transform before load), ELT (load then transform), Lambda (batch + speed layers), Kappa (stream-only), Lakehouse (unified)60- Design flow: sources → ingestion → processing → storage → serving61- Add observability touchpoints6263### 2. Ingestion Implementation64**Batch**65- Incremental loading with watermark columns66- Retry logic with exponential backoff67- Schema validation and dead letter queue for invalid records68- Metadata tracking (_extracted_at, _source)6970**Streaming**71- Kafka consumers with exactly-once semantics72- Manual offset commits within transactions73- Windowing for time-based aggregations74- Error handling and replay capability7576### 3. Orchestration77**Airflow**78- Task groups for logical organization79- XCom for inter-task communication80- SLA monitoring and email alerts81- Incremental execution with execution_date82- Retry with exponential backoff8384**Prefect**85- Task caching for idempotency86- Parallel execution with .submit()87- Artifacts for visibility88- Automatic retries with configurable delays8990### 4. Transformation with dbt91- Staging layer: incremental materialization, deduplication, late-arriving data handling92- Marts layer: dimensional models, aggregations, business logic93- Tests: unique, not_null, relationships, accepted_values, custom data quality tests94- Sources: freshness checks, loaded_at_field tracking95- Incremental strategy: merge or delete+insert9697### 5. Data Quality Framework98**Great Expectations**99- Table-level: row count, column count100- Column-level: uniqueness, nullability, type validation, value sets, ranges101- Checkpoints for validation execution102- Data docs for documentation103- Failure notifications104105**dbt Tests**106- Schema tests in YAML107- Custom data quality tests with dbt-expectations108- Test results tracked in metadata109110### 6. Storage Strategy111**Delta Lake**112- ACID transactions with append/overwrite/merge modes113- Upsert with predicate-based matching114- Time travel for historical queries115- Optimize: compact small files, Z-order clustering116- Vacuum to remove old files117118**Apache Iceberg**119- Partitioning and sort order optimization120- MERGE INTO for upserts121- Snapshot isolation and time travel122- File compaction with binpack strategy123- Snapshot expiration for cleanup124125### 7. Monitoring & Cost Optimization126**Monitoring**127- Track: records processed/failed, data size, execution time, success/failure rates128- CloudWatch metrics and custom namespaces129- SNS alerts for critical/warning/info events130- Data freshness checks131- Performance trend analysis132133**Cost Optimization**134- Partitioning: date/entity-based, avoid over-partitioning (keep >1GB)135- File sizes: 512MB-1GB for Parquet136- Lifecycle policies: hot (Standard) → warm (IA) → cold (Glacier)137- Compute: spot instances for batch, on-demand for streaming, serverless for adhoc138- Query optimization: partition pruning, clustering, predicate pushdown139140#### Imported: Requirements141142$ARGUMENTS143144## Examples145146### Example 1: Ask for the upstream workflow directly147148```text149Use @data-engineering-data-pipeline-v2 to handle <task>. Start from the copied upstream workflow, load only the files that change the outcome, and keep provenance visible in the answer.150```151152**Explanation:** This is the safest starting point when the operator needs the imported workflow, but not the entire repository.153154### Example 2: Ask for a provenance-grounded review155156```text157Review @data-engineering-data-pipeline-v2 against metadata.json and ORIGIN.md, then explain which copied upstream files you would load first and why.158```159160**Explanation:** Use this before review or troubleshooting when you need a precise, auditable explanation of origin and file selection.161162### Example 3: Narrow the copied support files before execution163164```text165Use @data-engineering-data-pipeline-v2 for <task>. Load only the copied references, examples, or scripts that change the outcome, and name the files explicitly before proceeding.166```167168**Explanation:** This keeps the skill aligned with progressive disclosure instead of loading the whole copied package by default.169170### Example 4: Build a reviewer packet171172```text173Review @data-engineering-data-pipeline-v2 using the copied upstream files plus provenance, then summarize any gaps before merge.174```175176**Explanation:** This is useful when the PR is waiting for human review and you want a repeatable audit packet.177178### Imported Usage Notes179180#### Imported: Example: Minimal Batch Pipeline181182```python183# Batch ingestion with validation184from batch_ingestion import BatchDataIngester185from storage.delta_lake_manager import DeltaLakeManager186from data_quality.expectations_suite import DataQualityFramework187188ingester = BatchDataIngester(config={})189190# Extract with incremental loading191df = ingester.extract_from_database(192 connection_string='postgresql://host:5432/db',193 query='SELECT * FROM orders',194 watermark_column='updated_at',195 last_watermark=last_run_timestamp196)197198# Validate199schema = {'required_fields': ['id', 'user_id'], 'dtypes': {'id': 'int64'}}200df = ingester.validate_and_clean(df, schema)201202# Data quality checks203dq = DataQualityFramework()204result = dq.validate_dataframe(df, suite_name='orders_suite', data_asset_name='orders')205206# Write to Delta Lake207delta_mgr = DeltaLakeManager(storage_path='s3://lake')208delta_mgr.create_or_update_table(209 df=df,210 table_name='orders',211 partition_columns=['order_date'],212 mode='append'213)214215# Save failed records216ingester.save_dead_letter_queue('s3://lake/dlq/orders')217```218219## Best Practices220221Treat the generated public skill as a reviewable packaging layer around the upstream repository. The goal is to keep provenance explicit and load only the copied source material that materially improves execution.222223- Keep the imported skill grounded in the upstream repository; do not invent steps that the source material cannot support.224- Prefer the smallest useful set of support files so the workflow stays auditable and fast to review.225- Keep provenance, source commit, and imported file paths visible in notes and PR descriptions.226- Point directly at the copied upstream files that justify the workflow instead of relying on generic review boilerplate.227- Treat generated examples as scaffolding; adapt them to the concrete task before execution.228- Route to a stronger native skill when architecture, debugging, design, or security concerns become dominant.229230231232## Troubleshooting233234### Problem: The operator skipped the imported context and answered too generically235236**Symptoms:** The result ignores the upstream workflow in `plugins/antigravity-awesome-skills/skills/data-engineering-data-pipeline`, fails to mention provenance, or does not use any copied source files at all.237**Solution:** Re-open `metadata.json`, `ORIGIN.md`, and the most relevant copied upstream files. Check the `external_source` block first, then restate the provenance before continuing.238239### Problem: The imported workflow feels incomplete during review240241**Symptoms:** Reviewers can see the generated `SKILL.md`, but they cannot quickly tell which references, examples, or scripts matter for the current task.242**Solution:** Point at the exact copied references, examples, scripts, or assets that justify the path you took. If the gap is still real, record it in the PR instead of hiding it.243244### Problem: The task drifted into a different specialization245246**Symptoms:** The imported skill starts in the right place, but the work turns into debugging, architecture, design, security, or release orchestration that a native skill handles better.247**Solution:** Use the related skills section to hand off deliberately. Keep the imported provenance visible so the next skill inherits the right context instead of starting blind.248249250251## Related Skills252253- `@00-andruia-consultant` - Use when the work is better handled by that native specialization after this imported skill establishes context.254- `@00-andruia-consultant-v2` - Use when the work is better handled by that native specialization after this imported skill establishes context.255- `@10-andruia-skill-smith` - Use when the work is better handled by that native specialization after this imported skill establishes context.256- `@10-andruia-skill-smith-v2` - Use when the work is better handled by that native specialization after this imported skill establishes context.257258## Additional Resources259260Use this support matrix and the linked files below as the operator packet for this imported skill. They should reflect real copied source material, not generic scaffolding.261262| Resource family | What it gives the reviewer | Example path |263| --- | --- | --- |264| `references` | copied reference notes, guides, or background material from upstream | `references/n/a` |265| `examples` | worked examples or reusable prompts copied from upstream | `examples/n/a` |266| `scripts` | upstream helper scripts that change execution or validation | `scripts/n/a` |267| `agents` | routing or delegation notes that are genuinely part of the imported package | `agents/n/a` |268| `assets` | supporting assets or schemas copied from the source package | `assets/n/a` |269270271272### Imported Reference Notes273274#### Imported: Core Capabilities275276- Design ETL/ELT, Lambda, Kappa, and Lakehouse architectures277- Implement batch and streaming data ingestion278- Build workflow orchestration with Airflow/Prefect279- Transform data using dbt and Spark280- Manage Delta Lake/Iceberg storage with ACID transactions281- Implement data quality frameworks (Great Expectations, dbt tests)282- Monitor pipelines with CloudWatch/Prometheus/Grafana283- Optimize costs through partitioning, lifecycle policies, and compute optimization284285#### Imported: Output Deliverables286287### 1. Architecture Documentation288- Architecture diagram with data flow289- Technology stack with justification290- Scalability analysis and growth patterns291- Failure modes and recovery strategies292293### 2. Implementation Code294- Ingestion: batch/streaming with error handling295- Transformation: dbt models (staging → marts) or Spark jobs296- Orchestration: Airflow/Prefect DAGs with dependencies297- Storage: Delta/Iceberg table management298- Data quality: Great Expectations suites and dbt tests299300### 3. Configuration Files301- Orchestration: DAG definitions, schedules, retry policies302- dbt: models, sources, tests, project config303- Infrastructure: Docker Compose, K8s manifests, Terraform304- Environment: dev/staging/prod configs305306### 4. Monitoring & Observability307- Metrics: execution time, records processed, quality scores308- Alerts: failures, performance degradation, data freshness309- Dashboards: Grafana/CloudWatch for pipeline health310- Logging: structured logs with correlation IDs311312### 5. Operations Guide313- Deployment procedures and rollback strategy314- Troubleshooting guide for common issues315- Scaling guide for increased volume316- Cost optimization strategies and savings317- Disaster recovery and backup procedures318319#### Imported: Success Criteria320321- Pipeline meets defined SLA (latency, throughput)322- Data quality checks pass with >99% success rate323- Automatic retry and alerting on failures324- Comprehensive monitoring shows health and performance325- Documentation enables team maintenance326- Cost optimization reduces infrastructure costs by 30-50%327- Schema evolution without downtime328- End-to-end data lineage tracked329330#### Imported: Limitations331332- Use this skill only when the task clearly matches the scope described above.333- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.334- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.