vaquarkhan
- 119 skills
- 0 followers
- 19 hours ago last updated
- ▌ Snowflake Native Pipelines And Governance · vaquarkhanGuides agents through Snowflake-native pipeline and governance workflows. Use when building or reviewing Snowflake pipelines with Streams, Tasks, Dynamic Tables, Snowpipe, Snowpark, masking policies, row access, secure sharing, and warehouse-native operational controls.
- ▌ Bigquery And Dataform Platform Engineering · vaquarkhanGuides agents through BigQuery- and Dataform-centered data engineering workflows. Use when designing BigQuery physical models, ingestion boundaries, Dataform transformation workflows, slot or cost controls, and platform decisions across BigQuery, Dataflow, Dataproc, and GCP orchestration services.
- ▌ Data Contract Testing With Schema Registry · vaquarkhanGuides agents through data-contract testing using schema registries and compatibility checks. Use when validating event contracts, stream schema evolution, consumer compatibility, or release gates for schema-managed systems.
- ▌ Data Platform CI CD And Release Management · vaquarkhanGuides agents through CI/CD and release management for data platforms. Use when promoting pipeline code, SQL models, contracts, infra, or configuration across environments with validation gates, staged rollout, and rollback awareness.
- ▌ Data Quality Platforms And Rule Management · vaquarkhanGuides agents through data-quality operating models and tool selection. Use when designing rule portfolios, severity levels, ownership, evidence, and enforcement across dbt tests, Great Expectations, Deequ, Cuallee, Soda, warehouse-native checks, and platform monitoring workflows.
- ▌ Data Reconciliation And Financial Controls · vaquarkhanGuides agents through reconciliation and control design for business-critical data. Use when validating financial, operational, or audit-sensitive metrics with source-to-target totals, control balances, exception tracking, or close-process dependencies.
- ▌ Terraform And Data Platform Infrastructure · vaquarkhanGuides agents through infrastructure as code for data platforms. Use when provisioning or modifying storage, roles, secrets, networking, orchestration resources, catalogs, compute, or environment-specific data platform foundations.
- ▌ Data Security Compliance And Regulated Data · vaquarkhanGuides agents through regulated-data security and compliance workflows for PII, PCI, HIPAA, PHI, and similar obligations. Use when data products handle sensitive fields, regulated records, control evidence, or audit-bound publish paths.
- ▌ Esg And Sustainability Regulatory Reporting · vaquarkhanGuides agents through ESG, sustainability, and regulatory reporting data products. Use when building governed metrics, traceable evidence, and audit-ready data pipelines for frameworks such as CSRD/ESRS, BRSR, climate disclosures, or similar sustainability reporting obligations.
- ▌ Microsoft Purview And Azure Data Governance · vaquarkhanGuides agents through Microsoft Purview and Azure-native data governance workflows. Use when designing collections, scans, classifications, lineage, policy boundaries, and governed publishing across ADLS, Synapse, Data Factory, Azure Databricks, Fabric, and Azure analytics estates.
- ▌ Warehouse Performance And Cost Optimization · vaquarkhanGuides agents through warehouse performance and cost decisions. Use when optimizing BigQuery, Snowflake, Redshift, Athena, Synapse, or lakehouse query patterns, storage layout, and workload isolation.
- ▌ Source Reliability And Extraction Resilience · vaquarkhanGuides agents through source reliability and extraction resilience. Use when upstream systems are flaky, slow, rate-limited, late, or operationally unreliable and ingestion behavior must remain safe and observable.
- ▌ Data Resiliency Testing And Failure Injection · vaquarkhanGuides agents through resiliency testing for data platforms. Use when designing or running failure drills, recovery validation, failover tests, replay-safety checks, dependency outage exercises, or fault injection for pipelines and publishes.
- ▌ Java Data Engineering And Integration Services · vaquarkhanGuides agents through Java-based data engineering services and processors. Use when building connectors, ingestion services, stream processors, metadata services, JVM batch tools, or operational integrations in Java.
- ▌ Lower Environment Data Masking And Obfuscation · vaquarkhanGuides agents through masking, obfuscating, and safely promoting production-like data into lower environments. Use when QA, development, or staging needs realistic data without exposing production-sensitive values.
- ▌ Python Data Engineering And Pipeline Packaging · vaquarkhanGuides agents through Python-based data engineering implementation. Use when building or modifying Python ingestion jobs, orchestration helpers, PySpark entry points, validation code, packaging, dependency management, or operational CLI workflows.
- ▌ Glue Data Catalog And Lake Formation Governance · vaquarkhanGuides agents through AWS-native data catalog and lake governance workflows. Use when designing or reviewing Glue Data Catalog, Lake Formation permissions, governed sharing, metadata quality, and access boundaries for S3, Athena, Redshift, EMR, or Glue pipelines.
- ▌ Enterprise Etl And Data Integration Modernization · vaquarkhanGuides agents through operating, hardening, and modernizing enterprise ETL and integration stacks such as Informatica, Talend, DataStage, SSIS, and Matillion. Use when legacy mappings, job orchestration, migration, or coexistence with modern lakehouse patterns must be handled safely.
- ▌ Spark Serverless Reliability And State Management · vaquarkhanEnforces timeout-aware rollbacks, resumable checkpoints, and orphan cleanup for serverless Spark workloads on AWS Lambda, Glue, and similar runtimes. Use when writing or reviewing Spark jobs in serverless environments, S3 checkpoint patterns, partial-failure recovery, or IceGuard-style state management.