When to Use
User needs to: extract data from sources (databases, APIs, files), clean and transform messy datasets, analyze and find patterns, visualize results, or automate recurring data tasks. Agent handles the full data workflow.
Quick Reference
| Area |
File |
Focus |
| Querying & Extraction |
querying.md |
SQL generation, API fetching, multi-source |
| Cleaning & Transformation |
cleaning.md |
Nulls, duplicates, normalization, joins |
| Analysis & Statistics |
analysis.md |
EDA, statistical tests, insights |
| Visualization & Reporting |
visualization.md |
Charts, dashboards, exports |
| Quality & Validation |
quality.md |
Data checks, anomaly detection, drift |
| Workflow Patterns |
patterns.md |
Common data workflows, automation |
Core Operations
Query generation: User describes what data they need → Agent writes SQL/query, handles joins, filters, aggregations → Returns results or explains execution plan.
Data cleaning: Load messy dataset → Detect issues (nulls, duplicates, outliers, inconsistent formats) → Apply appropriate fixes → Document transformations.
Exploratory analysis: New dataset arrives → Generate descriptive stats, distributions, correlations → Surface interesting patterns and anomalies → Produce summary with key findings.
Visualization: Analysis complete → Generate appropriate chart type → Export in requested format (PNG, SVG, interactive HTML) → Ready for stakeholders.
Recurring reports: Define report once → Agent runs on schedule → Updates charts and metrics → Delivers summary with highlights.
Critical Rules
- Always preview transformations before applying — show sample of what will change
- Document every data transformation with source, operation, and rationale
- Validate data types and ranges before analysis — garbage in, garbage out
- Use appropriate statistical tests — check assumptions first
- Generate reproducible outputs — include seeds, versions, timestamps
- Handle missing data explicitly — document chosen strategy (drop, impute, flag)
- Match chart type to data type — categorical, continuous, time series
User Modes
| Mode |
Focus |
Trigger |
| Analyst |
SQL, exploration, insights |
"What does this data tell us?" |
| Engineer |
Pipelines, transformations, quality |
"Clean this and load it there" |
| Business |
KPIs, dashboards, plain language |
"How are we doing vs last quarter?" |
| Researcher |
Statistical rigor, reproducibility |
"Is this difference significant?" |
| Developer |
Schema design, API data, types |
"Generate types from this JSON" |
See patterns.md for workflows per mode.
On First Use
- Identify data source (database, file, API)
- Establish connection or load file
- Initial EDA — shape, types, quality issues
- Clean and transform as needed
- Analyze or visualize per user goal
1---2name: data3description: Work with data across the full lifecycle from extraction and cleaning to analysis, visualization, and reporting.4---5
6## When to Use
7
8User needs to: extract data from sources (databases, APIs, files), clean and transform messy datasets, analyze and find patterns, visualize results, or automate recurring data tasks. Agent handles the full data workflow.
9
10## Quick Reference
11
12| Area | File | Focus |
13|------|------|-------|
14| Querying & Extraction | `querying.md` | SQL generation, API fetching, multi-source |
15| Cleaning & Transformation | `cleaning.md` | Nulls, duplicates, normalization, joins |
16| Analysis & Statistics | `analysis.md` | EDA, statistical tests, insights |
17| Visualization & Reporting | `visualization.md` | Charts, dashboards, exports |
18| Quality & Validation | `quality.md` | Data checks, anomaly detection, drift |
19| Workflow Patterns | `patterns.md` | Common data workflows, automation |
20
21## Core Operations
22
23**Query generation:** User describes what data they need → Agent writes SQL/query, handles joins, filters, aggregations → Returns results or explains execution plan.
24
25**Data cleaning:** Load messy dataset → Detect issues (nulls, duplicates, outliers, inconsistent formats) → Apply appropriate fixes → Document transformations.
26
27**Exploratory analysis:** New dataset arrives → Generate descriptive stats, distributions, correlations → Surface interesting patterns and anomalies → Produce summary with key findings.
28
29**Visualization:** Analysis complete → Generate appropriate chart type → Export in requested format (PNG, SVG, interactive HTML) → Ready for stakeholders.
30
31**Recurring reports:** Define report once → Agent runs on schedule → Updates charts and metrics → Delivers summary with highlights.
32
33## Critical Rules
34
35- Always preview transformations before applying — show sample of what will change
36- Document every data transformation with source, operation, and rationale
37- Validate data types and ranges before analysis — garbage in, garbage out
38- Use appropriate statistical tests — check assumptions first
39- Generate reproducible outputs — include seeds, versions, timestamps
40- Handle missing data explicitly — document chosen strategy (drop, impute, flag)
41- Match chart type to data type — categorical, continuous, time series
42
43## User Modes
44
45| Mode | Focus | Trigger |
46|------|-------|---------|
47| Analyst | SQL, exploration, insights | "What does this data tell us?" |
48| Engineer | Pipelines, transformations, quality | "Clean this and load it there" |
49| Business | KPIs, dashboards, plain language | "How are we doing vs last quarter?" |
50| Researcher | Statistical rigor, reproducibility | "Is this difference significant?" |
51| Developer | Schema design, API data, types | "Generate types from this JSON" |
52
53See `patterns.md` for workflows per mode.
54
55## On First Use
56
571. Identify data source (database, file, API)
582. Establish connection or load file
593. Initial EDA — shape, types, quality issues
604. Clean and transform as needed
615. Analyze or visualize per user goal