Data Analyzer
Capture-insights adapter (no KG)
Run from Agent Skills or chat ("analyze this dataset", "find patterns", "trend analysis").
- Deterministic EDA: profiles CSV/JSON in pursuit Studio +
competitive_intel_obligation.json when present.
- Market slice: DuckDB
usaspending_prime_awards summary for active NAICS (top recipients/agencies).
- Workflows: EDA (default), trend, comparative, anomaly — inferred from inquiry phrasing.
- Outputs:
02_intel/data_analysis.json + .md; optional LLM narrative when smart model on.
- No invented stats — every number traces to a file path or DuckDB query.
Expert data analysis agent that processes structured and unstructured datasets to extract
meaningful insights, identify patterns, detect anomalies, and generate data-driven
recommendations.
This is a domain-agnostic platform primitive. Let the data determine the analytical
approach — do not impose a domain lens. Govcon-specific variants (e.g., workload-analyzer)
build on this foundation via skill composition.
References (load when needed):
references/statistical-methods.md — Descriptive + inferential statistics, effect size measures
references/visualization-guide.md — Chart selection guide, design principles
references/pitfalls.md — Common analytical pitfalls and how to avoid them
assets/report_template.md — Report structure template for output
Workflow 1: Exploratory Data Analysis (EDA)
Objective: Understand dataset structure, quality, and preliminary patterns.
- Data Profiling — dimensions, column types, completeness, cardinality, ranges, summary statistics (mean, median, mode, std dev)
- Data Quality Assessment — missing data patterns (MCAR/MAR/MNAR), duplicates, outliers, consistency issues; document each with severity rating
- Univariate Analysis — distribution per variable, skewness/kurtosis, outlier detection (IQR, Z-score)
- Bivariate Analysis — correlations (Pearson, Spearman), scatter plots for continuous pairs, cross-tabulations for categorical pairs
- Multivariate Analysis — correlation matrices, dimensionality assessment, cluster tendency
- Initial Insights — key patterns, surprising findings, hypotheses for further investigation, data limitations
Load references/statistical-methods.md for formula reference and test selection guidance.
Deliverable: EDA report with summary statistics, visualizations, and preliminary insights.
Workflow 2: Pattern Detection & Trend Analysis
Objective: Identify meaningful patterns, trends, and relationships.
- Time Series Analysis (if temporal data) — trend direction, seasonality, cyclical patterns, anomaly detection, decompose into trend/seasonal/residual
- Segmentation Analysis — natural groupings, segment profiling and characterization, cross-metric comparison
- Correlation & Causation — correlation strength and significance, investigation of causal mechanisms, control for confounders; document the distinction explicitly
- Anomaly Detection — statistical outliers, contextual anomalies, point vs. collective anomalies; determine whether each is an error or a finding
- Pattern Validation — stability across subsets, sensitivity analysis, confidence intervals, significance testing
Deliverable: Pattern analysis report with validated findings.
Workflow 3: Statistical Hypothesis Testing
Objective: Rigorously test hypotheses using appropriate statistical methods.
- Hypothesis Formulation — define H0 and H1, specify significance level (α = 0.05 default), identify appropriate test
- Test Selection — means: t-test/ANOVA; proportions: chi-square/Fisher's; correlation: Pearson/Spearman; distribution: K-S/Shapiro-Wilk; load
references/statistical-methods.md for full decision tree
- Assumptions Checking — normality, homogeneity of variance, independence, sample size adequacy; use non-parametric alternatives if assumptions are violated
- Test Execution — compute test statistic, p-value, effect size (Cohen's d, η², R²), confidence intervals
- Result Interpretation — distinguish statistical significance from practical significance; translate to plain-language implications
Deliverable: Statistical test report with methodology, results, and interpretation.
Workflow 4: Comparative Analysis
Objective: Compare groups, segments, or time periods to identify differences and drivers.
- Define Comparison — groups to compare, metrics, baseline vs. target, success criteria
- Segment Performance — key metrics per segment, rank by performance, quantify gaps
- Driver Analysis — factors explaining differences, quantified contribution per driver, confounding controls
- Benchmarking — vs. historical performance, external standards, or best-in-class; calculate gaps to each benchmark
- Recommendations — close performance gaps; distinguish quick wins from strategic initiatives; quantify expected impact
Deliverable: Comparative analysis report with driver identification and action plan.
Workflow 5: Insight Synthesis & Storytelling
Objective: Transform analytical findings into clear, actionable insights.
- Insight Identification — review all findings; identify the "so what"; prioritize by impact; group into themes
- Insight Structuring — for each finding: Observation → Insight → Implication → Recommendation (answer first, then supporting detail)
- Evidence Assembly — key statistics, supporting visualizations, benchmarks, confidence levels
- Narrative Development — clear, jargon-free storyline with logical flow from problem to recommendation; anticipate counterarguments
- Visualization Design — appropriate chart types, annotated key insights; load
references/visualization-guide.md for chart selection
- Actionability — specific actions, owners, timelines, success metrics, quantified expected impact
Deliverable: Executive-ready insight report. Use assets/report_template.md as the output structure.
Quick Reference
| Action |
Trigger phrase |
| Full EDA |
"Analyze this dataset comprehensively" |
| Quick summary |
"Summarize key statistics" |
| Pattern detection |
"Find patterns in this dataset" |
| Hypothesis test |
"Test if [A] affects [B]" |
| Comparative analysis |
"Compare [group A] vs [group B]" |
| Correlation analysis |
"What correlates with [variable]?" |
| Anomaly detection |
"Find anomalies in this data" |
| Trend analysis |
"Analyze trends over time" |
Best Practices
- Start with questions — define what you are trying to learn before examining the data
- Document assumptions — be explicit about data limitations and analytical choices
- Visualize early and often — charts reveal patterns that tables hide
- Communicate uncertainty — use confidence intervals, p-values, and error bars
- Beware spurious correlations — correlation is not causation; see
references/pitfalls.md
- Validate with domain knowledge — ensure insights align with subject-matter expertise
- Iterate — analysis is rarely linear; loop back when findings raise new questions
Pre-Delivery Checklist
1---2name: data-analyzer3description: Advanced data analysis, pattern detection, and insight generation from structured and unstructured datasets. Use when the user wants to analyze data, perform statistical analysis, find insights, detect patterns, identify anomalies, compare segments, test hypotheses, or generate data-driven recommendations. Triggers on phrases like 'analyze data', 'data analysis', 'find insights', 'analyze dataset', 'statistical analysis', 'find patterns', 'compare groups', 'test hypothesis', 'correlation analysis', or 'trend analysis'.4license: MIT5---6# Data Analyzer78## Capture-insights adapter (no KG)910**Run** from Agent Skills or chat ("analyze this dataset", "find patterns", "trend analysis").1112- **Deterministic EDA:** profiles CSV/JSON in pursuit Studio + `competitive_intel_obligation.json` when present.13- **Market slice:** DuckDB `usaspending_prime_awards` summary for active NAICS (top recipients/agencies).14- **Workflows:** EDA (default), trend, comparative, anomaly — inferred from inquiry phrasing.15- **Outputs:** `02_intel/data_analysis.json` + `.md`; optional LLM narrative when smart model on.16- **No invented stats** — every number traces to a file path or DuckDB query.171819Expert data analysis agent that processes structured and unstructured datasets to extract20meaningful insights, identify patterns, detect anomalies, and generate data-driven21recommendations.2223**This is a domain-agnostic platform primitive.** Let the data determine the analytical24approach — do not impose a domain lens. Govcon-specific variants (e.g., `workload-analyzer`)25build on this foundation via skill composition.2627**References (load when needed):**2829- `references/statistical-methods.md` — Descriptive + inferential statistics, effect size measures30- `references/visualization-guide.md` — Chart selection guide, design principles31- `references/pitfalls.md` — Common analytical pitfalls and how to avoid them32- `assets/report_template.md` — Report structure template for output3334---3536## Workflow 1: Exploratory Data Analysis (EDA)3738**Objective:** Understand dataset structure, quality, and preliminary patterns.39401. **Data Profiling** — dimensions, column types, completeness, cardinality, ranges, summary statistics (mean, median, mode, std dev)412. **Data Quality Assessment** — missing data patterns (MCAR/MAR/MNAR), duplicates, outliers, consistency issues; document each with severity rating423. **Univariate Analysis** — distribution per variable, skewness/kurtosis, outlier detection (IQR, Z-score)434. **Bivariate Analysis** — correlations (Pearson, Spearman), scatter plots for continuous pairs, cross-tabulations for categorical pairs445. **Multivariate Analysis** — correlation matrices, dimensionality assessment, cluster tendency456. **Initial Insights** — key patterns, surprising findings, hypotheses for further investigation, data limitations4647> Load `references/statistical-methods.md` for formula reference and test selection guidance.4849**Deliverable:** EDA report with summary statistics, visualizations, and preliminary insights.5051---5253## Workflow 2: Pattern Detection & Trend Analysis5455**Objective:** Identify meaningful patterns, trends, and relationships.56571. **Time Series Analysis** (if temporal data) — trend direction, seasonality, cyclical patterns, anomaly detection, decompose into trend/seasonal/residual582. **Segmentation Analysis** — natural groupings, segment profiling and characterization, cross-metric comparison593. **Correlation & Causation** — correlation strength and significance, investigation of causal mechanisms, control for confounders; document the distinction explicitly604. **Anomaly Detection** — statistical outliers, contextual anomalies, point vs. collective anomalies; determine whether each is an error or a finding615. **Pattern Validation** — stability across subsets, sensitivity analysis, confidence intervals, significance testing6263**Deliverable:** Pattern analysis report with validated findings.6465---6667## Workflow 3: Statistical Hypothesis Testing6869**Objective:** Rigorously test hypotheses using appropriate statistical methods.70711. **Hypothesis Formulation** — define H0 and H1, specify significance level (α = 0.05 default), identify appropriate test722. **Test Selection** — means: t-test/ANOVA; proportions: chi-square/Fisher's; correlation: Pearson/Spearman; distribution: K-S/Shapiro-Wilk; load `references/statistical-methods.md` for full decision tree733. **Assumptions Checking** — normality, homogeneity of variance, independence, sample size adequacy; use non-parametric alternatives if assumptions are violated744. **Test Execution** — compute test statistic, p-value, effect size (Cohen's d, η², R²), confidence intervals755. **Result Interpretation** — distinguish statistical significance from practical significance; translate to plain-language implications7677**Deliverable:** Statistical test report with methodology, results, and interpretation.7879---8081## Workflow 4: Comparative Analysis8283**Objective:** Compare groups, segments, or time periods to identify differences and drivers.84851. **Define Comparison** — groups to compare, metrics, baseline vs. target, success criteria862. **Segment Performance** — key metrics per segment, rank by performance, quantify gaps873. **Driver Analysis** — factors explaining differences, quantified contribution per driver, confounding controls884. **Benchmarking** — vs. historical performance, external standards, or best-in-class; calculate gaps to each benchmark895. **Recommendations** — close performance gaps; distinguish quick wins from strategic initiatives; quantify expected impact9091**Deliverable:** Comparative analysis report with driver identification and action plan.9293---9495## Workflow 5: Insight Synthesis & Storytelling9697**Objective:** Transform analytical findings into clear, actionable insights.98991. **Insight Identification** — review all findings; identify the "so what"; prioritize by impact; group into themes1002. **Insight Structuring** — for each finding: Observation → Insight → Implication → Recommendation (answer first, then supporting detail)1013. **Evidence Assembly** — key statistics, supporting visualizations, benchmarks, confidence levels1024. **Narrative Development** — clear, jargon-free storyline with logical flow from problem to recommendation; anticipate counterarguments1035. **Visualization Design** — appropriate chart types, annotated key insights; load `references/visualization-guide.md` for chart selection1046. **Actionability** — specific actions, owners, timelines, success metrics, quantified expected impact105106**Deliverable:** Executive-ready insight report. Use `assets/report_template.md` as the output structure.107108---109110## Quick Reference111112| Action | Trigger phrase |113| -------------------- | -------------------------------------- |114| Full EDA | "Analyze this dataset comprehensively" |115| Quick summary | "Summarize key statistics" |116| Pattern detection | "Find patterns in this dataset" |117| Hypothesis test | "Test if [A] affects [B]" |118| Comparative analysis | "Compare [group A] vs [group B]" |119| Correlation analysis | "What correlates with [variable]?" |120| Anomaly detection | "Find anomalies in this data" |121| Trend analysis | "Analyze trends over time" |122123---124125## Best Practices126127- **Start with questions** — define what you are trying to learn before examining the data128- **Document assumptions** — be explicit about data limitations and analytical choices129- **Visualize early and often** — charts reveal patterns that tables hide130- **Communicate uncertainty** — use confidence intervals, p-values, and error bars131- **Beware spurious correlations** — correlation is not causation; see `references/pitfalls.md`132- **Validate with domain knowledge** — ensure insights align with subject-matter expertise133- **Iterate** — analysis is rarely linear; loop back when findings raise new questions134135## Pre-Delivery Checklist136137- [ ] Data quality assessed and documented138- [ ] Appropriate statistical tests selected and executed139- [ ] Test assumptions verified140- [ ] Statistical vs. practical significance both addressed141- [ ] Visualizations are clear, accurate, and annotated142- [ ] Limitations and caveats explicitly stated143- [ ] Recommendations are specific and quantifiable144- [ ] Report structured using `assets/report_template.md`