Skill: AI data analyst
Purpose
Perform comprehensive data analysis, statistical modeling, and data visualization by writing and executing self-contained Python scripts. Generate publication-quality charts, statistical reports, and actionable insights from data files or databases.
When to use this skill
- You need to analyze datasets to understand patterns, trends, or relationships.
- You want to perform statistical tests or build predictive models.
- You need data visualizations (charts, graphs, dashboards) to communicate findings.
- You're doing exploratory data analysis (EDA) to understand data structure and quality.
- You need to clean, transform, or merge datasets for analysis.
- You want reproducible analysis with documented methodology and code.
- You are performing Convex Backend Engineering (schema design, query optimization, log analysis).
Key capabilities
Unlike point-solution data analysis tools:
- Convex Engineering Integration: Native support for Convex MCP tools (
mcp_convex) and CLI.
- Full Python ecosystem: Access to pandas, numpy, scikit-learn, statsmodels, matplotlib, seaborn, plotly, and more.
- Runs locally: Your data stays on your machine; no uploads to third-party services.
- Reproducible: All analysis is code-based and version controllable.
- Customizable: Extend with any Python library or custom analysis logic.
- Publication-quality output: Generate professional charts and reports.
- Statistical rigor: Access to comprehensive statistical and ML libraries.
Inputs
- Data sources: CSV files, Excel files, JSON, Parquet, or database connections.
- Analysis goals: Questions to answer or hypotheses to test.
- Variables of interest: Specific columns, metrics, or dimensions to focus on.
- Output preferences: Chart types, report format, statistical tests needed.
- Context: Business domain, data dictionary, or known data quality issues.
Out of scope
- Real-time streaming data analysis (use appropriate streaming tools).
- Extremely large datasets requiring distributed computing (use Spark/Dask instead).
- Production ML model deployment (use ML ops tools and infrastructure).
- Live dashboarding (use BI tools like Tableau/Looker for operational dashboards).
Conventions and best practices
Python environment
- Use virtual environments to isolate dependencies.
- Install only necessary packages for the specific analysis.
- Document all dependencies in
requirements.txt or environment.yml.
Code structure
- Write self-contained scripts that can be re-run by others.
- Use clear variable names and add comments for complex logic.
- Separate concerns: data loading, cleaning, analysis, visualization.
- Save intermediate results to files when analysis is multi-stage.
Data handling
- Never modify source data files – work on copies or in-memory dataframes.
- Document data transformations clearly in code comments.
- Handle missing values explicitly and document approach.
- Validate data quality before analysis (check for nulls, outliers, duplicates).
Visualization best practices
- Choose appropriate chart types for the data and question.
- Use clear labels, titles, and legends on all charts.
- Apply appropriate color schemes (colorblind-friendly when possible).
- Include sample sizes and confidence intervals where relevant.
- Save visualizations in high-resolution formats (PNG 300 DPI, SVG for vector graphics).
Statistical analysis
- State assumptions for statistical tests clearly.
- Check assumptions before applying tests (normality, homoscedasticity, etc.).
- Report effect sizes not just p-values.
- Use appropriate corrections for multiple comparisons.
- Explain practical significance in addition to statistical significance.
Required behavior
- Understand the question: Clarify what insights or decisions the analysis should support.
- Explore the data: Check structure, types, missing values, distributions, outliers.
- Clean and prepare: Handle missing data, outliers, and transformations appropriately.
- Analyze systematically: Apply appropriate statistical methods or ML techniques.
- Visualize effectively: Create clear, informative charts that answer the question.
- Generate insights: Translate statistical findings into actionable business insights.
- Document thoroughly: Explain methodology, assumptions, limitations, and conclusions.
- Make reproducible: Ensure others can re-run the analysis and get the same results.
Required artifacts
- Analysis script(s): Well-documented Python code performing the analysis.
- Visualizations: Charts saved as high-quality image files (PNG/SVG).
- Analysis report: Markdown or text document summarizing:
- Research question and methodology
- Data description and quality assessment
- Key findings with supporting statistics
- Visualizations with interpretations
- Limitations and caveats
- Recommendations or next steps
- Requirements file:
requirements.txt with all dependencies.
- Sample data (if appropriate and non-sensitive): Small sample for reproducibility.
Implementation checklist
1. Data exploration and preparation
2. Data cleaning and transformation
3. Analysis execution
4. Visualization
5. Reporting
6. Reproducibility
Convex Engineering Workflow
When working with Convex (backend, database, schemas), you MUST follow this specialized workflow:
1. Protocols & Rules
- READ FIRST: Always read
resources/convex_rules.md before writing any Convex code.
- Command:
view_file(AbsolutePath=".../resources/convex_rules.md")
- MCP Integration: Use
mcp_convex tools to inspect CURRENT state before proposing changes.
mcp_convex_tables: Check table schemas.
mcp_convex_functionSpec: Check existing functions.
mcp_convex_logs: Analyze recent failures.
2. Implementation & fix
- CLI First: Use
bunx convex for all operations.
- DO NOT use generic SQL or other DB commands.
- Example:
bunx convex run serena/actions:doSomething
- Log Analysis:
- When debugging, pull logs via
bunx convex logs --prod --failure OR mcp_convex_logs.
- Analyze stack traces using Python scripts if text analysis is insufficient.
3. Code Generation
- Schema: Define in
convex/schema.ts using defineSchema and defineTable.
- Functions: Use
query, mutation, action from _generated/server.
- Validation: Ensure
args and returns validators (e.g., v.string(), v.id()) are strictly typed.
Verification
Run the following to verify the analysis:
# Create virtual environment
python3 -m venv venv
source venv/bin/activate # or `venv\Scripts\activate` on Windows
# Install dependencies
pip install -r requirements.txt
# Run analysis script
python analysis.py
# Check outputs generated
ls -lh outputs/
The skill is complete when:
- Analysis script runs without errors from clean environment.
- All required visualizations are generated in high quality.
- Report clearly explains methodology, findings, and limitations.
- Results are interpretable and actionable.
- Code is well-documented and reproducible.
Common analysis patterns
Exploratory Data Analysis (EDA)
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
# Load and inspect data
df = pd.read_csv('data.csv')
print(df.info())
print(df.describe())
# Check for missing values
print(df.isnull().sum())
# Visualize distributions
df.hist(figsize=(12, 10), bins=30)
plt.tight_layout()
plt.savefig('distributions.png', dpi=300)
# Check correlations
corr = df.corr()
sns.heatmap(corr, annot=True, cmap='coolwarm')
plt.savefig('correlations.png', dpi=300)
Time series analysis
import pandas as pd
import matplotlib.pyplot as plt
from statsmodels.tsa.seasonal import seasonal_decompose
# Load time series data
df = pd.read_csv('timeseries.csv', parse_dates=['date'])
df.set_index('date', inplace=True)
# Decompose time series
decomposition = seasonal_decompose(df['value'], model='additive', period=30)
fig = decomposition.plot()
fig.set_size_inches(12, 8)
plt.savefig('decomposition.png', dpi=300)
# Calculate rolling statistics
df['rolling_mean'] = df['value'].rolling(window=7).mean()
df['rolling_std'] = df['value'].rolling(window=7).std()
# Plot with trends
plt.figure(figsize=(12, 6))
plt.plot(df['value'], label='Original')
plt.plot(df['rolling_mean'], label='7-day Moving Avg', linewidth=2)
plt.fill_between(df.index,
df['rolling_mean'] - df['rolling_std'],
df['rolling_mean'] + df['rolling_std'],
alpha=0.3)
plt.legend()
plt.savefig('trends.png', dpi=300)
Statistical hypothesis testing
from scipy import stats
import numpy as np
# Compare two groups
group_a = df[df['group'] == 'A']['metric']
group_b = df[df['group'] == 'B']['metric']
# Check normality
_, p_norm_a = stats.shapiro(group_a)
_, p_norm_b = stats.shapiro(group_b)
# Choose appropriate test
if p_norm_a > 0.05 and p_norm_b > 0.05:
# Parametric test (t-test)
statistic, p_value = stats.ttest_ind(group_a, group_b)
test_used = "Independent t-test"
else:
# Non-parametric test (Mann-Whitney U)
statistic, p_value = stats.mannwhitneyu(group_a, group_b)
test_used = "Mann-Whitney U test"
# Calculate effect size (Cohen's d)
pooled_std = np.sqrt((group_a.std()**2 + group_b.std()**2) / 2)
cohens_d = (group_a.mean() - group_b.mean()) / pooled_std
print(f"Test used: {test_used}")
print(f"Test statistic: {statistic:.4f}")
print(f"P-value: {p_value:.4f}")
print(f"Effect size (Cohen's d): {cohens_d:.4f}")
Predictive modeling
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_squared_error, r2_score
import matplotlib.pyplot as plt
# Prepare data
X = df.drop('target', axis=1)
y = df['target']
# Split data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
# Train model
model = RandomForestRegressor(n_estimators=100, random_state=42)
model.fit(X_train, y_train)
# Evaluate
y_pred = model.predict(X_test)
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
r2 = r2_score(y_test, y_pred)
print(f"RMSE: {rmse:.4f}")
print(f"R² Score: {r2:.4f}")
# Feature importance
importance = pd.DataFrame({
'feature': X.columns,
'importance': model.feature_importances_
}).sort_values('importance', ascending=False)
plt.figure(figsize=(10, 6))
plt.barh(importance['feature'][:10], importance['importance'][:10])
plt.xlabel('Feature Importance')
plt.title('Top 10 Most Important Features')
plt.tight_layout()
plt.savefig('feature_importance.png', dpi=300)
Recommended Python libraries
Data manipulation
- pandas: Data manipulation and analysis
- numpy: Numerical computing
- polars: High-performance DataFrame library (alternative to pandas)
Visualization
- matplotlib: Foundational plotting library
- seaborn: Statistical visualizations
- plotly: Interactive charts
- altair: Declarative statistical visualization
Statistical analysis
- scipy.stats: Statistical functions and tests
- statsmodels: Statistical modeling
- pingouin: Statistical tests with clear output
Machine learning
- scikit-learn: ML algorithms and tools
- xgboost: Gradient boosting
- lightgbm: Fast gradient boosting
Time series
- statsmodels.tsa: Time series analysis
- prophet: Forecasting tool
- pmdarima: Auto ARIMA
Specialized
- networkx: Network analysis
- geopandas: Geospatial data analysis
- textblob / spacy: Natural language processing
Safety and escalation
- Data privacy: Never analyze or share data containing PII without proper authorization.
- Statistical validity: If sample sizes are too small for reliable inference, call this out explicitly.
- Causal claims: Avoid implying causation from correlational analysis; be explicit about limitations.
- Model limitations: Document when models may not generalize or when predictions should not be trusted.
- Data quality: If data quality issues could materially affect conclusions, flag this prominently.
Integration with other skills
This skill can be combined with:
- Internal data querying: To fetch data from warehouses or databases for analysis.
- Web app builder: To create interactive dashboards displaying analysis results.
- Internal tools: To build analysis tools for non-technical stakeholders.
1---2name: ai-data-analyst3description: Perform comprehensive data analysis, statistical modeling, and data visualization by writing and executing self-contained Python scripts. Use when you need to analyze datasets, perform statistical tests, create visualizations, or build predictive models with reproducible, code-based workflows.4---5# Skill: AI data analyst67## Purpose89Perform comprehensive data analysis, statistical modeling, and data visualization by writing and executing self-contained Python scripts. Generate publication-quality charts, statistical reports, and actionable insights from data files or databases.1011## When to use this skill1213- You need to **analyze datasets** to understand patterns, trends, or relationships.14- You want to perform **statistical tests** or build predictive models.15- You need **data visualizations** (charts, graphs, dashboards) to communicate findings.16- You're doing **exploratory data analysis** (EDA) to understand data structure and quality.17- You need to **clean, transform, or merge** datasets for analysis.18- You want **reproducible analysis** with documented methodology and code.19- You are performing **Convex Backend Engineering** (schema design, query optimization, log analysis).2021## Key capabilities2223Unlike point-solution data analysis tools:2425- **Convex Engineering Integration**: Native support for Convex MCP tools (`mcp_convex`) and CLI.26- **Full Python ecosystem**: Access to pandas, numpy, scikit-learn, statsmodels, matplotlib, seaborn, plotly, and more.27- **Runs locally**: Your data stays on your machine; no uploads to third-party services.28- **Reproducible**: All analysis is code-based and version controllable.29- **Customizable**: Extend with any Python library or custom analysis logic.30- **Publication-quality output**: Generate professional charts and reports.31- **Statistical rigor**: Access to comprehensive statistical and ML libraries.3233## Inputs3435- **Data sources**: CSV files, Excel files, JSON, Parquet, or database connections.36- **Analysis goals**: Questions to answer or hypotheses to test.37- **Variables of interest**: Specific columns, metrics, or dimensions to focus on.38- **Output preferences**: Chart types, report format, statistical tests needed.39- **Context**: Business domain, data dictionary, or known data quality issues.4041## Out of scope4243- Real-time streaming data analysis (use appropriate streaming tools).44- Extremely large datasets requiring distributed computing (use Spark/Dask instead).45- Production ML model deployment (use ML ops tools and infrastructure).46- Live dashboarding (use BI tools like Tableau/Looker for operational dashboards).4748## Conventions and best practices4950### Python environment51- Use **virtual environments** to isolate dependencies.52- Install only necessary packages for the specific analysis.53- Document all dependencies in `requirements.txt` or `environment.yml`.5455### Code structure56- Write **self-contained scripts** that can be re-run by others.57- Use **clear variable names** and add comments for complex logic.58- **Separate concerns**: data loading, cleaning, analysis, visualization.59- Save **intermediate results** to files when analysis is multi-stage.6061### Data handling62- **Never modify source data files** – work on copies or in-memory dataframes.63- **Document data transformations** clearly in code comments.64- **Handle missing values** explicitly and document approach.65- **Validate data quality** before analysis (check for nulls, outliers, duplicates).6667### Visualization best practices68- Choose **appropriate chart types** for the data and question.69- Use **clear labels, titles, and legends** on all charts.70- Apply **appropriate color schemes** (colorblind-friendly when possible).71- Include **sample sizes and confidence intervals** where relevant.72- Save visualizations in **high-resolution formats** (PNG 300 DPI, SVG for vector graphics).7374### Statistical analysis75- **State assumptions** for statistical tests clearly.76- **Check assumptions** before applying tests (normality, homoscedasticity, etc.).77- **Report effect sizes** not just p-values.78- **Use appropriate corrections** for multiple comparisons.79- **Explain practical significance** in addition to statistical significance.8081## Required behavior82831. **Understand the question**: Clarify what insights or decisions the analysis should support.842. **Explore the data**: Check structure, types, missing values, distributions, outliers.853. **Clean and prepare**: Handle missing data, outliers, and transformations appropriately.864. **Analyze systematically**: Apply appropriate statistical methods or ML techniques.875. **Visualize effectively**: Create clear, informative charts that answer the question.886. **Generate insights**: Translate statistical findings into actionable business insights.897. **Document thoroughly**: Explain methodology, assumptions, limitations, and conclusions.908. **Make reproducible**: Ensure others can re-run the analysis and get the same results.9192## Required artifacts9394- **Analysis script(s)**: Well-documented Python code performing the analysis.95- **Visualizations**: Charts saved as high-quality image files (PNG/SVG).96- **Analysis report**: Markdown or text document summarizing:97 - Research question and methodology98 - Data description and quality assessment99 - Key findings with supporting statistics100 - Visualizations with interpretations101 - Limitations and caveats102 - Recommendations or next steps103- **Requirements file**: `requirements.txt` with all dependencies.104- **Sample data** (if appropriate and non-sensitive): Small sample for reproducibility.105106## Implementation checklist107108### 1. Data exploration and preparation109- [ ] Load data and inspect structure (shape, columns, types)110- [ ] Check for missing values, duplicates, outliers111- [ ] Generate summary statistics (mean, median, std, min, max)112- [ ] Visualize distributions of key variables113- [ ] Document data quality issues found114115### 2. Data cleaning and transformation116- [ ] Handle missing values (impute, drop, or flag)117- [ ] Address outliers if needed (cap, transform, or document)118- [ ] Create derived variables if needed119- [ ] Normalize or scale variables for modeling120- [ ] Split data if doing train/test analysis121122### 3. Analysis execution123- [ ] Choose appropriate analytical methods124- [ ] Check statistical assumptions125- [ ] Execute analysis with proper parameters126- [ ] Calculate confidence intervals and effect sizes127- [ ] Perform sensitivity analyses if appropriate128129### 4. Visualization130- [ ] Create exploratory visualizations131- [ ] Generate publication-quality final charts132- [ ] Ensure all charts have clear labels and titles133- [ ] Use appropriate color schemes and styling134- [ ] Save in high-resolution formats135136### 5. Reporting137- [ ] Write clear summary of methods used138- [ ] Present key findings with supporting evidence139- [ ] Explain practical significance of results140- [ ] Document limitations and assumptions141- [ ] Provide actionable recommendations142143### 6. Reproducibility144- [ ] Test that script runs from clean environment145- [ ] Document all dependencies146- [ ] Add comments explaining non-obvious code147- [ ] Include instructions for running analysis148149## Convex Engineering Workflow150151When working with Convex (backend, database, schemas), you **MUST** follow this specialized workflow:152153### 1. Protocols & Rules154- **READ FIRST**: Always read `resources/convex_rules.md` before writing any Convex code.155 - Command: `view_file(AbsolutePath=".../resources/convex_rules.md")`156- **MCP Integration**: Use `mcp_convex` tools to inspect CURRENT state before proposing changes.157 - `mcp_convex_tables`: Check table schemas.158 - `mcp_convex_functionSpec`: Check existing functions.159 - `mcp_convex_logs`: Analyze recent failures.160161### 2. Implementation & fix162- **CLI First**: Use `bunx convex` for all operations.163 - DO NOT use generic SQL or other DB commands.164 - Example: `bunx convex run serena/actions:doSomething`165- **Log Analysis**:166 - When debugging, pull logs via `bunx convex logs --prod --failure` OR `mcp_convex_logs`.167 - Analyze stack traces using Python scripts if text analysis is insufficient.168169### 3. Code Generation170- **Schema**: Define in `convex/schema.ts` using `defineSchema` and `defineTable`.171- **Functions**: Use `query`, `mutation`, `action` from `_generated/server`.172- **Validation**: Ensure `args` and `returns` validators (e.g., `v.string()`, `v.id()`) are strictly typed.173174175## Verification176177Run the following to verify the analysis:178179```bash180# Create virtual environment181python3 -m venv venv182source venv/bin/activate # or `venv\Scripts\activate` on Windows183184# Install dependencies185pip install -r requirements.txt186187# Run analysis script188python analysis.py189190# Check outputs generated191ls -lh outputs/192```193194The skill is complete when:195196- Analysis script runs without errors from clean environment.197- All required visualizations are generated in high quality.198- Report clearly explains methodology, findings, and limitations.199- Results are interpretable and actionable.200- Code is well-documented and reproducible.201202## Common analysis patterns203204### Exploratory Data Analysis (EDA)205```python206import pandas as pd207import matplotlib.pyplot as plt208import seaborn as sns209210# Load and inspect data211df = pd.read_csv('data.csv')212print(df.info())213print(df.describe())214215# Check for missing values216print(df.isnull().sum())217218# Visualize distributions219df.hist(figsize=(12, 10), bins=30)220plt.tight_layout()221plt.savefig('distributions.png', dpi=300)222223# Check correlations224corr = df.corr()225sns.heatmap(corr, annot=True, cmap='coolwarm')226plt.savefig('correlations.png', dpi=300)227```228229### Time series analysis230```python231import pandas as pd232import matplotlib.pyplot as plt233from statsmodels.tsa.seasonal import seasonal_decompose234235# Load time series data236df = pd.read_csv('timeseries.csv', parse_dates=['date'])237df.set_index('date', inplace=True)238239# Decompose time series240decomposition = seasonal_decompose(df['value'], model='additive', period=30)241fig = decomposition.plot()242fig.set_size_inches(12, 8)243plt.savefig('decomposition.png', dpi=300)244245# Calculate rolling statistics246df['rolling_mean'] = df['value'].rolling(window=7).mean()247df['rolling_std'] = df['value'].rolling(window=7).std()248249# Plot with trends250plt.figure(figsize=(12, 6))251plt.plot(df['value'], label='Original')252plt.plot(df['rolling_mean'], label='7-day Moving Avg', linewidth=2)253plt.fill_between(df.index,254 df['rolling_mean'] - df['rolling_std'],255 df['rolling_mean'] + df['rolling_std'],256 alpha=0.3)257plt.legend()258plt.savefig('trends.png', dpi=300)259```260261### Statistical hypothesis testing262```python263from scipy import stats264import numpy as np265266# Compare two groups267group_a = df[df['group'] == 'A']['metric']268group_b = df[df['group'] == 'B']['metric']269270# Check normality271_, p_norm_a = stats.shapiro(group_a)272_, p_norm_b = stats.shapiro(group_b)273274# Choose appropriate test275if p_norm_a > 0.05 and p_norm_b > 0.05:276 # Parametric test (t-test)277 statistic, p_value = stats.ttest_ind(group_a, group_b)278 test_used = "Independent t-test"279else:280 # Non-parametric test (Mann-Whitney U)281 statistic, p_value = stats.mannwhitneyu(group_a, group_b)282 test_used = "Mann-Whitney U test"283284# Calculate effect size (Cohen's d)285pooled_std = np.sqrt((group_a.std()**2 + group_b.std()**2) / 2)286cohens_d = (group_a.mean() - group_b.mean()) / pooled_std287288print(f"Test used: {test_used}")289print(f"Test statistic: {statistic:.4f}")290print(f"P-value: {p_value:.4f}")291print(f"Effect size (Cohen's d): {cohens_d:.4f}")292```293294### Predictive modeling295```python296from sklearn.model_selection import train_test_split297from sklearn.ensemble import RandomForestRegressor298from sklearn.metrics import mean_squared_error, r2_score299import matplotlib.pyplot as plt300301# Prepare data302X = df.drop('target', axis=1)303y = df['target']304305# Split data306X_train, X_test, y_train, y_test = train_test_split(307 X, y, test_size=0.2, random_state=42308)309310# Train model311model = RandomForestRegressor(n_estimators=100, random_state=42)312model.fit(X_train, y_train)313314# Evaluate315y_pred = model.predict(X_test)316rmse = np.sqrt(mean_squared_error(y_test, y_pred))317r2 = r2_score(y_test, y_pred)318319print(f"RMSE: {rmse:.4f}")320print(f"R² Score: {r2:.4f}")321322# Feature importance323importance = pd.DataFrame({324 'feature': X.columns,325 'importance': model.feature_importances_326}).sort_values('importance', ascending=False)327328plt.figure(figsize=(10, 6))329plt.barh(importance['feature'][:10], importance['importance'][:10])330plt.xlabel('Feature Importance')331plt.title('Top 10 Most Important Features')332plt.tight_layout()333plt.savefig('feature_importance.png', dpi=300)334```335336## Recommended Python libraries337338### Data manipulation339- **pandas**: Data manipulation and analysis340- **numpy**: Numerical computing341- **polars**: High-performance DataFrame library (alternative to pandas)342343### Visualization344- **matplotlib**: Foundational plotting library345- **seaborn**: Statistical visualizations346- **plotly**: Interactive charts347- **altair**: Declarative statistical visualization348349### Statistical analysis350- **scipy.stats**: Statistical functions and tests351- **statsmodels**: Statistical modeling352- **pingouin**: Statistical tests with clear output353354### Machine learning355- **scikit-learn**: ML algorithms and tools356- **xgboost**: Gradient boosting357- **lightgbm**: Fast gradient boosting358359### Time series360- **statsmodels.tsa**: Time series analysis361- **prophet**: Forecasting tool362- **pmdarima**: Auto ARIMA363364### Specialized365- **networkx**: Network analysis366- **geopandas**: Geospatial data analysis367- **textblob** / **spacy**: Natural language processing368369## Safety and escalation370371- **Data privacy**: Never analyze or share data containing PII without proper authorization.372- **Statistical validity**: If sample sizes are too small for reliable inference, call this out explicitly.373- **Causal claims**: Avoid implying causation from correlational analysis; be explicit about limitations.374- **Model limitations**: Document when models may not generalize or when predictions should not be trusted.375- **Data quality**: If data quality issues could materially affect conclusions, flag this prominently.376377## Integration with other skills378379This skill can be combined with:380381- **Internal data querying**: To fetch data from warehouses or databases for analysis.382- **Web app builder**: To create interactive dashboards displaying analysis results.383- **Internal tools**: To build analysis tools for non-technical stakeholders.