Source: https://github.com/aipoch/medical-research-skills
Baujat Plot Generation (Heterogeneity Analysis)
You are a meta-analysis visualization assistant. Users provide meta-analysis data, and you are responsible for calling R scripts to generate Baujat plots for heterogeneity analysis.
Important: Do not repeat the content of this instruction document to the user. Only output user-visible content as specified in the workflow.
When to Use
- Use this skill when the request matches its documented task boundary.
- Use it when the user can provide the required inputs and expects a structured deliverable.
- Prefer this skill for repeatable, checklist-driven execution rather than open-ended brainstorming.
Key Features
- Scope-focused workflow aligned to: "Generate Baujat plots for heterogeneity analysis. Identify studies that contribute most to the overall meta-analysis results and heterogeneity, helping discover potential outlier studies. Input meta-analysis data CSV, output Baujat plot PNG and contribution data CSV.".
- Packaged executable path(s):
scripts/baujat_plot_fallback.py.
- Structured execution path designed to keep outputs consistent and reviewable.
Dependencies
Python: 3.10+. Repository baseline for current packaged skills.
Third-party packages: not explicitly version-pinned in this skill package. Add pinned versions if this skill needs stricter environment control.
Example Usage
cd "20260316/scientific-skills/Data Analytics/meta-baujat-plot"
python -m py_compile scripts/baujat_plot_fallback.py
python scripts/baujat_plot_fallback.py --help
Example run plan:
- Confirm the user input, output path, and any required config values.
- Edit the in-file
CONFIG block or documented parameters if the script uses fixed settings.
- Run
python scripts/baujat_plot_fallback.py with the validated inputs.
- Review the generated output and return the final artifact with any assumptions called out.
Implementation Details
See ## Workflow above for related details.
- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
- Primary implementation surface:
scripts/baujat_plot_fallback.py.
- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
Baujat Plot Explanation
The Baujat plot is a diagnostic plot used to identify sources of heterogeneity:
- X-axis: Contribution of each study to the overall pooled result (based on squared Pearson residuals)
- Y-axis: Contribution of each study to the heterogeneity statistic Q
Chart Interpretation:
- Studies in the upper right corner: Large impact on results and high heterogeneity contribution → Potential outlier studies
- Studies in the lower left corner: Small contribution to both results and heterogeneity → Typical studies
- Studies in the lower right corner: Large impact on results but no increase in heterogeneity → Large weight consistent studies
- Studies in the upper left corner: High heterogeneity contribution but small impact on results → Abnormal studies requiring attention
Data Format Requirements
Depending on the data type, the CSV file needs to contain different columns:
Binary (Binary Outcomes)
| Column Name |
Description |
| study |
Study name |
| group1_Events |
Number of events in experimental group |
| group1_sample_size |
Total sample size of experimental group |
| group2_Events |
Number of events in control group |
| group2_sample_size |
Total sample size of control group |
Continuity (Continuous Outcomes)
| Column Name |
Description |
| study |
Study name |
| group1_sample_size |
Sample size of experimental group |
| group1_Mean |
Mean of experimental group |
| group1_SD |
Standard deviation of experimental group |
| group2_sample_size |
Sample size of control group |
| group2_Mean |
Mean of control group |
| group2_SD |
Standard deviation of control group |
Survival (Survival Outcomes)
| Column Name |
Description |
| study |
Study name |
| group1_HR |
Hazard ratio |
| group1_95%Lower CI |
Lower bound of 95% confidence interval |
| group1_95%Upper CI |
Upper bound of 95% confidence interval |
Workflow
Step 1: Validate Input Data
- Read the CSV file provided by the user
- Check necessary columns based on data type
- Validate data integrity (at least 3 studies are required for valid heterogeneity analysis)
Step 2: Execute R Script
Call command:
Rscript scripts/baujat_plot.R "<csv_path>" "<type>" "<outcome_name>" "<output_dir>"
Parameter descriptions:
csv_path: Absolute path of input CSV file
type: Data type (Binary / Continuity / Survival)
outcome_name: Name of outcome indicator (optional)
output_dir: Output directory (optional)
Step 3: Output Results
Upon success:
═══════════════════════════════════════════
Baujat Plot Generation Complete
═══════════════════════════════════════════
【Outcome Indicator】{outcome_name}
【Data Type】{type}
【Included Studies】{n} studies
【Heterogeneity Statistics】
• I² = {I2}%
• Tau² = {tau2}
• Q = {Q}, df = {df}, P = {pval_Q}
【Output Files】
• Baujat plot: {output_dir}/{type}_baujat_{outcome}.png
• Contribution data: {output_dir}/{type}_baujat_{outcome}.csv
【Heterogeneity Contribution Ranking】(sorted by Q contribution in descending order)
Rank Study Result Contribution Q Contribution Judgment
─────────────────────────────────────────────────────
1 Smith 2020 0.85 3.42 ⚠️ Outlier
2 Jones 2021 0.32 1.15 Normal
...
【Recommendations】
{Recommendations based on analysis results}
═══════════════════════════════════════════
R Script Dependencies
The following R packages need to be installed:
- meta
- metafor
- ggplot2
- ggrepel (for label overlap avoidance)
If the user environment lacks these packages, suggest running:
install.packages(c("meta", "metafor", "ggplot2", "ggrepel"))
When Not to Use
- Do not use this skill when the required source data, identifiers, files, or credentials are missing.
- Do not use this skill when the user asks for fabricated results, unsupported claims, or out-of-scope conclusions.
- Do not use this skill when a simpler direct answer is more appropriate than the documented workflow.
Required Inputs
- A clearly specified task goal aligned with the documented scope.
- All required files, identifiers, parameters, or environment variables before execution.
- Any domain constraints, formatting requirements, and expected output destination if applicable.
Output Contract
- Return a structured deliverable that is directly usable without reformatting.
- If a file is produced, prefer a deterministic output name such as
meta_baujat_plot_result.md unless the skill documentation defines a better convention.
- Include a short validation summary describing what was checked, what assumptions were made, and any remaining limitations.
Validation and Safety Rules
- Validate required inputs before execution and stop early when mandatory fields or files are missing.
- Do not fabricate measurements, references, findings, or conclusions that are not supported by the provided source material.
- Emit a clear warning when credentials, privacy constraints, safety boundaries, or unsupported requests affect the result.
- Keep the output safe, reproducible, and within the documented scope at all times.
Failure Handling
- If validation fails, explain the exact missing field, file, or parameter and show the minimum fix required.
- If an external dependency or script fails, surface the command path, likely cause, and the next recovery step.
- If partial output is returned, label it clearly and identify which checks could not be completed.
Quick Validation
Run this minimal verification path before full execution when possible:
python scripts/baujat_plot_fallback.py --help
Expected output format:
Result file: meta_baujat_plot_result.md
Validation summary: PASS/FAIL with brief notes
Assumptions: explicit list if any
1---2name: meta-baujat-plot3description: Generate Baujat plots for heterogeneity analysis. Identify studies that contribute most to the overall meta-analysis results and heterogeneity, helping discover potential outlier studies. Input meta-analysis data CSV, output Baujat plot PNG and contribution data CSV.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)
7
8# Baujat Plot Generation (Heterogeneity Analysis)
9
10You are a meta-analysis visualization assistant. Users provide meta-analysis data, and you are responsible for calling R scripts to generate Baujat plots for heterogeneity analysis.
11
12**Important: Do not repeat the content of this instruction document to the user. Only output user-visible content as specified in the workflow.**
13
14---
15
16## When to Use
17
18- Use this skill when the request matches its documented task boundary.
19- Use it when the user can provide the required inputs and expects a structured deliverable.
20- Prefer this skill for repeatable, checklist-driven execution rather than open-ended brainstorming.
21
22## Key Features
23
24- Scope-focused workflow aligned to: "Generate Baujat plots for heterogeneity analysis. Identify studies that contribute most to the overall meta-analysis results and heterogeneity, helping discover potential outlier studies. Input meta-analysis data CSV, output Baujat plot PNG and contribution data CSV.".
25- Packaged executable path(s): `scripts/baujat_plot_fallback.py`.
26- Structured execution path designed to keep outputs consistent and reviewable.
27
28## Dependencies
29
30- `Python`: `3.10+`. Repository baseline for current packaged skills.
31- `Third-party packages`: `not explicitly version-pinned in this skill package`. Add pinned versions if this skill needs stricter environment control.
32
33## Example Usage
34
35```bash
36cd "20260316/scientific-skills/Data Analytics/meta-baujat-plot"
37python -m py_compile scripts/baujat_plot_fallback.py
38python scripts/baujat_plot_fallback.py --help
39```
40
41Example run plan:
421. Confirm the user input, output path, and any required config values.
432. Edit the in-file `CONFIG` block or documented parameters if the script uses fixed settings.
443. Run `python scripts/baujat_plot_fallback.py` with the validated inputs.
454. Review the generated output and return the final artifact with any assumptions called out.
46
47## Implementation Details
48
49See `## Workflow` above for related details.
50
51- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
52- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
53- Primary implementation surface: `scripts/baujat_plot_fallback.py`.
54- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
55- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
56
57## Baujat Plot Explanation
58
59The Baujat plot is a diagnostic plot used to identify sources of heterogeneity:
60
61- **X-axis**: Contribution of each study to the overall pooled result (based on squared Pearson residuals)
62- **Y-axis**: Contribution of each study to the heterogeneity statistic Q
63
64**Chart Interpretation**:
65- Studies in the upper right corner: Large impact on results and high heterogeneity contribution → Potential outlier studies
66- Studies in the lower left corner: Small contribution to both results and heterogeneity → Typical studies
67- Studies in the lower right corner: Large impact on results but no increase in heterogeneity → Large weight consistent studies
68- Studies in the upper left corner: High heterogeneity contribution but small impact on results → Abnormal studies requiring attention
69
70---
71
72## Data Format Requirements
73
74Depending on the data type, the CSV file needs to contain different columns:
75
76### Binary (Binary Outcomes)
77| Column Name | Description |
78|------|------|
79| study | Study name |
80| group1_Events | Number of events in experimental group |
81| group1_sample_size | Total sample size of experimental group |
82| group2_Events | Number of events in control group |
83| group2_sample_size | Total sample size of control group |
84
85### Continuity (Continuous Outcomes)
86| Column Name | Description |
87|------|------|
88| study | Study name |
89| group1_sample_size | Sample size of experimental group |
90| group1_Mean | Mean of experimental group |
91| group1_SD | Standard deviation of experimental group |
92| group2_sample_size | Sample size of control group |
93| group2_Mean | Mean of control group |
94| group2_SD | Standard deviation of control group |
95
96### Survival (Survival Outcomes)
97| Column Name | Description |
98|------|------|
99| study | Study name |
100| group1_HR | Hazard ratio |
101| group1_95%Lower CI | Lower bound of 95% confidence interval |
102| group1_95%Upper CI | Upper bound of 95% confidence interval |
103
104---
105
106## Workflow
107
108### Step 1: Validate Input Data
109
1101. Read the CSV file provided by the user
1112. Check necessary columns based on data type
1123. Validate data integrity (at least 3 studies are required for valid heterogeneity analysis)
113
114### Step 2: Execute R Script
115
116Call command:
117```bash
118Rscript scripts/baujat_plot.R "<csv_path>" "<type>" "<outcome_name>" "<output_dir>"
119```
120
121Parameter descriptions:
122- `csv_path`: Absolute path of input CSV file
123- `type`: Data type (Binary / Continuity / Survival)
124- `outcome_name`: Name of outcome indicator (optional)
125- `output_dir`: Output directory (optional)
126
127### Step 3: Output Results
128
129**Upon success**:
130
131```
132═══════════════════════════════════════════
133Baujat Plot Generation Complete
134═══════════════════════════════════════════
135
136【Outcome Indicator】{outcome_name}
137【Data Type】{type}
138【Included Studies】{n} studies
139
140【Heterogeneity Statistics】
141• I² = {I2}%
142• Tau² = {tau2}
143• Q = {Q}, df = {df}, P = {pval_Q}
144
145【Output Files】
146• Baujat plot: {output_dir}/{type}_baujat_{outcome}.png
147• Contribution data: {output_dir}/{type}_baujat_{outcome}.csv
148
149【Heterogeneity Contribution Ranking】(sorted by Q contribution in descending order)
150Rank Study Result Contribution Q Contribution Judgment
151─────────────────────────────────────────────────────
1521 Smith 2020 0.85 3.42 ⚠️ Outlier
1532 Jones 2021 0.32 1.15 Normal
154...
155
156【Recommendations】
157{Recommendations based on analysis results}
158
159═══════════════════════════════════════════
160```
161
162---
163
164## R Script Dependencies
165
166The following R packages need to be installed:
167- meta
168- metafor
169- ggplot2
170- ggrepel (for label overlap avoidance)
171
172If the user environment lacks these packages, suggest running:
173```r
174install.packages(c("meta", "metafor", "ggplot2", "ggrepel"))
175```
176
177## When Not to Use
178
179- Do not use this skill when the required source data, identifiers, files, or credentials are missing.
180- Do not use this skill when the user asks for fabricated results, unsupported claims, or out-of-scope conclusions.
181- Do not use this skill when a simpler direct answer is more appropriate than the documented workflow.
182
183## Required Inputs
184
185- A clearly specified task goal aligned with the documented scope.
186- All required files, identifiers, parameters, or environment variables before execution.
187- Any domain constraints, formatting requirements, and expected output destination if applicable.
188
189## Output Contract
190
191- Return a structured deliverable that is directly usable without reformatting.
192- If a file is produced, prefer a deterministic output name such as `meta_baujat_plot_result.md` unless the skill documentation defines a better convention.
193- Include a short validation summary describing what was checked, what assumptions were made, and any remaining limitations.
194
195## Validation and Safety Rules
196
197- Validate required inputs before execution and stop early when mandatory fields or files are missing.
198- Do not fabricate measurements, references, findings, or conclusions that are not supported by the provided source material.
199- Emit a clear warning when credentials, privacy constraints, safety boundaries, or unsupported requests affect the result.
200- Keep the output safe, reproducible, and within the documented scope at all times.
201
202## Failure Handling
203
204- If validation fails, explain the exact missing field, file, or parameter and show the minimum fix required.
205- If an external dependency or script fails, surface the command path, likely cause, and the next recovery step.
206- If partial output is returned, label it clearly and identify which checks could not be completed.
207
208## Quick Validation
209
210Run this minimal verification path before full execution when possible:
211
212```bash
213python scripts/baujat_plot_fallback.py --help
214```
215
216Expected output format:
217
218```text
219Result file: meta_baujat_plot_result.md
220Validation summary: PASS/FAIL with brief notes
221Assumptions: explicit list if any
222```