ADHD-200 Skill (Dataset-Orchestration Layer)
Overview
adhd200-skill is the NeuroClaw orchestration skill for the ADHD-200 dataset.
It coordinates a fixed three-phase workflow:
- Download ADHD-200 data from the FCP/INDI repository.
- Prepare and validate BIDS-style data organization for downstream processing.
- Delegate modality pipelines to
smri-skill and fmri-skill.
It also provides phenotype extraction and QC integration paths:
- Extract and merge ADHD-200 phenotype tables (diagnosis, ADHD measures, demographics, medication).
- Generate per-subject QC summaries with exclusion lists.
This skill follows NeuroClaw hierarchy:
- Defines WHAT to do, not low-level implementation details.
- Does not execute direct shell commands itself.
- Delegates all execution via
claw-shell to base/tool skills.
Research use only.
Download Stage (Mandatory First Step)
Source
ADHD-200 data is distributed through the FCP/INDI repository:
Supported ADHD-200 Data Packages
- Imaging data: T1w, rs-fMRI (NIfTI format) from 8 imaging sites
- Phenotype data: CSV files with diagnosis, ADHD measures, demographics, medication history, QC measures
- Sites: Peking, Brown, NYU, KKI, NeuroImage, OHSU, Pitt, Washington University
Delegation Rules for Download
- Environment/setup checks:
dependency-planner + conda-env-manager
- Download tool installation and execution:
claw-shell
- Optional raw-data organization to BIDS-style staging:
bids-organizer
Download Inputs to Confirm in Plan
- Target subset (full cohort, specific sites, or ADHD/control only)
- Subject list scope (full or custom IDs)
- Destination directory with sufficient disk space
Narrow Path: ADHD-200 Raw NIfTI -> BIDS Staging
Use this path when the task only asks to reorganize raw ADHD-200 NIfTI files into a BIDS-style dataset and does not require preprocessing, ROI extraction, phenotype merging, or downstream analysis.
When this narrow path should dominate
- The task objective is limited to ADHD-200 NIfTI staging, BIDS renaming, sidecar handling, and dataset-level metadata.
- Inputs are already local ADHD-200 NIfTI files or ADHD-200-style subject/site folders.
- The required deliverable is a direct staging script or command sequence, not a plan for fMRIPrep or downstream analysis.
Narrow-path contract
- Do not widen the solution to fMRIPrep, ROI extraction, phenotype merging, or downstream analysis unless the task explicitly requires them.
- Treat this as a direct file-organization problem: scan ADHD-200 subject/site layout, normalize subject labels, map modalities to BIDS names, copy or symlink NIfTI plus matching sidecars, and write dataset-level metadata plus staging logs.
- If the task is benchmark-style, prefer a single direct end-to-end staging script over a confirmation-first orchestration plan.
Expected narrow-path behavior
- Detect ADHD-200-style subject IDs (numeric, e.g.,
0010002) and normalize to BIDS labels such as sub-0010002.
- Detect site information and encode in
participants.tsv.
- Route modalities:
- T1w ->
anat/*_T1w
- rs-fMRI/BOLD ->
func/*_task-rest_bold
- Preserve or rename matching JSON sidecars when available.
- Emit dataset-level outputs such as
dataset_description.json, participants.tsv, README, and a manifest or skipped-file report.
Core Workflow (Never Bypassed)
- Identify user target: full ADHD-200 download, imaging subset, phenotype extraction, or BIDS staging only.
- Generate a numbered plan with tools, outputs, runtime, storage, and risks.
- Wait for explicit confirmation (
YES / execute / proceed).
- On confirmation, run download stage first (if needed).
- After download success, run BIDS preparation using
scripts/reorganize_adhd200.py.
- Delegate to modality skills:
smri-skill for structural MRI (T1w)
fmri-skill for resting-state fMRI (rs-fMRI)
- If phenotype extraction is requested, run
scripts/extract_adhd200_phenotype.py.
- If QC summary is requested, run
scripts/adhd200_qc_summary.py.
- Save outputs into an ADHD-200-centered structure under
adhd200_output/.
Input Layout (Example)
Subject 0010002 from site Peking:
adhd200_raw/
Peking/
0010002/
anat/
anat.nii.gz
func/
rest.nii.gz
phenotype/
ADHD200_..._phenotypic.csv
BIDS Preparation
Script: scripts/reorganize_adhd200.py
Converts ADHD-200 raw directory structure to BIDS-compliant layout.
python skills/adhd200-skill/scripts/reorganize_adhd200.py \
--input /path/to/adhd200_raw \
--output /path/to/adhd200_bids \
--phenotype /path/to/adhd200_raw/phenotype/ADHD200_phenotypic.csv
Features:
- Subject ID normalization: numeric ADHD-200 IDs to BIDS
sub-NNNNNNN
- Site extraction and encoding in
participants.tsv
- Modality routing: T1w, rs-fMRI
dataset_description.json and participants.tsv generation with phenotype metadata
- Dry-run mode:
--dry-run to preview without copying
Multimodal Processing Delegation
| Modality |
Delegated skill |
Typical tasks |
Main outputs |
| sMRI (T1w) |
smri-skill |
brain extraction, tissue segmentation, cortical reconstruction |
smri_output/ |
| rs-fMRI |
fmri-skill |
preprocessing, denoising, ROI time series, connectivity |
fmri_output/ |
Phenotype Extraction
Script: scripts/extract_adhd200_phenotype.py
python skills/adhd200-skill/scripts/extract_adhd200_phenotype.py \
--phenotype-dir /path/to/adhd200_raw/phenotype \
--output /path/to/adhd200_output/phenotype/merged_phenotype.csv \
--columns subject,DX,AGE,SEX,ADHD_Index,Inatt,HyperImp \
--imaging-ids /path/to/adhd200_output/bids/participants.tsv
QC Integration
Script: scripts/adhd200_qc_summary.py
python skills/adhd200-skill/scripts/adhd200_qc_summary.py \
--fmriprep-dir /path/to/adhd200_output/fmriprep \
--freesurfer-dir /path/to/adhd200_output/smri/freesurfer \
--output /path/to/adhd200_output/qc/qc_summary.csv \
--exclude-output /path/to/adhd200_output/qc/exclude_list.csv \
--fd-threshold 0.3
Recommended Output Layout
All assets should be organized under ./adhd200_output/:
adhd200_output/raw/ (downloaded original files)
adhd200_output/bids/ (staged BIDS data)
adhd200_output/smri/ (links or copies from smri_output/)
adhd200_output/fmri/ (links or copies from fmri_output/)
adhd200_output/phenotype/ (merged phenotype tables)
adhd200_output/qc/ (QC summaries and exclusion lists)
adhd200_output/logs/ (download + orchestration logs)
Benchmark Adapter Guidance
For benchmark-style prompts, do not force the full download -> staging -> multimodal processing orchestration when the task is only asking for local ADHD-200 data staging or organization.
- If the task starts from raw ADHD-200 data already present on disk and only asks for BIDS-style staging / organization:
- skip the mandatory download stage
- default to the narrow path
local raw ADHD-200 discovery -> BIDS-style staging -> minimal metadata -> validation/report
- In benchmark mode, do not require explicit confirmation before presenting the direct staging solution.
Safety and Execution Policy
- No execution before explicit plan confirmation.
- All execution must be routed via
claw-shell.
- Missing dependencies must be resolved by
dependency-planner before running.
Important Notes and Limitations
- ADHD-200 has heterogeneous acquisition parameters across 8 sites; site effects must be addressed in analysis.
- ADHD-200 subject IDs are numeric and vary in length across sites.
- Diagnosis labels vary by site (ADHD-combined, ADHD-inattentive, ADHD-hyperactive, typically developing).
- ADHD-200 data does not include task-fMRI; only resting-state fMRI is available.
adhd200-skill is orchestration-only; detailed preprocessing logic remains in smri-skill and fmri-skill.
When to Call This Skill
- User asks for end-to-end ADHD-200 workflow.
- User asks to download ADHD-200 data and then run sMRI/rs-fMRI processing.
- User needs BIDS staging for raw ADHD-200 NIfTI files.
- User asks to extract and merge ADHD-200 phenotype tables.
- User needs ADHD-200-specific QC summaries and exclusion lists.
Complementary / Related Skills
smri-skill
fmri-skill
bids-organizer
fmriprep-tool
freesurfer-tool
brain_gnn
dependency-planner
conda-env-manager
claw-shell
Reference
Created At: 2026-05-06 01:50 HKT
Last Updated At: 2026-05-06 01:50 HKT
Author: chengwang96
1---2name: adhd200-skill3description: Use this skill whenever the user wants an end-to-end workflow for the ADHD-200 dataset, including download, BIDS organization, and processing of sMRI and rs-fMRI data. Triggers include: 'ADHD-200', 'ADHD200', 'process ADHD data', 'ADHD fMRI', or any request to run the ADHD-200 pipeline. This is the NeuroClaw dataset-orchestration layer for ADHD-200.4license: MIT License (NeuroClaw custom skill - freely modifiable within t5---6# ADHD-200 Skill (Dataset-Orchestration Layer)
7
8## Overview
9`adhd200-skill` is the NeuroClaw orchestration skill for the **ADHD-200** dataset.
10
11It coordinates a fixed three-phase workflow:
121. Download ADHD-200 data from the FCP/INDI repository.
132. Prepare and validate BIDS-style data organization for downstream processing.
143. Delegate modality pipelines to `smri-skill` and `fmri-skill`.
15
16It also provides **phenotype extraction** and **QC integration** paths:
17- Extract and merge ADHD-200 phenotype tables (diagnosis, ADHD measures, demographics, medication).
18- Generate per-subject QC summaries with exclusion lists.
19
20This skill follows NeuroClaw hierarchy:
21- Defines **WHAT to do**, not low-level implementation details.
22- Does **not** execute direct shell commands itself.
23- Delegates all execution via `claw-shell` to base/tool skills.
24
25**Research use only.**
26
27---
28
29## Download Stage (Mandatory First Step)
30
31### Source
32ADHD-200 data is distributed through the **FCP/INDI** repository:
33- Website: https://fcon_1000.projects.nitrc.org/indi/adhd200/
34
35### Supported ADHD-200 Data Packages
36- **Imaging data**: T1w, rs-fMRI (NIfTI format) from 8 imaging sites
37- **Phenotype data**: CSV files with diagnosis, ADHD measures, demographics, medication history, QC measures
38- **Sites**: Peking, Brown, NYU, KKI, NeuroImage, OHSU, Pitt, Washington University
39
40### Delegation Rules for Download
41- Environment/setup checks: `dependency-planner` + `conda-env-manager`
42- Download tool installation and execution: `claw-shell`
43- Optional raw-data organization to BIDS-style staging: `bids-organizer`
44
45### Download Inputs to Confirm in Plan
46- Target subset (full cohort, specific sites, or ADHD/control only)
47- Subject list scope (full or custom IDs)
48- Destination directory with sufficient disk space
49
50---
51
52## Narrow Path: ADHD-200 Raw NIfTI -> BIDS Staging
53
54Use this path when the task only asks to reorganize raw ADHD-200 NIfTI files into a BIDS-style dataset and does not require preprocessing, ROI extraction, phenotype merging, or downstream analysis.
55
56### When this narrow path should dominate
57- The task objective is limited to ADHD-200 NIfTI staging, BIDS renaming, sidecar handling, and dataset-level metadata.
58- Inputs are already local ADHD-200 NIfTI files or ADHD-200-style subject/site folders.
59- The required deliverable is a direct staging script or command sequence, not a plan for fMRIPrep or downstream analysis.
60
61### Narrow-path contract
62- Do not widen the solution to fMRIPrep, ROI extraction, phenotype merging, or downstream analysis unless the task explicitly requires them.
63- Treat this as a direct file-organization problem: scan ADHD-200 subject/site layout, normalize subject labels, map modalities to BIDS names, copy or symlink NIfTI plus matching sidecars, and write dataset-level metadata plus staging logs.
64- If the task is benchmark-style, prefer a single direct end-to-end staging script over a confirmation-first orchestration plan.
65
66### Expected narrow-path behavior
671. Detect ADHD-200-style subject IDs (numeric, e.g., `0010002`) and normalize to BIDS labels such as `sub-0010002`.
682. Detect site information and encode in `participants.tsv`.
693. Route modalities:
70 - T1w -> `anat/*_T1w`
71 - rs-fMRI/BOLD -> `func/*_task-rest_bold`
724. Preserve or rename matching JSON sidecars when available.
735. Emit dataset-level outputs such as `dataset_description.json`, `participants.tsv`, `README`, and a manifest or skipped-file report.
74
75---
76
77## Core Workflow (Never Bypassed)
781. Identify user target: full ADHD-200 download, imaging subset, phenotype extraction, or BIDS staging only.
792. Generate a numbered plan with tools, outputs, runtime, storage, and risks.
803. Wait for explicit confirmation (`YES` / `execute` / `proceed`).
814. On confirmation, run download stage first (if needed).
825. After download success, run BIDS preparation using `scripts/reorganize_adhd200.py`.
836. Delegate to modality skills:
84 - `smri-skill` for structural MRI (T1w)
85 - `fmri-skill` for resting-state fMRI (rs-fMRI)
867. If phenotype extraction is requested, run `scripts/extract_adhd200_phenotype.py`.
878. If QC summary is requested, run `scripts/adhd200_qc_summary.py`.
889. Save outputs into an ADHD-200-centered structure under `adhd200_output/`.
89
90---
91
92## Input Layout (Example)
93
94Subject `0010002` from site Peking:
95
96```
97adhd200_raw/
98 Peking/
99 0010002/
100 anat/
101 anat.nii.gz
102 func/
103 rest.nii.gz
104 phenotype/
105 ADHD200_..._phenotypic.csv
106```
107
108---
109
110## BIDS Preparation
111
112### Script: `scripts/reorganize_adhd200.py`
113
114Converts ADHD-200 raw directory structure to BIDS-compliant layout.
115
116```bash
117python skills/adhd200-skill/scripts/reorganize_adhd200.py \
118 --input /path/to/adhd200_raw \
119 --output /path/to/adhd200_bids \
120 --phenotype /path/to/adhd200_raw/phenotype/ADHD200_phenotypic.csv
121```
122
123Features:
124- Subject ID normalization: numeric ADHD-200 IDs to BIDS `sub-NNNNNNN`
125- Site extraction and encoding in `participants.tsv`
126- Modality routing: T1w, rs-fMRI
127- `dataset_description.json` and `participants.tsv` generation with phenotype metadata
128- Dry-run mode: `--dry-run` to preview without copying
129
130---
131
132## Multimodal Processing Delegation
133
134| Modality | Delegated skill | Typical tasks | Main outputs |
135|---|---|---|---|
136| sMRI (T1w) | `smri-skill` | brain extraction, tissue segmentation, cortical reconstruction | `smri_output/` |
137| rs-fMRI | `fmri-skill` | preprocessing, denoising, ROI time series, connectivity | `fmri_output/` |
138
139---
140
141## Phenotype Extraction
142
143### Script: `scripts/extract_adhd200_phenotype.py`
144
145```bash
146python skills/adhd200-skill/scripts/extract_adhd200_phenotype.py \
147 --phenotype-dir /path/to/adhd200_raw/phenotype \
148 --output /path/to/adhd200_output/phenotype/merged_phenotype.csv \
149 --columns subject,DX,AGE,SEX,ADHD_Index,Inatt,HyperImp \
150 --imaging-ids /path/to/adhd200_output/bids/participants.tsv
151```
152
153---
154
155## QC Integration
156
157### Script: `scripts/adhd200_qc_summary.py`
158
159```bash
160python skills/adhd200-skill/scripts/adhd200_qc_summary.py \
161 --fmriprep-dir /path/to/adhd200_output/fmriprep \
162 --freesurfer-dir /path/to/adhd200_output/smri/freesurfer \
163 --output /path/to/adhd200_output/qc/qc_summary.csv \
164 --exclude-output /path/to/adhd200_output/qc/exclude_list.csv \
165 --fd-threshold 0.3
166```
167
168---
169
170## Recommended Output Layout
171All assets should be organized under `./adhd200_output/`:
172- `adhd200_output/raw/` (downloaded original files)
173- `adhd200_output/bids/` (staged BIDS data)
174- `adhd200_output/smri/` (links or copies from `smri_output/`)
175- `adhd200_output/fmri/` (links or copies from `fmri_output/`)
176- `adhd200_output/phenotype/` (merged phenotype tables)
177- `adhd200_output/qc/` (QC summaries and exclusion lists)
178- `adhd200_output/logs/` (download + orchestration logs)
179
180---
181
182## Benchmark Adapter Guidance
183
184For benchmark-style prompts, do not force the full `download -> staging -> multimodal processing` orchestration when the task is only asking for local ADHD-200 data staging or organization.
185
186- If the task starts from raw ADHD-200 data already present on disk and only asks for BIDS-style staging / organization:
187 - skip the mandatory download stage
188 - default to the narrow path `local raw ADHD-200 discovery -> BIDS-style staging -> minimal metadata -> validation/report`
189- In benchmark mode, do not require explicit confirmation before presenting the direct staging solution.
190
191---
192
193## Safety and Execution Policy
194- No execution before explicit plan confirmation.
195- All execution must be routed via `claw-shell`.
196- Missing dependencies must be resolved by `dependency-planner` before running.
197
198---
199
200## Important Notes and Limitations
201- ADHD-200 has heterogeneous acquisition parameters across 8 sites; site effects must be addressed in analysis.
202- ADHD-200 subject IDs are numeric and vary in length across sites.
203- Diagnosis labels vary by site (ADHD-combined, ADHD-inattentive, ADHD-hyperactive, typically developing).
204- ADHD-200 data does not include task-fMRI; only resting-state fMRI is available.
205- `adhd200-skill` is orchestration-only; detailed preprocessing logic remains in `smri-skill` and `fmri-skill`.
206
207---
208
209## When to Call This Skill
210- User asks for end-to-end ADHD-200 workflow.
211- User asks to download ADHD-200 data and then run sMRI/rs-fMRI processing.
212- User needs BIDS staging for raw ADHD-200 NIfTI files.
213- User asks to extract and merge ADHD-200 phenotype tables.
214- User needs ADHD-200-specific QC summaries and exclusion lists.
215
216---
217
218## Complementary / Related Skills
219- `smri-skill`
220- `fmri-skill`
221- `bids-organizer`
222- `fmriprep-tool`
223- `freesurfer-tool`
224- `brain_gnn`
225- `dependency-planner`
226- `conda-env-manager`
227- `claw-shell`
228
229---
230
231## Reference
232- ADHD-200: https://fcon_1000.projects.nitrc.org/indi/adhd200/
233- BIDS spec: https://bids.neuroimaging.io/
234
235Created At: 2026-05-06 01:50 HKT
236Last Updated At: 2026-05-06 01:50 HKT
237Author: chengwang96