ABIDE Skill (Dataset-Orchestration Layer)
Overview
abide-skill is the NeuroClaw orchestration skill for the ABIDE (Autism Brain Imaging Data Exchange) dataset.
It coordinates a fixed three-phase workflow:
- Download ABIDE data from the FCP/INDI repository or NITRC.
- Prepare and validate BIDS-style data organization for downstream processing.
- Delegate modality pipelines to
smri-skill and fmri-skill.
It also provides phenotype extraction and QC integration paths:
- Extract and merge ABIDE phenotype tables (diagnosis, age, sex, site, FIQ, ADOS, etc.).
- Generate per-subject QC summaries with exclusion lists.
This skill follows NeuroClaw hierarchy:
- Defines WHAT to do, not low-level implementation details.
- Does not execute direct shell commands itself.
- Delegates all execution via
claw-shell to base/tool skills.
Research use only.
Download Stage (Mandatory First Step)
Source
ABIDE data is distributed through the FCP/INDI repository:
Supported ABIDE Data Packages
- ABIDE I: 1,112 subjects from 17 international sites (539 ASD, 573 controls)
- ABIDE II: 1,044 subjects from 19 sites
- Phenotype data: CSV files with demographics, diagnosis, cognitive scores
- Preprocessed derivatives (optional): CPAC, DPARSF, CCS, NeuroMark pipelines
Delegation Rules for Download
- Environment/setup checks:
dependency-planner + conda-env-manager
- Download tool installation and execution:
claw-shell
- Optional raw-data organization to BIDS-style staging:
bids-organizer
Download Inputs to Confirm in Plan
- Target ABIDE version (I, II, or both)
- Target subset (full cohort, specific sites, or ASD/control only)
- Subject list scope (full or custom IDs)
- Whether to download raw data or preprocessed derivatives
- Destination directory with sufficient disk space
Narrow Path: ABIDE Raw NIfTI -> BIDS Staging
Use this path when the task only asks to reorganize raw ABIDE NIfTI files into a BIDS-style dataset and does not require preprocessing, ROI extraction, phenotype merging, or downstream analysis.
When this narrow path should dominate
- The task objective is limited to ABIDE NIfTI staging, BIDS renaming, sidecar handling, and dataset-level metadata.
- Inputs are already local ABIDE NIfTI files or ABIDE-style subject/site folders.
- The required deliverable is a direct staging script or command sequence, not a plan for fMRIPrep or downstream analysis.
Narrow-path contract
- Do not widen the solution to fMRIPrep, ROI extraction, phenotype merging, or downstream analysis unless the task explicitly requires them.
- Treat this as a direct file-organization problem: scan ABIDE subject/site layout, normalize subject labels, map modalities to BIDS names, copy or symlink NIfTI plus matching sidecars, and write dataset-level metadata plus staging logs.
- If the task is benchmark-style, prefer a single direct end-to-end staging script over a confirmation-first orchestration plan.
Expected narrow-path behavior
- Detect ABIDE-style subject IDs (numeric, e.g.,
0050642) and normalize to BIDS labels such as sub-0050642.
- Detect site information and encode as BIDS session or metadata (e.g.,
ses-NYU, or site column in participants.tsv).
- Route modalities:
- T1w ->
anat/*_T1w
- rs-fMRI/BOLD ->
func/*_task-rest_bold
- Preserve or rename matching JSON sidecars when available; if metadata is absent, create only the minimal dataset files required by the task and log the limitation.
- Emit dataset-level outputs such as
dataset_description.json, participants.tsv, README, and a manifest or skipped-file report.
Core Workflow (Never Bypassed)
- Identify user target: full ABIDE download, imaging subset, phenotype extraction, or BIDS staging only.
- Generate a numbered plan with tools, outputs, runtime, storage, and risks.
- Wait for explicit confirmation (
YES / execute / proceed).
- On confirmation, run download stage first (if needed).
- After download success, run BIDS preparation using
scripts/reorganize_abide.py.
- Delegate to modality skills:
smri-skill for structural MRI (T1w)
fmri-skill for resting-state fMRI (rs-fMRI)
- If phenotype extraction is requested, run
scripts/extract_abide_phenotype.py.
- If QC summary is requested, run
scripts/abide_qc_summary.py.
- Save outputs into an ABIDE-centered structure under
abide_output/.
Input Layout (Example)
Subject 0050642 from site NYU:
abide_raw/
NYU/
0050642/
session_1/
anat_1/
anat.nii.gz
func_1/
func.nii.gz
phenotype/
ABIDE_phenotypic.csv
Or flat layout:
abide_raw/
0050642/
anat/
T1w.nii.gz
func/
rest_bold.nii.gz
BIDS Preparation
Script: scripts/reorganize_abide.py
Converts ABIDE raw directory structure to BIDS-compliant layout.
python skills/abide-skill/scripts/reorganize_abide.py \
--input /path/to/abide_raw \
--output /path/to/abide_bids \
--phenotype /path/to/abide_raw/phenotype/ABIDE_phenotypic.csv
Features:
- Subject ID normalization: numeric ABIDE IDs to BIDS
sub-NNNNNNN
- Site extraction and encoding in
participants.tsv
- Modality routing: T1w, rs-fMRI
- Sidecar JSON preservation and validation
dataset_description.json and participants.tsv generation with phenotype metadata
- Dry-run mode:
--dry-run to preview without copying
Multimodal Processing Delegation
After BIDS staging completes, abide-skill delegates by modality:
| Modality |
Delegated skill |
Typical tasks |
Main outputs |
| sMRI (T1w) |
smri-skill |
brain extraction, tissue segmentation, cortical reconstruction, ROI morphometry |
smri_output/ derivatives and stats |
| rs-fMRI |
fmri-skill |
preprocessing, denoising, ROI time series, connectivity |
fmri_output/ derivatives, timeseries, connectivity |
Delegation Strategy
- If user asks for full ABIDE analysis: run sMRI -> fMRI in ordered phases.
- If user asks for one modality only: call only the corresponding modality skill.
- If compute resources are adequate and the user approves parallel runs: run modality pipelines in parallel.
Phenotype Extraction
Script: scripts/extract_abide_phenotype.py
Extracts and merges ABIDE phenotype tables for downstream analysis.
python skills/abide-skill/scripts/extract_abide_phenotype.py \
--phenotype-dir /path/to/abide_raw/phenotype \
--output /path/to/abide_output/phenotype/merged_phenotype.csv \
--columns subject,DX_GROUP,AGE_AT_SCAN,SEX,FIQ,VIQ,PIQ,site \
--imaging-ids /path/to/abide_output/bids/participants.tsv
Features:
- Reads ABIDE phenotype CSV files (ABIDE I and II compatible)
- Standardizes column names (DX_GROUP: 1=ASD, 2=control)
- Column selection and renaming
- Site encoding and grouping
- Missing value handling
- Cross-reference with imaging subject list
- Outputs merged CSV ready for statistical analysis or model training
QC Integration
Script: scripts/abide_qc_summary.py
Generates per-subject QC summaries and exclusion lists.
python skills/abide-skill/scripts/abide_qc_summary.py \
--fmriprep-dir /path/to/abide_output/fmriprep \
--freesurfer-dir /path/to/abide_output/smri/freesurfer \
--raw-qc /path/to/abide_raw/phenotype/ABIDE_phenotypic.csv \
--output /path/to/abide_output/qc/qc_summary.csv \
--exclude-output /path/to/abide_output/qc/exclude_list.csv \
--fd-threshold 0.3 \
--coverage-threshold 0.8
Features:
- Reads fMRIPrep confounds (framewise displacement, DVARS)
- Reads FreeSurfer recon-all QC metrics
- Incorporates ABIDE QC flags (QC_RATER_1, func_perc_fd, anat_rater_1, etc.)
- Applies exclusion criteria: motion threshold (FD), coverage threshold, structural quality
- Per-site QC summary for site-effect assessment
- Outputs per-subject QC summary CSV and exclusion list CSV
Recommended Output Layout
All assets should be organized under ./abide_output/:
abide_output/raw/ (downloaded original ABIDE files)
abide_output/bids/ (staged BIDS data)
abide_output/staging/ (optional normalized staging intermediate)
abide_output/smri/ (links or copies from smri_output/)
abide_output/fmri/ (links or copies from fmri_output/)
abide_output/phenotype/ (merged phenotype tables)
abide_output/qc/ (QC summaries and exclusion lists)
abide_output/logs/ (download + orchestration logs)
Benchmark Adapter Guidance
For benchmark-style prompts, do not force the full download -> staging -> multimodal processing orchestration when the task is only asking for local ABIDE data staging or organization.
- If the task starts from raw ABIDE data already present on disk and only asks for BIDS-style staging / organization:
- skip the mandatory download stage
- do not automatically delegate to
smri-skill or fmri-skill
- default to the narrow path
local raw ABIDE discovery -> BIDS-style staging -> minimal metadata -> validation/report
- In benchmark mode, do not require explicit confirmation before presenting the direct staging solution.
- Preserve the ABIDE-centered output contract under
abide_output/bids/ when the task is specifically a staging benchmark.
- Only use the full multimodal orchestration and confirmation-heavy workflow when the prompt explicitly asks for download, end-to-end ABIDE processing, or post-staging structural / functional analysis.
Safety and Execution Policy
- No execution before explicit plan confirmation.
- All execution must be routed via
claw-shell.
- Missing dependencies must be resolved by
dependency-planner before running.
- If download fails for partial subjects, continue batch with clear failure report and retry list.
Important Notes and Limitations
- ABIDE data from different sites may have varying acquisition parameters; site effects should be accounted for in analysis.
- ABIDE subject IDs are numeric and vary in length across sites.
- ABIDE I and II have different phenotype table formats; the extraction script handles both.
- ABIDE provides preprocessed derivatives from multiple pipelines (CPAC, DPARSF, CCS); raw data processing via fMRIPrep is recommended for reproducibility.
- ABIDE data does not include task-fMRI; only resting-state fMRI is available.
abide-skill is orchestration-only; detailed preprocessing logic remains in smri-skill and fmri-skill.
When to Call This Skill
- User asks for end-to-end ABIDE workflow.
- User asks to download ABIDE data and then run sMRI/rs-fMRI processing.
- User needs BIDS staging for raw ABIDE NIfTI files.
- User asks to extract and merge ABIDE phenotype tables.
- User asks for ABIDE-specific QC summaries and exclusion lists.
- User needs a single entry point for ABIDE multimodal orchestration.
Complementary / Related Skills
smri-skill
fmri-skill
bids-organizer
fmriprep-tool
freesurfer-tool
neurostorm
brain_gnn
dependency-planner
conda-env-manager
claw-shell
Reference
Created At: 2026-05-06 01:45 HKT
Last Updated At: 2026-05-06 01:45 HKT
Author: chengwang96
1---2name: abide-skill3description: Use this skill whenever the user wants an end-to-end workflow for the ABIDE (Autism Brain Imaging Data Exchange) dataset, including download, BIDS organization, and processing of sMRI and rs-fMRI data. Triggers include: 'ABIDE', 'ABIDE data', 'process ABIDE', 'ABIDE fMRI', 'ABIDE sMRI', 'autism imaging', or any request to run the ABIDE pipeline. This is the NeuroClaw dataset-orchestration layer for ABIDE.4license: MIT License (NeuroClaw custom skill - freely modifiable within t5---6# ABIDE Skill (Dataset-Orchestration Layer)
7
8## Overview
9`abide-skill` is the NeuroClaw orchestration skill for the **ABIDE (Autism Brain Imaging Data Exchange)** dataset.
10
11It coordinates a fixed three-phase workflow:
121. Download ABIDE data from the FCP/INDI repository or NITRC.
132. Prepare and validate BIDS-style data organization for downstream processing.
143. Delegate modality pipelines to `smri-skill` and `fmri-skill`.
15
16It also provides **phenotype extraction** and **QC integration** paths:
17- Extract and merge ABIDE phenotype tables (diagnosis, age, sex, site, FIQ, ADOS, etc.).
18- Generate per-subject QC summaries with exclusion lists.
19
20This skill follows NeuroClaw hierarchy:
21- Defines **WHAT to do**, not low-level implementation details.
22- Does **not** execute direct shell commands itself.
23- Delegates all execution via `claw-shell` to base/tool skills.
24
25**Research use only.**
26
27---
28
29## Download Stage (Mandatory First Step)
30
31### Source
32ABIDE data is distributed through the **FCP/INDI** repository:
33- ABIDE I: https://fcon_1000.projects.nitrc.org/indi/abide/
34- ABIDE II: https://fcon_1000.projects.nitrc.org/indi/abide_II.html
35- NITRC mirror: https://www.nitrc.org/projects/fcp_indi/
36
37### Supported ABIDE Data Packages
38- **ABIDE I**: 1,112 subjects from 17 international sites (539 ASD, 573 controls)
39- **ABIDE II**: 1,044 subjects from 19 sites
40- **Phenotype data**: CSV files with demographics, diagnosis, cognitive scores
41- **Preprocessed derivatives** (optional): CPAC, DPARSF, CCS, NeuroMark pipelines
42
43### Delegation Rules for Download
44- Environment/setup checks: `dependency-planner` + `conda-env-manager`
45- Download tool installation and execution: `claw-shell`
46- Optional raw-data organization to BIDS-style staging: `bids-organizer`
47
48### Download Inputs to Confirm in Plan
49- Target ABIDE version (I, II, or both)
50- Target subset (full cohort, specific sites, or ASD/control only)
51- Subject list scope (full or custom IDs)
52- Whether to download raw data or preprocessed derivatives
53- Destination directory with sufficient disk space
54
55---
56
57## Narrow Path: ABIDE Raw NIfTI -> BIDS Staging
58
59Use this path when the task only asks to reorganize raw ABIDE NIfTI files into a BIDS-style dataset and does not require preprocessing, ROI extraction, phenotype merging, or downstream analysis.
60
61### When this narrow path should dominate
62- The task objective is limited to ABIDE NIfTI staging, BIDS renaming, sidecar handling, and dataset-level metadata.
63- Inputs are already local ABIDE NIfTI files or ABIDE-style subject/site folders.
64- The required deliverable is a direct staging script or command sequence, not a plan for fMRIPrep or downstream analysis.
65
66### Narrow-path contract
67- Do not widen the solution to fMRIPrep, ROI extraction, phenotype merging, or downstream analysis unless the task explicitly requires them.
68- Treat this as a direct file-organization problem: scan ABIDE subject/site layout, normalize subject labels, map modalities to BIDS names, copy or symlink NIfTI plus matching sidecars, and write dataset-level metadata plus staging logs.
69- If the task is benchmark-style, prefer a single direct end-to-end staging script over a confirmation-first orchestration plan.
70
71### Expected narrow-path behavior
721. Detect ABIDE-style subject IDs (numeric, e.g., `0050642`) and normalize to BIDS labels such as `sub-0050642`.
732. Detect site information and encode as BIDS session or metadata (e.g., `ses-NYU`, or site column in `participants.tsv`).
743. Route modalities:
75 - T1w -> `anat/*_T1w`
76 - rs-fMRI/BOLD -> `func/*_task-rest_bold`
774. Preserve or rename matching JSON sidecars when available; if metadata is absent, create only the minimal dataset files required by the task and log the limitation.
785. Emit dataset-level outputs such as `dataset_description.json`, `participants.tsv`, `README`, and a manifest or skipped-file report.
79
80---
81
82## Core Workflow (Never Bypassed)
831. Identify user target: full ABIDE download, imaging subset, phenotype extraction, or BIDS staging only.
842. Generate a numbered plan with tools, outputs, runtime, storage, and risks.
853. Wait for explicit confirmation (`YES` / `execute` / `proceed`).
864. On confirmation, run download stage first (if needed).
875. After download success, run BIDS preparation using `scripts/reorganize_abide.py`.
886. Delegate to modality skills:
89 - `smri-skill` for structural MRI (T1w)
90 - `fmri-skill` for resting-state fMRI (rs-fMRI)
917. If phenotype extraction is requested, run `scripts/extract_abide_phenotype.py`.
928. If QC summary is requested, run `scripts/abide_qc_summary.py`.
939. Save outputs into an ABIDE-centered structure under `abide_output/`.
94
95---
96
97## Input Layout (Example)
98
99Subject `0050642` from site NYU:
100
101```
102abide_raw/
103 NYU/
104 0050642/
105 session_1/
106 anat_1/
107 anat.nii.gz
108 func_1/
109 func.nii.gz
110 phenotype/
111 ABIDE_phenotypic.csv
112```
113
114Or flat layout:
115
116```
117abide_raw/
118 0050642/
119 anat/
120 T1w.nii.gz
121 func/
122 rest_bold.nii.gz
123```
124
125---
126
127## BIDS Preparation
128
129### Script: `scripts/reorganize_abide.py`
130
131Converts ABIDE raw directory structure to BIDS-compliant layout.
132
133```bash
134python skills/abide-skill/scripts/reorganize_abide.py \
135 --input /path/to/abide_raw \
136 --output /path/to/abide_bids \
137 --phenotype /path/to/abide_raw/phenotype/ABIDE_phenotypic.csv
138```
139
140Features:
141- Subject ID normalization: numeric ABIDE IDs to BIDS `sub-NNNNNNN`
142- Site extraction and encoding in `participants.tsv`
143- Modality routing: T1w, rs-fMRI
144- Sidecar JSON preservation and validation
145- `dataset_description.json` and `participants.tsv` generation with phenotype metadata
146- Dry-run mode: `--dry-run` to preview without copying
147
148---
149
150## Multimodal Processing Delegation
151
152After BIDS staging completes, `abide-skill` delegates by modality:
153
154| Modality | Delegated skill | Typical tasks | Main outputs |
155|---|---|---|---|
156| sMRI (T1w) | `smri-skill` | brain extraction, tissue segmentation, cortical reconstruction, ROI morphometry | `smri_output/` derivatives and stats |
157| rs-fMRI | `fmri-skill` | preprocessing, denoising, ROI time series, connectivity | `fmri_output/` derivatives, timeseries, connectivity |
158
159### Delegation Strategy
160- If user asks for full ABIDE analysis: run sMRI -> fMRI in ordered phases.
161- If user asks for one modality only: call only the corresponding modality skill.
162- If compute resources are adequate and the user approves parallel runs: run modality pipelines in parallel.
163
164---
165
166## Phenotype Extraction
167
168### Script: `scripts/extract_abide_phenotype.py`
169
170Extracts and merges ABIDE phenotype tables for downstream analysis.
171
172```bash
173python skills/abide-skill/scripts/extract_abide_phenotype.py \
174 --phenotype-dir /path/to/abide_raw/phenotype \
175 --output /path/to/abide_output/phenotype/merged_phenotype.csv \
176 --columns subject,DX_GROUP,AGE_AT_SCAN,SEX,FIQ,VIQ,PIQ,site \
177 --imaging-ids /path/to/abide_output/bids/participants.tsv
178```
179
180Features:
181- Reads ABIDE phenotype CSV files (ABIDE I and II compatible)
182- Standardizes column names (DX_GROUP: 1=ASD, 2=control)
183- Column selection and renaming
184- Site encoding and grouping
185- Missing value handling
186- Cross-reference with imaging subject list
187- Outputs merged CSV ready for statistical analysis or model training
188
189---
190
191## QC Integration
192
193### Script: `scripts/abide_qc_summary.py`
194
195Generates per-subject QC summaries and exclusion lists.
196
197```bash
198python skills/abide-skill/scripts/abide_qc_summary.py \
199 --fmriprep-dir /path/to/abide_output/fmriprep \
200 --freesurfer-dir /path/to/abide_output/smri/freesurfer \
201 --raw-qc /path/to/abide_raw/phenotype/ABIDE_phenotypic.csv \
202 --output /path/to/abide_output/qc/qc_summary.csv \
203 --exclude-output /path/to/abide_output/qc/exclude_list.csv \
204 --fd-threshold 0.3 \
205 --coverage-threshold 0.8
206```
207
208Features:
209- Reads fMRIPrep confounds (framewise displacement, DVARS)
210- Reads FreeSurfer recon-all QC metrics
211- Incorporates ABIDE QC flags (QC_RATER_1, func_perc_fd, anat_rater_1, etc.)
212- Applies exclusion criteria: motion threshold (FD), coverage threshold, structural quality
213- Per-site QC summary for site-effect assessment
214- Outputs per-subject QC summary CSV and exclusion list CSV
215
216---
217
218## Recommended Output Layout
219All assets should be organized under `./abide_output/`:
220- `abide_output/raw/` (downloaded original ABIDE files)
221- `abide_output/bids/` (staged BIDS data)
222- `abide_output/staging/` (optional normalized staging intermediate)
223- `abide_output/smri/` (links or copies from `smri_output/`)
224- `abide_output/fmri/` (links or copies from `fmri_output/`)
225- `abide_output/phenotype/` (merged phenotype tables)
226- `abide_output/qc/` (QC summaries and exclusion lists)
227- `abide_output/logs/` (download + orchestration logs)
228
229---
230
231## Benchmark Adapter Guidance
232
233For benchmark-style prompts, do not force the full `download -> staging -> multimodal processing` orchestration when the task is only asking for local ABIDE data staging or organization.
234
235- If the task starts from raw ABIDE data already present on disk and only asks for BIDS-style staging / organization:
236 - skip the mandatory download stage
237 - do not automatically delegate to `smri-skill` or `fmri-skill`
238 - default to the narrow path `local raw ABIDE discovery -> BIDS-style staging -> minimal metadata -> validation/report`
239- In benchmark mode, do not require explicit confirmation before presenting the direct staging solution.
240- Preserve the ABIDE-centered output contract under `abide_output/bids/` when the task is specifically a staging benchmark.
241- Only use the full multimodal orchestration and confirmation-heavy workflow when the prompt explicitly asks for download, end-to-end ABIDE processing, or post-staging structural / functional analysis.
242
243---
244
245## Safety and Execution Policy
246- No execution before explicit plan confirmation.
247- All execution must be routed via `claw-shell`.
248- Missing dependencies must be resolved by `dependency-planner` before running.
249- If download fails for partial subjects, continue batch with clear failure report and retry list.
250
251---
252
253## Important Notes and Limitations
254- ABIDE data from different sites may have varying acquisition parameters; site effects should be accounted for in analysis.
255- ABIDE subject IDs are numeric and vary in length across sites.
256- ABIDE I and II have different phenotype table formats; the extraction script handles both.
257- ABIDE provides preprocessed derivatives from multiple pipelines (CPAC, DPARSF, CCS); raw data processing via fMRIPrep is recommended for reproducibility.
258- ABIDE data does not include task-fMRI; only resting-state fMRI is available.
259- `abide-skill` is orchestration-only; detailed preprocessing logic remains in `smri-skill` and `fmri-skill`.
260
261---
262
263## When to Call This Skill
264- User asks for end-to-end ABIDE workflow.
265- User asks to download ABIDE data and then run sMRI/rs-fMRI processing.
266- User needs BIDS staging for raw ABIDE NIfTI files.
267- User asks to extract and merge ABIDE phenotype tables.
268- User asks for ABIDE-specific QC summaries and exclusion lists.
269- User needs a single entry point for ABIDE multimodal orchestration.
270
271---
272
273## Complementary / Related Skills
274- `smri-skill`
275- `fmri-skill`
276- `bids-organizer`
277- `fmriprep-tool`
278- `freesurfer-tool`
279- `neurostorm`
280- `brain_gnn`
281- `dependency-planner`
282- `conda-env-manager`
283- `claw-shell`
284
285---
286
287## Reference
288- ABIDE I: https://fcon_1000.projects.nitrc.org/indi/abide/
289- ABIDE II: https://fcon_1000.projects.nitrc.org/indi/abide_II.html
290- Di Martino et al., 2014, *The autism brain imaging data exchange: towards a large-scale evaluation of the intrinsic brain architecture in autism*
291- BIDS spec: https://bids.neuroimaging.io/
292
293Created At: 2026-05-06 01:45 HKT
294Last Updated At: 2026-05-06 01:45 HKT
295Author: chengwang96