HCP-YA Skill (Dataset-Orchestration Layer)
Overview
hcpya-skill is the NeuroClaw orchestration skill for the HCP Young Adult (HCP-YA / HCP1200) dataset.
It strictly follows the NeuroClaw hierarchical design principles:
- This skill only describes WHAT needs to be done and which tool skill to delegate to.
- It contains no implementation code or concrete commands.
- All concrete execution is delegated to existing base/tool skills via
claw-shell.
- Companion scripts in
scripts/ provide reference implementations for data reorganization, phenotype extraction, and QC.
Core workflow (never bypassed):
- Identify input HCP-YA data and target modalities.
- Generate a numbered execution plan clearly stating WHAT needs to be done and which tool skill will handle each step.
- Present the full plan, estimated runtime, resource requirements, and risks to the user and wait for explicit confirmation ("YES" / "execute" / "proceed").
- On confirmation, delegate every step to the appropriate skill via
claw-shell.
- After execution, save all outputs in a clean directory structure (
hcpya_output/).
Research use only.
Quick Reference
| Task |
What needs to be done |
Delegate to |
Expected output |
| Data download |
Download HCP-YA from ConnectomeDB via NeuroSTORM scripts |
claw-shell |
Raw HCP-YA files |
| BIDS staging |
Reorganize HCP-YA native layout to BIDS |
scripts/reorganize_hcpya.py |
BIDS-compliant dataset |
| sMRI processing |
Brain extraction, tissue segmentation, cortical reconstruction |
smri-skill |
smri_output/ derivatives |
| fMRI processing |
Preprocessing, denoising, connectivity, task GLM |
fmri-skill |
fmri_output/ derivatives |
| dMRI processing |
Eddy correction, tensor metrics, tractography |
dwi-skill |
dwi_output/ metrics |
| Phenotype extraction |
Cognitive, behavioral, demographic data |
scripts/extract_hcpya_phenotype.py |
Merged phenotype CSV |
| QC summary |
Per-subject quality control |
scripts/hcpya_qc_summary.py |
QC summary + exclusion list |
Download Stage (Mandatory First Step)
Source
HCP-YA data is distributed through ConnectomeDB:
Supported Download Entry Scripts
download_HCP_1200_all.py (all modalities)
download_HCP_1200_rfMRI.py (resting-state fMRI)
download_HCP_1200_tfMRI.py (task fMRI)
download_HCP_1200_t1t2.py (structural T1w/T2w)
all_pid.pkl (subject list metadata)
Download Inputs to Confirm in Plan
- ConnectomeDB credentials/token
- Target subset (
all, rfMRI, tfMRI, t1t2)
- Subject list scope (full 1,200 or custom subset)
- Destination directory with sufficient disk space (~80 TB for full dataset)
HCP-YA Task Paradigms
| Task |
Description |
Duration |
| MOTOR |
Finger tapping, toe movement, tongue movement |
~3 min |
| EMOTION |
Faces and shapes matching |
~2 min |
| GAMBLING |
Card guessing with reward/loss |
~3 min |
| LANGUAGE |
Story comprehension and math |
~4 min |
| RELATIONAL |
Relational reasoning matching |
~3 min |
| SOCIAL |
Social cognition (mentalizing) movie clips |
~3 min |
| WM |
Working memory (faces, places, tools, body parts) |
~5 min |
| REST |
Resting-state (eyes open) |
~15 min × 4 runs |
BIDS Preparation
Script: scripts/reorganize_hcpya.py
Converts HCP-YA native directory structure to BIDS-compliant layout.
python skills/hcpya-skill/scripts/reorganize_hcpya.py \
--input /path/to/HCPYA/raw \
--output /path/to/HCPYA/bids \
--participants /path/to/subject_list.txt
Features:
- Subject ID normalization: HCP format (
100307) to BIDS (sub-100307)
- Modality routing: T1w, T2w, dMRI, rs-fMRI, task-fMRI (7 tasks)
- Sidecar JSON generation from HCP metadata
dataset_description.json and participants.tsv generation
- Dry-run mode:
--dry-run to preview without copying
Core Workflow (Never Bypassed)
- Identify user target: full HCP-YA processing, imaging subset, phenotype extraction, or BIDS staging only.
- Generate a numbered plan with tools, outputs, runtime, storage, and risks.
- Wait for explicit confirmation (
YES / execute / proceed).
- On confirmation, run download stage first (if needed).
- After download success, run BIDS preparation using
scripts/reorganize_hcpya.py.
- Delegate to
smri-skill for structural MRI processing.
- Delegate to
fmri-skill for functional MRI processing.
- Delegate to
dwi-skill for diffusion MRI processing.
- If phenotype extraction is requested, run
scripts/extract_hcpya_phenotype.py.
- If QC summary is requested, run
scripts/hcpya_qc_summary.py.
- Save outputs into
hcpya_output/.
Modality Processing Delegation
| Modality |
Delegated skill |
Typical tasks |
Main outputs |
| sMRI (T1w/T2w) |
smri-skill |
brain extraction, tissue segmentation, cortical reconstruction, ROI morphometry |
smri_output/ derivatives |
| fMRI (rs-fMRI/task-fMRI) |
fmri-skill |
preprocessing, denoising, ROI time series, connectivity, task GLM |
fmri_output/ derivatives |
| dMRI (DWI) |
dwi-skill |
eddy correction, tensor metrics, tractography, connectome |
dwi_output/ metrics |
Standard Output Layout
hcpya_output/
├── raw/ # Downloaded original HCP-YA files
├── bids/ # BIDS-staged data
├── smri/ # Structural MRI derivatives
├── fmri/ # Functional MRI derivatives
├── dwi/ # Diffusion MRI derivatives
├── phenotype/ # Merged phenotype tables
├── qc/ # QC summaries and exclusion lists
└── logs/ # Download + orchestration logs
Benchmark Adapter Guidance
For benchmark-style prompts, do not force the full download -> staging -> multimodal processing orchestration when the task only asks for local HCP-YA data staging or organization.
- If the task starts from raw HCP-YA data already present on disk and only asks for BIDS-style staging:
- Skip the mandatory download stage
- Default to the narrow path
local raw HCP-YA discovery -> BIDS-style staging -> minimal metadata -> validation/report
- In benchmark mode, do not require explicit confirmation before presenting the direct staging solution.
- Only use the full multimodal orchestration when the prompt explicitly asks for download or end-to-end processing.
Safety and Execution Policy
- No execution before explicit plan confirmation.
- All execution must be routed via
claw-shell.
- Missing dependencies must be resolved by
dependency-planner before running.
- If download fails for partial subjects, continue batch with clear failure report and retry list.
Important Notes and Limitations
- HCP-YA processing is resource intensive (CPU, RAM, and storage).
- Full HCP-YA dataset is ~80 TB; plan storage accordingly.
- HCP-YA has 1,200 subjects with complete multimodal data.
- Age range: 22-35 years.
- For HCP-native preprocessing (minimal preprocessing pipelines), optionally delegate to
hcppipeline-tool.
hcpya-skill is orchestration-only; detailed preprocessing logic remains in modality skills.
When to Call This Skill
- User asks for end-to-end HCP Young Adult workflow.
- User asks to download HCP1200 and run sMRI/fMRI/DTI processing.
- User needs BIDS staging for HCP-YA data.
- User asks to extract HCP-YA phenotype data (cognitive, behavioral, demographic).
Complementary / Related Skills
smri-skill → structural MRI preprocessing
fmri-skill → functional MRI preprocessing and analysis
dwi-skill → diffusion MRI preprocessing and analysis
hcppipeline-tool → HCP-native minimal preprocessing pipelines
bids-organizer → BIDS validation and organization
brain-visualization → visualization of derivatives
dependency-planner → dependency resolution
conda-env-manager → environment management
claw-shell → command execution
Reference
Created At: 2026-05-06 13:02 HKT
Last Updated At: 2026-05-06 13:02 HKT
Author: chengwang96
1---2name: hcpya-skill3description: Use this skill whenever the user wants an end-to-end workflow for the HCP Young Adult (HCP-YA / HCP1200) dataset, including dataset download, BIDS organization, and multimodal processing of sMRI, fMRI, and dMRI. Triggers include: 'HCP Young Adult', 'HCP-YA', 'HCP1200', 'process HCP data', 'HCP sMRI fMRI DTI', or any request to run the HCP-YA multimodal pipeline.4license: MIT License (NeuroClaw custom skill - freely modifiable within t5---6# HCP-YA Skill (Dataset-Orchestration Layer)
7
8## Overview
9
10`hcpya-skill` is the NeuroClaw orchestration skill for the **HCP Young Adult (HCP-YA / HCP1200)** dataset.
11
12It strictly follows the NeuroClaw hierarchical design principles:
13- This skill **only describes WHAT needs to be done** and **which tool skill to delegate to**.
14- It contains **no implementation code or concrete commands**.
15- All concrete execution is delegated to existing base/tool skills via `claw-shell`.
16- Companion scripts in `scripts/` provide reference implementations for data reorganization, phenotype extraction, and QC.
17
18**Core workflow (never bypassed):**
191. Identify input HCP-YA data and target modalities.
202. Generate a **numbered execution plan** clearly stating WHAT needs to be done and which tool skill will handle each step.
213. Present the full plan, estimated runtime, resource requirements, and risks to the user and wait for explicit confirmation ("YES" / "execute" / "proceed").
224. On confirmation, delegate every step to the appropriate skill via `claw-shell`.
235. After execution, save all outputs in a clean directory structure (`hcpya_output/`).
24
25**Research use only.**
26
27---
28
29## Quick Reference
30
31| Task | What needs to be done | Delegate to | Expected output |
32|---|---|---|---|
33| Data download | Download HCP-YA from ConnectomeDB via NeuroSTORM scripts | `claw-shell` | Raw HCP-YA files |
34| BIDS staging | Reorganize HCP-YA native layout to BIDS | `scripts/reorganize_hcpya.py` | BIDS-compliant dataset |
35| sMRI processing | Brain extraction, tissue segmentation, cortical reconstruction | `smri-skill` | `smri_output/` derivatives |
36| fMRI processing | Preprocessing, denoising, connectivity, task GLM | `fmri-skill` | `fmri_output/` derivatives |
37| dMRI processing | Eddy correction, tensor metrics, tractography | `dwi-skill` | `dwi_output/` metrics |
38| Phenotype extraction | Cognitive, behavioral, demographic data | `scripts/extract_hcpya_phenotype.py` | Merged phenotype CSV |
39| QC summary | Per-subject quality control | `scripts/hcpya_qc_summary.py` | QC summary + exclusion list |
40
41---
42
43## Download Stage (Mandatory First Step)
44
45### Source
46HCP-YA data is distributed through **ConnectomeDB**:
47- Website: https://db.humanconnectome.org/
48- Requires ConnectomeDB account and data use agreement
49- NeuroSTORM download scripts available at: https://github.com/CUHK-AIM-Group/NeuroSTORM/tree/main/scripts/dataset_download
50
51### Supported Download Entry Scripts
52- `download_HCP_1200_all.py` (all modalities)
53- `download_HCP_1200_rfMRI.py` (resting-state fMRI)
54- `download_HCP_1200_tfMRI.py` (task fMRI)
55- `download_HCP_1200_t1t2.py` (structural T1w/T2w)
56- `all_pid.pkl` (subject list metadata)
57
58### Download Inputs to Confirm in Plan
59- ConnectomeDB credentials/token
60- Target subset (`all`, `rfMRI`, `tfMRI`, `t1t2`)
61- Subject list scope (full 1,200 or custom subset)
62- Destination directory with sufficient disk space (~80 TB for full dataset)
63
64---
65
66## HCP-YA Task Paradigms
67
68| Task | Description | Duration |
69|---|---|---|
70| MOTOR | Finger tapping, toe movement, tongue movement | ~3 min |
71| EMOTION | Faces and shapes matching | ~2 min |
72| GAMBLING | Card guessing with reward/loss | ~3 min |
73| LANGUAGE | Story comprehension and math | ~4 min |
74| RELATIONAL | Relational reasoning matching | ~3 min |
75| SOCIAL | Social cognition (mentalizing) movie clips | ~3 min |
76| WM | Working memory (faces, places, tools, body parts) | ~5 min |
77| REST | Resting-state (eyes open) | ~15 min × 4 runs |
78
79---
80
81## BIDS Preparation
82
83### Script: `scripts/reorganize_hcpya.py`
84
85Converts HCP-YA native directory structure to BIDS-compliant layout.
86
87```bash
88python skills/hcpya-skill/scripts/reorganize_hcpya.py \
89 --input /path/to/HCPYA/raw \
90 --output /path/to/HCPYA/bids \
91 --participants /path/to/subject_list.txt
92```
93
94Features:
95- Subject ID normalization: HCP format (`100307`) to BIDS (`sub-100307`)
96- Modality routing: T1w, T2w, dMRI, rs-fMRI, task-fMRI (7 tasks)
97- Sidecar JSON generation from HCP metadata
98- `dataset_description.json` and `participants.tsv` generation
99- Dry-run mode: `--dry-run` to preview without copying
100
101---
102
103## Core Workflow (Never Bypassed)
104
1051. Identify user target: full HCP-YA processing, imaging subset, phenotype extraction, or BIDS staging only.
1062. Generate a numbered plan with tools, outputs, runtime, storage, and risks.
1073. Wait for explicit confirmation (`YES` / `execute` / `proceed`).
1084. On confirmation, run download stage first (if needed).
1095. After download success, run BIDS preparation using `scripts/reorganize_hcpya.py`.
1106. Delegate to `smri-skill` for structural MRI processing.
1117. Delegate to `fmri-skill` for functional MRI processing.
1128. Delegate to `dwi-skill` for diffusion MRI processing.
1139. If phenotype extraction is requested, run `scripts/extract_hcpya_phenotype.py`.
11410. If QC summary is requested, run `scripts/hcpya_qc_summary.py`.
11511. Save outputs into `hcpya_output/`.
116
117---
118
119## Modality Processing Delegation
120
121| Modality | Delegated skill | Typical tasks | Main outputs |
122|---|---|---|---|
123| sMRI (T1w/T2w) | `smri-skill` | brain extraction, tissue segmentation, cortical reconstruction, ROI morphometry | `smri_output/` derivatives |
124| fMRI (rs-fMRI/task-fMRI) | `fmri-skill` | preprocessing, denoising, ROI time series, connectivity, task GLM | `fmri_output/` derivatives |
125| dMRI (DWI) | `dwi-skill` | eddy correction, tensor metrics, tractography, connectome | `dwi_output/` metrics |
126
127---
128
129## Standard Output Layout
130
131```
132hcpya_output/
133├── raw/ # Downloaded original HCP-YA files
134├── bids/ # BIDS-staged data
135├── smri/ # Structural MRI derivatives
136├── fmri/ # Functional MRI derivatives
137├── dwi/ # Diffusion MRI derivatives
138├── phenotype/ # Merged phenotype tables
139├── qc/ # QC summaries and exclusion lists
140└── logs/ # Download + orchestration logs
141```
142
143---
144
145## Benchmark Adapter Guidance
146
147For benchmark-style prompts, do not force the full `download -> staging -> multimodal processing` orchestration when the task only asks for local HCP-YA data staging or organization.
148
149- If the task starts from raw HCP-YA data already present on disk and only asks for BIDS-style staging:
150 - Skip the mandatory download stage
151 - Default to the narrow path `local raw HCP-YA discovery -> BIDS-style staging -> minimal metadata -> validation/report`
152- In benchmark mode, do not require explicit confirmation before presenting the direct staging solution.
153- Only use the full multimodal orchestration when the prompt explicitly asks for download or end-to-end processing.
154
155---
156
157## Safety and Execution Policy
158- No execution before explicit plan confirmation.
159- All execution must be routed via `claw-shell`.
160- Missing dependencies must be resolved by `dependency-planner` before running.
161- If download fails for partial subjects, continue batch with clear failure report and retry list.
162
163---
164
165## Important Notes and Limitations
166- HCP-YA processing is resource intensive (CPU, RAM, and storage).
167- Full HCP-YA dataset is ~80 TB; plan storage accordingly.
168- HCP-YA has 1,200 subjects with complete multimodal data.
169- Age range: 22-35 years.
170- For HCP-native preprocessing (minimal preprocessing pipelines), optionally delegate to `hcppipeline-tool`.
171- `hcpya-skill` is orchestration-only; detailed preprocessing logic remains in modality skills.
172
173---
174
175## When to Call This Skill
176- User asks for end-to-end HCP Young Adult workflow.
177- User asks to download HCP1200 and run sMRI/fMRI/DTI processing.
178- User needs BIDS staging for HCP-YA data.
179- User asks to extract HCP-YA phenotype data (cognitive, behavioral, demographic).
180
181---
182
183## Complementary / Related Skills
184- `smri-skill` → structural MRI preprocessing
185- `fmri-skill` → functional MRI preprocessing and analysis
186- `dwi-skill` → diffusion MRI preprocessing and analysis
187- `hcppipeline-tool` → HCP-native minimal preprocessing pipelines
188- `bids-organizer` → BIDS validation and organization
189- `brain-visualization` → visualization of derivatives
190- `dependency-planner` → dependency resolution
191- `conda-env-manager` → environment management
192- `claw-shell` → command execution
193
194---
195
196## Reference
197- HCP-YA: https://www.humanconnectome.org/study/hcp-young-adult
198- ConnectomeDB: https://db.humanconnectome.org/
199- NeuroSTORM download scripts: https://github.com/CUHK-AIM-Group/NeuroSTORM/tree/main/scripts/dataset_download
200- Glasser et al. (2013): The Human Connectome Project minimally preprocessed pipelines
201
202Created At: 2026-05-06 13:02 HKT
203Last Updated At: 2026-05-06 13:02 HKT
204Author: chengwang96