Run Models Skill (Model Entry Layer)
Overview
run_models is the NeuroClaw entry skill for model-level inference workflows.
This skill is responsible for:
- Maintaining a model registry (name, paper, source code, input/output, doc file path)
- Selecting the correct model skill under
skills/<model-name>/SKILL.md
- Coordinating required data preparation before model execution
- Delegating modality preprocessing to
fmri-skill and smri-skill
It supports both:
- deep learning model routes for phenotype prediction
- non-deep-learning statistical / unsupervised / classical machine-learning routes such as first-level and second-level task-fMRI GLM, resting-state ICA, resting-state DictLearning, disease classification with SVM, disease classification with SpaceNet, brain parcellation with K-means, brain parcellation with Hierarchical clustering, temporal filtering, and detrending
This skill does not hardcode detailed install/run commands for each model. Those details are stored in model-specific markdown files.
Research use only.
Core Workflow (Never Bypassed)
- Identify requested model and task (classification/regression phenotype prediction).
- Locate the corresponding model skill under
skills/<model-name>/SKILL.md.
- Verify required inputs (ROI features, optional sMRI features).
- If inputs are not ready, delegate preprocessing to modality skills:
fmri-skill for ROI extraction from fMRI
smri-skill when model additionally requires structural features
- Generate a numbered execution plan and wait for explicit user confirmation (
YES / execute / proceed).
- On confirmation, execute via
claw-shell following model doc instructions.
Model Registry (Current)
| Model |
Paper |
Code |
Input |
Output |
Model Doc |
| BrainGNN |
Li et al., 2020, Braingnn: Interpretable brain graph neural network for fmri analysis |
https://github.com/xxlya/BrainGNN_Pytorch/tree/main |
fMRI ROI features (graph/node-level ROI representation) |
Phenotype prediction (classification/regression) + interpretable graph indicators |
skills/brain_gnn/SKILL.md |
| BNT |
Kan et al., 2022, BrainNetworkTransformer |
https://github.com/Wayfear/BrainNetworkTransformer |
fMRI ROI FC matrix (dense [N, N], no PyG) |
Phenotype prediction (classification/regression) + attention weights + DEC cluster assignments |
skills/bnt/SKILL.md |
| FM-APP |
He et al., 2024, FM-APP: Foundation model for any phenotype prediction via fMRI to sMRI knowledge transfer |
https://github.com/ZhibinHe/FM-APP |
fMRI ROI features + sMRI features |
Phenotype prediction (any-phenotype setting) |
skills/fm_app/SKILL.md |
| NeuroStorm |
NeuroClaw model entry for storm-related phenotype prediction workflows |
see skills/neurostorm/SKILL.md |
Multi-modal neuroimaging features as specified in the model doc |
Phenotype prediction / downstream inference as specified in the model doc |
skills/neurostorm/SKILL.md |
| GLM |
Classical first-level and second-level task-fMRI general linear model |
Nilearn / SPM-style implementation route |
Preprocessed task fMRI, events, optional confounds, and optional subject-level contrast maps for group inference |
Task activation contrasts, group z maps, and statistical inference outputs |
skills/glm/SKILL.md |
| ICA |
Classical resting-state network decomposition method |
Nilearn decomposition implementation route |
Preprocessed resting-state fMRI, optional mask, optional confounds |
Intrinsic connectivity component maps, subject time series, optional connectomes |
skills/ica/SKILL.md |
| DictLearning |
Classical sparse resting-state network decomposition method |
Nilearn decomposition implementation route |
Preprocessed resting-state fMRI, optional mask, optional confounds |
Sparse component maps, subject time series, optional connectomes |
skills/dictlearning/SKILL.md |
| SVM |
Classical disease classification method for neuroimaging |
Nilearn / scikit-learn style decoding route |
Preprocessed ROI features, labels, optional covariates |
Predicted labels, decision scores, CV metrics |
skills/svm/SKILL.md |
| SpaceNet |
Classical voxel-wise disease classification method for neuroimaging |
Nilearn decoding implementation route |
Aligned voxel maps, labels, optional covariates, optional mask |
Predicted labels, decision scores, CV metrics, coefficient maps |
skills/spacenet/SKILL.md |
| K-means |
Classical brain parcellation method for neuroimaging |
Nilearn / clustering-based parcellation route |
Preprocessed feature maps or image lists, optional mask, requested parcel count |
Parcel labels, cluster summaries, optional centroid outputs |
skills/kmeans/SKILL.md |
| Hierarchical |
Classical hierarchical brain parcellation method for neuroimaging |
Nilearn / clustering-based parcellation route |
Preprocessed feature maps or image lists, optional mask, requested parcel count |
Parcel labels, cluster summaries, optional dendrogram outputs |
skills/hierarchical/SKILL.md |
| Filtering |
Classical signal denoising method for neuroimaging time series |
Nilearn / preprocessing route |
Preprocessed BOLD image or time series, TR, optional confounds, optional mask |
Denoised BOLD, cleaned time series, optional QC summaries |
skills/filtering/SKILL.md |
| Detrending |
Classical signal denoising method for neuroimaging time series |
Nilearn / preprocessing route |
Preprocessed BOLD image or time series, TR, optional confounds, optional mask |
Cleaned BOLD, cleaned time series, optional QC summaries |
skills/detrending/SKILL.md |
Cross-Cutting Tools (Apply Across Models)
These are not models. They are horizontal layers that any model in the registry above can opt into without changing model code.
| Tool |
Purpose |
When to invoke |
Tool Doc |
| harmonization-tool |
Remove site/scanner/batch effects from features before model training; supports ComBat / ComBat-GAM / CovBat / site-as-covariate; ships site-stratified and leave-site-out splitters; required for honest mega-analysis across multi-site cohorts |
Any multi-site or multi-dataset run (ABIDE, ADHD-200, ABCD, multi-cohort pooling); user mentions ComBat / harmonize / site effect / cross-site / mega-analysis |
skills/harmonization-tool/SKILL.md |
Insertion point: between dataset-skill output (feature matrix + meta) and model-skill input. Models read harmonized features identically to raw features.
Citation Notes
- BrainGNN:
- Li X, Zhou Y, Dvornek N, Zhang M, Gao S, Zhuang J, Scheinost D, Staib L, Ventola P, Duncan J. 2020.
- BNT:
- Kan X, Dai W, Cui H, Zhang Z, Guo Y, He L. 2022. BrainNetworkTransformer. NeurIPS.
- FM-APP:
- He Z, Li W, Liu Y, et al. FM-APP. IEEE TMI, 2024, 44(10): 4010-4022.
- NeuroStorm:
- See
skills/neurostorm/SKILL.md for the current model card, citation, and execution details.
- GLM:
- Classical first-level and second-level general linear model for task-evoked activation analysis and group-level inference; see
skills/glm/SKILL.md.
- ICA:
- Classical resting-state network decomposition route based on independent component analysis; see
skills/ica/SKILL.md.
- DictLearning:
- Classical sparse resting-state network decomposition route; see
skills/dictlearning/SKILL.md.
- SVM:
- Classical disease classification route for ROI-level or tabular decoding; see
skills/svm/SKILL.md.
- SpaceNet:
- Classical voxel-wise disease classification route with sparse coefficient maps; see
skills/spacenet/SKILL.md.
- K-means:
- Classical brain parcellation route for fixed-K parcel discovery; see
skills/kmeans/SKILL.md.
- Hierarchical:
- Classical brain parcellation route for multi-scale parcel discovery; see
skills/hierarchical/SKILL.md.
- Filtering:
- Classical signal denoising route for temporal filtering; see
skills/filtering/SKILL.md.
- Detrending:
- Classical signal denoising route for temporal drift removal; see
skills/detrending/SKILL.md.
Harness-Aware Model Registration (Declarative + Testing + Drift Detection)
Model Specification Format (Extended)
Every model integrated into run_models must include a model specification file in JSON format alongside its Markdown documentation:
File: skills/{model_name}/{model_name}_spec.json
{
"model_name": "brain_gnn",
"version": "1.0.0",
"paper": "Li et al., 2020",
"code_repo": "https://github.com/xxlya/BrainGNN_Pytorch",
"required_dependencies": {
"torch": ">=1.9.0,<2.1.0",
"numpy": ">=1.21.0",
"scipy": ">=1.7.0",
"networkx": ">=2.6.0"
},
"input_spec": {
"modality": "fMRI",
"format": "ROI time-series (N_nodes, T_timepoints)",
"expected_shape": [116, null],
"value_range": [-5.0, 5.0],
"required_preprocessing": ["z-score normalization"]
},
"output_spec": {
"type": "classification|regression",
"classes": null,
"value_range": null
},
"validation_checksums": {
"weights_sha256": "abc123...",
"test_data_sha256": "def456..."
}
}
Test Suite Requirements
Every model must include an automated test suite covering:
- Input validation: verify input dimensions, data types, value ranges
- Determinism check: seed control + verify identical outputs with same seed (tolerance: 1e-6)
- Performance regression: compare inference speed and memory usage against baseline
- Output coherence: verify outputs lie within expected value range, no NaN/Inf values
- Backward compatibility: test model against previous version checksum (if available)
Test execution:
python -m pytest run_models/tests/test_{model_name}.py -v --harness-report
Output: run_models_test_report_{model_name}_{timestamp}.json with pass/fail status and metrics
Drift Detection Protocol
Monitor production/inference results for concept drift (distribution shift in data or model behavior):
Automated monitoring per 100 inferences:
- Input distribution shift (KL divergence against reference data): flag if deviation > 0.1
- Output distribution shift (prediction probability / regression output quantiles): flag if shift detected
- Latency drift (average inference time): alert if >20% increase
- Failure rate monitoring (predictions with NaN/Inf / out-of-range): flag if >1% failures
Logging output: run_models_drift_log.json (append-only, timestamped entries)
Example entry:
{
"timestamp": "2026-04-05T14:32:00Z",
"model": "brain_gnn",
"inference_count": 100,
"input_kl_divergence": 0.045,
"output_mean_shift": 0.002,
"latency_ms": 45.2,
"failure_rate": 0.0,
"status": "healthy"
}
Alert thresholds:
- KL divergence > 0.1 → generate warning
- Output shift > 5% std dev → investigation recommended
- Latency drift > 20% → check computational resource bottleneck
- Failure rate > 1% → stop inference, require manual review
Model Card Template (Minimum Required Metadata)
Each model must include a model card in skills/{model_name}/SKILL.md documenting:
## Model Card: {model_name}
### Model Details
- **Model name**: {name}
- **Version**: {X.Y.Z}
- **Date**: {YYYY-MM-DD}
- **Source repository**: {repo_url}
- **Paper**: {citation}
### Intended Use
- **Primary use case**: [e.g., fMRI-based phenotype classification]
- **Input modalities**: [fMRI, sMRI, etc.]
- **Supported tasks**: [classification, regression, interpretability]
### Known Limitations
- [e.g., "Trained on N subjects aged 18-65; generalization to pediatric/geriatric populations not validated"]
- [e.g., "Sensitive to head motion artifacts; recommend ICA-FIX preprocessing"]
### Validation Results
- **Test set performance**: [accuracy/AUC/RMSE with confidence intervals]
- **Cross-site validation**: [performance on held-out sites, if applicable]
- **Robustness checks**: [drift detection history, adversarial perturbation results]
### Dependencies & Versioning
- **Required libraries**: [see {model_name}_spec.json]
- **Hash (model weights)**: {SHA256}
- **Last verified**: {date}
Delegation Rules
BrainGNN Route
- Required modality preprocessing:
fmri-skill
- Typical upstream outputs expected: ROI matrices/time-series converted to model-required feature tensors
BNT Route
- Required modality preprocessing:
fmri-skill
- Typical upstream outputs expected: Same ROI .pt files as BrainGNN (shared data source under
data/braingnn_input/)
FM-APP Route
- Required modality preprocessing:
fmri-skill + smri-skill
- Typical upstream outputs expected: fMRI ROI features plus structural MRI-derived features
NeuroStorm Route
- Required modality preprocessing: follow the model doc in
skills/neurostorm/SKILL.md
- Typical upstream outputs expected: inputs and features specified by the NeuroStorm model card
GLM Route
- Required modality preprocessing:
fmri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- first-level GLM: preprocessed task fMRI, events, optional confounds, named contrasts
- second-level GLM: subject-level contrast maps, group design matrix, group contrast definition
ICA Route
- Required modality preprocessing:
fmri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- preprocessed resting-state fMRI image list
- optional mask and confounds
- requested component count
DictLearning Route
- Required modality preprocessing:
fmri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- preprocessed resting-state fMRI image list
- optional mask and confounds
- requested component count
SVM Route
- Required modality preprocessing:
fmri-skill and/or smri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- ROI/tabular feature matrix, diagnosis labels, optional covariates
SpaceNet Route
- Required modality preprocessing:
fmri-skill and/or smri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- aligned subject image list, diagnosis labels, mask image, optional covariates
K-means Route
- Required modality preprocessing:
fmri-skill and/or smri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- feature matrix or aligned image list for parcel discovery
- optional mask
- target parcel count
Hierarchical Route
- Required modality preprocessing:
fmri-skill and/or smri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- feature matrix or aligned image list for parcel discovery
- optional mask or similarity structure
- target parcel count
Filtering Route
- Required modality preprocessing:
fmri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- preprocessed BOLD image or extracted time series
- TR, optional confounds, optional mask
- optional frequency settings
Detrending Route
- Required modality preprocessing:
fmri-skill
- Concrete model/tool execution:
nilearn-tool
- Typical upstream outputs expected:
- preprocessed BOLD image or extracted time series
- TR, optional confounds, optional mask
- detrending request and optional standardization settings
Shared Execution Routing
- Environment/dependency planning:
dependency-planner + conda-env-manager
- Actual model run command execution:
claw-shell
Input and Output Contract (Entry-Level)
Inputs expected by this skill
- Model selection (
brain_gnn, bnt, fm_app, neurostorm, glm, ica, dictlearning, svm, spacenet, kmeans, hierarchical, filtering, or detrending)
- Data split / subject list
- Phenotype target definition
- Optional compute constraints (GPU/CPU, memory, batch size)
For GLM routes, the required task definition should be expressed as:
- task name
- events file
- contrast(s) of interest
- optional group-level analysis scope
- whether the request is first-level GLM or second-level GLM
- if second-level GLM: contrast map list and group design matrix
For ICA routes, the required decomposition definition should be expressed as:
- resting-state image list or subject list
- number of components
- optional mask and confounds
For DictLearning routes, the required decomposition definition should be expressed as:
- resting-state image list or subject list
- number of components
- optional mask and confounds
For SVM routes, the required classification definition should be expressed as:
- diagnosis target / label column
- feature type (
roi/tabular)
- subject list or split definition
- optional covariates
For SpaceNet routes, the required classification definition should be expressed as:
- diagnosis target / label column
- feature type (
voxel-wise)
- subject list or split definition
- optional covariates and mask
For K-means routes, the required parcellation definition should be expressed as:
- image list or feature matrix
- target parcel / cluster count
- optional mask
For Hierarchical routes, the required parcellation definition should be expressed as:
- image list or feature matrix
- target parcel / cluster count
- optional mask, similarity structure, or adjacency constraint
For Filtering routes, the required denoising definition should be expressed as:
- input BOLD image or time series
- TR
- optional confounds, mask, and frequency settings
For Detrending routes, the required denoising definition should be expressed as:
- input BOLD image or time series
- TR
- optional confounds, mask, and standardization settings
Outputs produced by this skill
- A confirmed, numbered run plan
- Pointers to the model-specific instruction file
- Delegated preprocessing plan for required modalities
- Structured output location recommendations
Recommended Output Layout
All model-running artifacts should be managed under ./run_models_output/:
run_models_output/preprocessed/
fmri/ (from fmri-skill)
smri/ (from smri-skill, if required)
run_models_output/brain_gnn/
run_models_output/bnt/
run_models_output/fm_app/
run_models_output/neurostorm/
run_models_output/glm/
run_models_output/ica/
run_models_output/dictlearning/
run_models_output/svm/
run_models_output/spacenet/
run_models_output/kmeans/
run_models_output/hierarchical/
run_models_output/filtering/
run_models_output/detrending/
run_models_output/logs/
run_models_output/reports/
Safety and Execution Policy
- No execution before explicit user confirmation of the numbered plan.
- All run/install actions must go through
claw-shell.
- If model skills are missing in
skills/<model-name>/, stop and request or create them before execution.
- Keep train/val/test split and target definition explicit to avoid leakage.
When to Call This Skill
- User asks to run BrainGNN or FM-APP.
- User asks to run BNT (BrainNetworkTransformer).
- User asks to run NeuroStorm.
- User asks to run classical task activation analysis with GLM.
- User asks to run group-level inference with second-level GLM.
- User asks to perform resting-state network decomposition with ICA.
- User asks to perform resting-state network decomposition with DictLearning.
- User asks to perform disease classification with SVM.
- User asks to perform disease classification with SpaceNet.
- User asks to perform brain parcellation with K-means.
- User asks to perform brain parcellation with Hierarchical clustering.
- User asks to perform signal denoising with filtering.
- User asks to perform signal denoising with detrending.
- User asks which phenotype model to use for fMRI/sMRI ROI data.
- User asks for a unified entry point to model introduction + run routing.
Complementary / Related Skills
fmri-skill
smri-skill
dependency-planner
conda-env-manager
claw-shell
Reference
- BrainGNN paper and code:
- FM-APP paper and code:
- Nilearn GLM documentation:
Created At: 2026-03-28 20:38 HKT
Last Updated At: 2026-04-14 00:28 HKT
Author: chengwang96
1---2name: run-models3description: Use this skill whenever the user wants to run phenotype-prediction models, browse model cards, map model inputs/outputs, or choose an execution route for fMRI/sMRI based models. This is a model-entry orchestration skill: it routes requests to model-specific docs and delegates preprocessing to modality skills.4license: MIT License (NeuroClaw custom skill - freely modifiable within t5---6# Run Models Skill (Model Entry Layer)
7
8## Overview
9`run_models` is the NeuroClaw entry skill for model-level inference workflows.
10
11This skill is responsible for:
12- Maintaining a model registry (name, paper, source code, input/output, doc file path)
13- Selecting the correct model skill under `skills/<model-name>/SKILL.md`
14- Coordinating required data preparation before model execution
15- Delegating modality preprocessing to `fmri-skill` and `smri-skill`
16
17It supports both:
18- deep learning model routes for phenotype prediction
19- non-deep-learning statistical / unsupervised / classical machine-learning routes such as first-level and second-level task-fMRI GLM, resting-state ICA, resting-state DictLearning, disease classification with SVM, disease classification with SpaceNet, brain parcellation with K-means, brain parcellation with Hierarchical clustering, temporal filtering, and detrending
20
21This skill does not hardcode detailed install/run commands for each model. Those details are stored in model-specific markdown files.
22
23**Research use only.**
24
25---
26
27## Core Workflow (Never Bypassed)
281. Identify requested model and task (classification/regression phenotype prediction).
292. Locate the corresponding model skill under `skills/<model-name>/SKILL.md`.
303. Verify required inputs (ROI features, optional sMRI features).
314. If inputs are not ready, delegate preprocessing to modality skills:
32 - `fmri-skill` for ROI extraction from fMRI
33 - `smri-skill` when model additionally requires structural features
345. Generate a numbered execution plan and wait for explicit user confirmation (`YES` / `execute` / `proceed`).
356. On confirmation, execute via `claw-shell` following model doc instructions.
36
37---
38
39## Model Registry (Current)
40
41| Model | Paper | Code | Input | Output | Model Doc |
42|---|---|---|---|---|---|
43| BrainGNN | Li et al., 2020, *Braingnn: Interpretable brain graph neural network for fmri analysis* | https://github.com/xxlya/BrainGNN_Pytorch/tree/main | fMRI ROI features (graph/node-level ROI representation) | Phenotype prediction (classification/regression) + interpretable graph indicators | `skills/brain_gnn/SKILL.md` |
44| BNT | Kan et al., 2022, *BrainNetworkTransformer* | https://github.com/Wayfear/BrainNetworkTransformer | fMRI ROI FC matrix (dense [N, N], no PyG) | Phenotype prediction (classification/regression) + attention weights + DEC cluster assignments | `skills/bnt/SKILL.md` |
45| FM-APP | He et al., 2024, *FM-APP: Foundation model for any phenotype prediction via fMRI to sMRI knowledge transfer* | https://github.com/ZhibinHe/FM-APP | fMRI ROI features + sMRI features | Phenotype prediction (any-phenotype setting) | `skills/fm_app/SKILL.md` |
46| NeuroStorm | NeuroClaw model entry for storm-related phenotype prediction workflows | see `skills/neurostorm/SKILL.md` | Multi-modal neuroimaging features as specified in the model doc | Phenotype prediction / downstream inference as specified in the model doc | `skills/neurostorm/SKILL.md` |
47| GLM | Classical first-level and second-level task-fMRI general linear model | Nilearn / SPM-style implementation route | Preprocessed task fMRI, events, optional confounds, and optional subject-level contrast maps for group inference | Task activation contrasts, group z maps, and statistical inference outputs | `skills/glm/SKILL.md` |
48| ICA | Classical resting-state network decomposition method | Nilearn decomposition implementation route | Preprocessed resting-state fMRI, optional mask, optional confounds | Intrinsic connectivity component maps, subject time series, optional connectomes | `skills/ica/SKILL.md` |
49| DictLearning | Classical sparse resting-state network decomposition method | Nilearn decomposition implementation route | Preprocessed resting-state fMRI, optional mask, optional confounds | Sparse component maps, subject time series, optional connectomes | `skills/dictlearning/SKILL.md` |
50| SVM | Classical disease classification method for neuroimaging | Nilearn / scikit-learn style decoding route | Preprocessed ROI features, labels, optional covariates | Predicted labels, decision scores, CV metrics | `skills/svm/SKILL.md` |
51| SpaceNet | Classical voxel-wise disease classification method for neuroimaging | Nilearn decoding implementation route | Aligned voxel maps, labels, optional covariates, optional mask | Predicted labels, decision scores, CV metrics, coefficient maps | `skills/spacenet/SKILL.md` |
52| K-means | Classical brain parcellation method for neuroimaging | Nilearn / clustering-based parcellation route | Preprocessed feature maps or image lists, optional mask, requested parcel count | Parcel labels, cluster summaries, optional centroid outputs | `skills/kmeans/SKILL.md` |
53| Hierarchical | Classical hierarchical brain parcellation method for neuroimaging | Nilearn / clustering-based parcellation route | Preprocessed feature maps or image lists, optional mask, requested parcel count | Parcel labels, cluster summaries, optional dendrogram outputs | `skills/hierarchical/SKILL.md` |
54| Filtering | Classical signal denoising method for neuroimaging time series | Nilearn / preprocessing route | Preprocessed BOLD image or time series, TR, optional confounds, optional mask | Denoised BOLD, cleaned time series, optional QC summaries | `skills/filtering/SKILL.md` |
55| Detrending | Classical signal denoising method for neuroimaging time series | Nilearn / preprocessing route | Preprocessed BOLD image or time series, TR, optional confounds, optional mask | Cleaned BOLD, cleaned time series, optional QC summaries | `skills/detrending/SKILL.md` |
56
57### Cross-Cutting Tools (Apply Across Models)
58These are not models. They are horizontal layers that any model in the registry above can opt into without changing model code.
59
60| Tool | Purpose | When to invoke | Tool Doc |
61|---|---|---|---|
62| harmonization-tool | Remove site/scanner/batch effects from features before model training; supports ComBat / ComBat-GAM / CovBat / site-as-covariate; ships site-stratified and leave-site-out splitters; required for honest mega-analysis across multi-site cohorts | Any multi-site or multi-dataset run (ABIDE, ADHD-200, ABCD, multi-cohort pooling); user mentions ComBat / harmonize / site effect / cross-site / mega-analysis | `skills/harmonization-tool/SKILL.md` |
63
64Insertion point: between dataset-skill output (feature matrix + meta) and model-skill input. Models read harmonized features identically to raw features.
65
66### Citation Notes
67- BrainGNN:
68 - Li X, Zhou Y, Dvornek N, Zhang M, Gao S, Zhuang J, Scheinost D, Staib L, Ventola P, Duncan J. 2020.
69- BNT:
70 - Kan X, Dai W, Cui H, Zhang Z, Guo Y, He L. 2022. BrainNetworkTransformer. NeurIPS.
71- FM-APP:
72 - He Z, Li W, Liu Y, et al. FM-APP. IEEE TMI, 2024, 44(10): 4010-4022.
73- NeuroStorm:
74 - See `skills/neurostorm/SKILL.md` for the current model card, citation, and execution details.
75- GLM:
76 - Classical first-level and second-level general linear model for task-evoked activation analysis and group-level inference; see `skills/glm/SKILL.md`.
77- ICA:
78 - Classical resting-state network decomposition route based on independent component analysis; see `skills/ica/SKILL.md`.
79- DictLearning:
80 - Classical sparse resting-state network decomposition route; see `skills/dictlearning/SKILL.md`.
81- SVM:
82 - Classical disease classification route for ROI-level or tabular decoding; see `skills/svm/SKILL.md`.
83- SpaceNet:
84 - Classical voxel-wise disease classification route with sparse coefficient maps; see `skills/spacenet/SKILL.md`.
85- K-means:
86 - Classical brain parcellation route for fixed-K parcel discovery; see `skills/kmeans/SKILL.md`.
87- Hierarchical:
88 - Classical brain parcellation route for multi-scale parcel discovery; see `skills/hierarchical/SKILL.md`.
89- Filtering:
90 - Classical signal denoising route for temporal filtering; see `skills/filtering/SKILL.md`.
91- Detrending:
92 - Classical signal denoising route for temporal drift removal; see `skills/detrending/SKILL.md`.
93
94## Harness-Aware Model Registration (Declarative + Testing + Drift Detection)
95
96### Model Specification Format (Extended)
97Every model integrated into run_models **must** include a **model specification file** in JSON format alongside its Markdown documentation:
98
99**File**: `skills/{model_name}/{model_name}_spec.json`
100
101```json
102{
103 "model_name": "brain_gnn",
104 "version": "1.0.0",
105 "paper": "Li et al., 2020",
106 "code_repo": "https://github.com/xxlya/BrainGNN_Pytorch",
107 "required_dependencies": {
108 "torch": ">=1.9.0,<2.1.0",
109 "numpy": ">=1.21.0",
110 "scipy": ">=1.7.0",
111 "networkx": ">=2.6.0"
112 },
113 "input_spec": {
114 "modality": "fMRI",
115 "format": "ROI time-series (N_nodes, T_timepoints)",
116 "expected_shape": [116, null],
117 "value_range": [-5.0, 5.0],
118 "required_preprocessing": ["z-score normalization"]
119 },
120 "output_spec": {
121 "type": "classification|regression",
122 "classes": null,
123 "value_range": null
124 },
125 "validation_checksums": {
126 "weights_sha256": "abc123...",
127 "test_data_sha256": "def456..."
128 }
129}
130```
131
132### Test Suite Requirements
133Every model **must** include an automated test suite covering:
134
1351. **Input validation**: verify input dimensions, data types, value ranges
1362. **Determinism check**: seed control + verify identical outputs with same seed (tolerance: 1e-6)
1373. **Performance regression**: compare inference speed and memory usage against baseline
1384. **Output coherence**: verify outputs lie within expected value range, no NaN/Inf values
1395. **Backward compatibility**: test model against previous version checksum (if available)
140
141**Test execution**:
142```bash
143python -m pytest run_models/tests/test_{model_name}.py -v --harness-report
144```
145
146Output: `run_models_test_report_{model_name}_{timestamp}.json` with pass/fail status and metrics
147
148### Drift Detection Protocol
149Monitor production/inference results for concept drift (distribution shift in data or model behavior):
150
151**Automated monitoring per 100 inferences**:
152- **Input distribution shift** (KL divergence against reference data): flag if deviation > 0.1
153- **Output distribution shift** (prediction probability / regression output quantiles): flag if shift detected
154- **Latency drift** (average inference time): alert if >20% increase
155- **Failure rate monitoring** (predictions with NaN/Inf / out-of-range): flag if >1% failures
156
157**Logging output**: `run_models_drift_log.json` (append-only, timestamped entries)
158
159Example entry:
160```json
161{
162 "timestamp": "2026-04-05T14:32:00Z",
163 "model": "brain_gnn",
164 "inference_count": 100,
165 "input_kl_divergence": 0.045,
166 "output_mean_shift": 0.002,
167 "latency_ms": 45.2,
168 "failure_rate": 0.0,
169 "status": "healthy"
170}
171```
172
173**Alert thresholds**:
174- KL divergence > 0.1 → generate warning
175- Output shift > 5% std dev → investigation recommended
176- Latency drift > 20% → check computational resource bottleneck
177- Failure rate > 1% → stop inference, require manual review
178
179### Model Card Template (Minimum Required Metadata)
180Each model must include a model card in `skills/{model_name}/SKILL.md` documenting:
181
182```markdown
183## Model Card: {model_name}
184
185### Model Details
186- **Model name**: {name}
187- **Version**: {X.Y.Z}
188- **Date**: {YYYY-MM-DD}
189- **Source repository**: {repo_url}
190- **Paper**: {citation}
191
192### Intended Use
193- **Primary use case**: [e.g., fMRI-based phenotype classification]
194- **Input modalities**: [fMRI, sMRI, etc.]
195- **Supported tasks**: [classification, regression, interpretability]
196
197### Known Limitations
198- [e.g., "Trained on N subjects aged 18-65; generalization to pediatric/geriatric populations not validated"]
199- [e.g., "Sensitive to head motion artifacts; recommend ICA-FIX preprocessing"]
200
201### Validation Results
202- **Test set performance**: [accuracy/AUC/RMSE with confidence intervals]
203- **Cross-site validation**: [performance on held-out sites, if applicable]
204- **Robustness checks**: [drift detection history, adversarial perturbation results]
205
206### Dependencies & Versioning
207- **Required libraries**: [see {model_name}_spec.json]
208- **Hash (model weights)**: {SHA256}
209- **Last verified**: {date}
210```
211
212---
213
214## Delegation Rules
215
216### BrainGNN Route
217- Required modality preprocessing: `fmri-skill`
218- Typical upstream outputs expected: ROI matrices/time-series converted to model-required feature tensors
219
220### BNT Route
221- Required modality preprocessing: `fmri-skill`
222- Typical upstream outputs expected: Same ROI .pt files as BrainGNN (shared data source under `data/braingnn_input/`)
223
224### FM-APP Route
225- Required modality preprocessing: `fmri-skill` + `smri-skill`
226- Typical upstream outputs expected: fMRI ROI features plus structural MRI-derived features
227
228### NeuroStorm Route
229- Required modality preprocessing: follow the model doc in `skills/neurostorm/SKILL.md`
230- Typical upstream outputs expected: inputs and features specified by the NeuroStorm model card
231
232### GLM Route
233- Required modality preprocessing: `fmri-skill`
234- Concrete model/tool execution: `nilearn-tool`
235- Typical upstream outputs expected:
236 - first-level GLM: preprocessed task fMRI, events, optional confounds, named contrasts
237 - second-level GLM: subject-level contrast maps, group design matrix, group contrast definition
238
239### ICA Route
240- Required modality preprocessing: `fmri-skill`
241- Concrete model/tool execution: `nilearn-tool`
242- Typical upstream outputs expected:
243 - preprocessed resting-state fMRI image list
244 - optional mask and confounds
245 - requested component count
246
247### DictLearning Route
248- Required modality preprocessing: `fmri-skill`
249- Concrete model/tool execution: `nilearn-tool`
250- Typical upstream outputs expected:
251 - preprocessed resting-state fMRI image list
252 - optional mask and confounds
253 - requested component count
254
255### SVM Route
256- Required modality preprocessing: `fmri-skill` and/or `smri-skill`
257- Concrete model/tool execution: `nilearn-tool`
258- Typical upstream outputs expected:
259 - ROI/tabular feature matrix, diagnosis labels, optional covariates
260
261### SpaceNet Route
262- Required modality preprocessing: `fmri-skill` and/or `smri-skill`
263- Concrete model/tool execution: `nilearn-tool`
264- Typical upstream outputs expected:
265 - aligned subject image list, diagnosis labels, mask image, optional covariates
266
267### K-means Route
268- Required modality preprocessing: `fmri-skill` and/or `smri-skill`
269- Concrete model/tool execution: `nilearn-tool`
270- Typical upstream outputs expected:
271 - feature matrix or aligned image list for parcel discovery
272 - optional mask
273 - target parcel count
274
275### Hierarchical Route
276- Required modality preprocessing: `fmri-skill` and/or `smri-skill`
277- Concrete model/tool execution: `nilearn-tool`
278- Typical upstream outputs expected:
279 - feature matrix or aligned image list for parcel discovery
280 - optional mask or similarity structure
281 - target parcel count
282
283### Filtering Route
284- Required modality preprocessing: `fmri-skill`
285- Concrete model/tool execution: `nilearn-tool`
286- Typical upstream outputs expected:
287 - preprocessed BOLD image or extracted time series
288 - TR, optional confounds, optional mask
289 - optional frequency settings
290
291### Detrending Route
292- Required modality preprocessing: `fmri-skill`
293- Concrete model/tool execution: `nilearn-tool`
294- Typical upstream outputs expected:
295 - preprocessed BOLD image or extracted time series
296 - TR, optional confounds, optional mask
297 - detrending request and optional standardization settings
298
299### Shared Execution Routing
300- Environment/dependency planning: `dependency-planner` + `conda-env-manager`
301- Actual model run command execution: `claw-shell`
302
303---
304
305## Input and Output Contract (Entry-Level)
306
307### Inputs expected by this skill
308- Model selection (`brain_gnn`, `bnt`, `fm_app`, `neurostorm`, `glm`, `ica`, `dictlearning`, `svm`, `spacenet`, `kmeans`, `hierarchical`, `filtering`, or `detrending`)
309- Data split / subject list
310- Phenotype target definition
311- Optional compute constraints (GPU/CPU, memory, batch size)
312
313For GLM routes, the required task definition should be expressed as:
314- task name
315- events file
316- contrast(s) of interest
317- optional group-level analysis scope
318- whether the request is first-level GLM or second-level GLM
319- if second-level GLM: contrast map list and group design matrix
320
321For ICA routes, the required decomposition definition should be expressed as:
322- resting-state image list or subject list
323- number of components
324- optional mask and confounds
325
326For DictLearning routes, the required decomposition definition should be expressed as:
327- resting-state image list or subject list
328- number of components
329- optional mask and confounds
330
331For SVM routes, the required classification definition should be expressed as:
332- diagnosis target / label column
333- feature type (`roi/tabular`)
334- subject list or split definition
335- optional covariates
336
337For SpaceNet routes, the required classification definition should be expressed as:
338- diagnosis target / label column
339- feature type (`voxel-wise`)
340- subject list or split definition
341- optional covariates and mask
342
343For K-means routes, the required parcellation definition should be expressed as:
344- image list or feature matrix
345- target parcel / cluster count
346- optional mask
347
348For Hierarchical routes, the required parcellation definition should be expressed as:
349- image list or feature matrix
350- target parcel / cluster count
351- optional mask, similarity structure, or adjacency constraint
352
353For Filtering routes, the required denoising definition should be expressed as:
354- input BOLD image or time series
355- TR
356- optional confounds, mask, and frequency settings
357
358For Detrending routes, the required denoising definition should be expressed as:
359- input BOLD image or time series
360- TR
361- optional confounds, mask, and standardization settings
362
363### Outputs produced by this skill
364- A confirmed, numbered run plan
365- Pointers to the model-specific instruction file
366- Delegated preprocessing plan for required modalities
367- Structured output location recommendations
368
369---
370
371## Recommended Output Layout
372All model-running artifacts should be managed under `./run_models_output/`:
373- `run_models_output/preprocessed/`
374 - `fmri/` (from `fmri-skill`)
375 - `smri/` (from `smri-skill`, if required)
376- `run_models_output/brain_gnn/`
377- `run_models_output/bnt/`
378- `run_models_output/fm_app/`
379- `run_models_output/neurostorm/`
380- `run_models_output/glm/`
381- `run_models_output/ica/`
382- `run_models_output/dictlearning/`
383- `run_models_output/svm/`
384- `run_models_output/spacenet/`
385- `run_models_output/kmeans/`
386- `run_models_output/hierarchical/`
387- `run_models_output/filtering/`
388- `run_models_output/detrending/`
389- `run_models_output/logs/`
390- `run_models_output/reports/`
391
392---
393
394## Safety and Execution Policy
395- No execution before explicit user confirmation of the numbered plan.
396- All run/install actions must go through `claw-shell`.
397- If model skills are missing in `skills/<model-name>/`, stop and request or create them before execution.
398- Keep train/val/test split and target definition explicit to avoid leakage.
399
400---
401
402## When to Call This Skill
403- User asks to run BrainGNN or FM-APP.
404- User asks to run BNT (BrainNetworkTransformer).
405- User asks to run NeuroStorm.
406- User asks to run classical task activation analysis with GLM.
407- User asks to run group-level inference with second-level GLM.
408- User asks to perform resting-state network decomposition with ICA.
409- User asks to perform resting-state network decomposition with DictLearning.
410- User asks to perform disease classification with SVM.
411- User asks to perform disease classification with SpaceNet.
412- User asks to perform brain parcellation with K-means.
413- User asks to perform brain parcellation with Hierarchical clustering.
414- User asks to perform signal denoising with filtering.
415- User asks to perform signal denoising with detrending.
416- User asks which phenotype model to use for fMRI/sMRI ROI data.
417- User asks for a unified entry point to model introduction + run routing.
418
419---
420
421## Complementary / Related Skills
422- `fmri-skill`
423- `smri-skill`
424- `dependency-planner`
425- `conda-env-manager`
426- `claw-shell`
427
428---
429
430## Reference
431- BrainGNN paper and code:
432 - https://github.com/xxlya/BrainGNN_Pytorch/tree/main
433- FM-APP paper and code:
434 - https://github.com/ZhibinHe/FM-APP
435- Nilearn GLM documentation:
436 - https://nilearn.github.io/stable/glm/index.html
437
438Created At: 2026-03-28 20:38 HKT
439Last Updated At: 2026-04-14 00:28 HKT
440Author: chengwang96