BioResearch Agent — Disease Case Study Skill
Capability
A workflow-composition + validation skill. It runs a real, public GEO dataset through the framework's analysis engine (differential expression → pathway enrichment → candidate ranking, with optional literature and causal-inference stages) and produces a blind benchmark evaluation:
- Known-gene recovery — how many established disease genes surface in the top-ranked candidates (e.g. SNCA / LRRK2 / PARK7 / PINK1 / PRKN / GBA for Parkinson's).
- Pathway sanity — whether disease-relevant pathways (dopamine / mitochondrial / oxidative / immune) are enriched.
- Reproducibility — commit hash, environment, and data
sha256recorded in an Evidence Package.
Case Study 1 ships with the suite: Parkinson's disease / GSE7621 (GPL570, 25 samples).
See bio-research-os/eval/case_study_pd.py and the Biomedical Workflow Validation Suite README.
Run
python bio-research-os/eval/case_study_pd.py \
--matrix-path downloads/gse7621_matrix.txt.gz \
--annot-path downloads/gpl570.annot.gz \
--output-dir docs/case-study
(If paths are omitted the runner downloads the real GSE7621 matrix + GPL570 annotation itself.)
Outputs (in docs/case-study/)
GSE7621_deg.csv— differential expression resultsGSE7621_top_candidates.csv— ranked biomarker candidatesGSE7621_pathway_enrichment.csv— enriched KEGG / GO termsGSE7621_volcano.png— volcano plotGSE7621_report.txt— structured report incl. blind benchmarkGSE7621_evidence_package.json— provenance + benchmark + limitations
Note
This skill is a composition & validation interface. It adds no statistics of its own — every number comes from the framework's workflow modules. It is the honest way to demonstrate the framework works on real data (not synthetic injection). Requires no LLM key.