Aggregate results, train ML models, and produce reports with validated references.
Instructions
Join outputs in DuckDB v1.1+ and build feature tables. Arrow / DuckLake integration is the recommended bridge into ML pipelines for large datasets.
Train baseline models and evaluate with cross-validation.
CPU baseline: scikit-learn v1.5+ for linear/tree/clustering baselines; XGBoost v2.1.4+ for gradient boosting.
GPU node available (CUDA): set device="cuda" on XGBoost (native since v2.0) by default. For sklearn-compatible estimators (random forest, k-means, PCA, UMAP), use RAPIDS cuML as a drop-in replacement and record the device in the run log.
Generate reports and validate references.
For exploratory omics projects, aggregate discovery evidence across the literature-derived analysis playbook, annotation, phylogenomics, viromics, and comparative-genomics outputs.
Comparative-axes rollup — join the per-axis comparison artifacts produced by upstream skills into a single comparative_axes_summary.tsv. The rollup must have one row per (query genome, axis) and include:
genome-property frontier (size, gene count, etc. — link to relative_genome_metrics.tsv and genome_size_frontier.tsv)
marker-gene census (link to marker_census.tsv)
family copy-number expansions/contractions (link to family_copy_number_comparison.tsv and family_expansion_candidates.tsv)
synteny / conserved neighborhoods (link to conserved_neighborhoods.tsv)
non-coding RNA census (link to ncRNA_census.tsv)
Each row records observation, comparison baseline, literature reference, status (notable / conserved / artifact / negative), and a follow-up test.
Produce an interesting-findings section that ranks candidate discoveries relative to the literature-derived baseline and separates:
strong candidates with multiple evidence types
plausible candidates needing validation
likely artifacts or conserved lineage features
explicit negative findings where nothing notable was detected
Include the comparison baseline, literature context, confidence, and next discriminating analyses for each candidate.
Quick Reference
Task
Action
Run workflow
Follow the steps in this skill and capture outputs.
Validate inputs
Confirm required inputs and reference data exist.
Review outputs
Inspect reports and QC gates before proceeding.
Tool docs
See docs/README.md.
Input Requirements
Prerequisites:
Tools available in the active environment (Pixi/conda/system). See docs/README.md for expected tools.
Results tables and metadata are available.
Inputs:
On failure: retry with alternative parameters; if still failing, record in report and exit non-zero.
Verify input tables are readable and schema-consistent.
Discovery summary joins candidate genes/features to annotation evidence, comparison baseline, literature context, and confidence.
comparative_axes_summary.tsv covers all five mandatory axes (genome-property frontier, marker-gene census, family copy-number, synteny/neighborhoods, ncRNA census) for every query genome, with rows for axes that produced negative findings.
Final report states what is interesting, what is conserved/expected, what is likely artifact, and what should be tested next.
Examples
Example 1: Expected input layout
results/*.parquet or results/*.tsv
metadata.tsv
Troubleshooting
Issue: Missing inputs or reference databases
Solution: Verify paths and permissions before running the workflow.
Issue: Low-quality results or failed QC gates
Solution: Review reports, adjust parameters, and re-run the affected step.
1---2name: bio-stats-ml-reporting3description: Aggregate results, train ML models, and produce reports with validated references.4---56# Bio Stats ML Reporting
78Aggregate results, train ML models, and produce reports with validated references.
910## Instructions
11121. Join outputs in DuckDB v1.1+ and build feature tables. Arrow / DuckLake integration is the recommended bridge into ML pipelines for large datasets.
132. Train baseline models and evaluate with cross-validation.
14 - CPU baseline: scikit-learn v1.5+ for linear/tree/clustering baselines; XGBoost v2.1.4+ for gradient boosting.
15 - GPU node available (CUDA): set `device="cuda"` on XGBoost (native since v2.0) by default. For sklearn-compatible estimators (random forest, k-means, PCA, UMAP), use **RAPIDS cuML** as a drop-in replacement and record the device in the run log.
163. Generate reports and validate references.
174. For exploratory omics projects, aggregate discovery evidence across the literature-derived analysis playbook, annotation, phylogenomics, viromics, and comparative-genomics outputs.
185. **Comparative-axes rollup** — join the per-axis comparison artifacts produced by upstream skills into a single `comparative_axes_summary.tsv`. The rollup must have one row per (query genome, axis) and include:
19 - `genome-property frontier` (size, gene count, etc. — link to `relative_genome_metrics.tsv` and `genome_size_frontier.tsv`)
20 - `marker-gene census` (link to `marker_census.tsv`)
21 - `family copy-number expansions/contractions` (link to `family_copy_number_comparison.tsv` and `family_expansion_candidates.tsv`)
22 - `synteny / conserved neighborhoods` (link to `conserved_neighborhoods.tsv`)
23 - `non-coding RNA census` (link to `ncRNA_census.tsv`)
24 Each row records observation, comparison baseline, literature reference, status (notable / conserved / artifact / negative), and a follow-up test.
256. Produce an interesting-findings section that ranks candidate discoveries relative to the literature-derived baseline and separates:
26 - strong candidates with multiple evidence types
27 - plausible candidates needing validation
28 - likely artifacts or conserved lineage features
29 - explicit negative findings where nothing notable was detected
307. Include the comparison baseline, literature context, confidence, and next discriminating analyses for each candidate.
3132## Quick Reference
3334| Task | Action |
35|------|--------|
36| Run workflow | Follow the steps in this skill and capture outputs. |
37| Validate inputs | Confirm required inputs and reference data exist. |
38| Review outputs | Inspect reports and QC gates before proceeding. |
39| Tool docs | See `docs/README.md`. |
4041## Input Requirements
4243Prerequisites:
44- Tools available in the active environment (Pixi/conda/system). See `docs/README.md` for expected tools.
45- Results tables and metadata are available.
46Inputs:
47- results/*.parquet or results/*.tsv
48- metadata.tsv
4950## Output
5152- results/bio-stats-ml-reporting/models/
53- results/bio-stats-ml-reporting/metrics.tsv
54- results/bio-stats-ml-reporting/comparative_axes_summary.tsv
55- results/bio-stats-ml-reporting/discovery_summary.tsv
56- results/bio-stats-ml-reporting/report.md
57- results/bio-stats-ml-reporting/logs/
5859## Quality Gates
6061- [ ] Model performance sanity checks pass.
62- [ ] Reference validation passes.
63- [ ] On failure: retry with alternative parameters; if still failing, record in report and exit non-zero.
64- [ ] Verify input tables are readable and schema-consistent.
65- [ ] Discovery summary joins candidate genes/features to annotation evidence, comparison baseline, literature context, and confidence.
66- [ ] `comparative_axes_summary.tsv` covers all five mandatory axes (genome-property frontier, marker-gene census, family copy-number, synteny/neighborhoods, ncRNA census) for every query genome, with rows for axes that produced negative findings.
67- [ ] Final report states what is interesting, what is conserved/expected, what is likely artifact, and what should be tested next.
6869## Examples
7071### Example 1: Expected input layout
7273```text
74results/*.parquet or results/*.tsv
75metadata.tsv
76```
7778## Troubleshooting
7980**Issue**: Missing inputs or reference databases
81**Solution**: Verify paths and permissions before running the workflow.
8283**Issue**: Low-quality results or failed QC gates
84**Solution**: Review reports, adjust parameters, and re-run the affected step.
Run npx skillmds@latest add gabrielmoreira/bio-stats-ml-reporting in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.
Aggregate results, train ML models, and produce reports with validated references. It is listed under AI & ML on SkillMD.
This skill has not completed SkillMD's automated safety review yet. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: docs only. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.
This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.
Yes. Installing skills from SkillMD is free, and the skill stays under its author's original license.
gabrielmoreira (@gabrielmoreira) published this skill. Their other Agent Skills are listed on their SkillMD profile.