Data & Analytics
Data agent skills make AI agents useful for data work: writing SQL, cleaning datasets, building pipelines, working with spreadsheets, and producing analyses. Each skill is a reviewed SKILL.md file that teaches the agent one workflow well, ready to install in seconds.
-
holobiomicslab Skill Metabolomics Parameter ExtractionUse when you have raw untargeted metabolomics data in mzML, mzXML, or CDF format from qTOF, Orbitrap, or FTICR mass analyzers, at least 3 samples, a sample metadata spreadsheet linking filenames to experimental factors, and need to generate optimized processing parameters for XCMS or MZmine2.
-
holobiomicslab Skill Polarity Based Compound FilteringUse when you have a multi-polarity compound target list (e.g., a .xlsx file with a polarity or ionization mode column indicating positive or negative ESI mode) and you are about to perform targeted peak detection in a single LC–MS acquisition mode (e.g., positive-ion mode only).
-
holobiomicslab Skill Regulatory Relationship InferenceUse when you have matched multiomics measurements (CNV, mutations, DNA methylation, histone PTMs, coding/noncoding transcripts, miRNA, lncRNA, proteomics, phosphoproteomics) across a cohort of cell lines or samples and want to identify which upstream omics features predict metabolite abundance.
-
holobiomicslab Skill Annotaterc Function ParameterizationUse when you have LC–MS all-ion fragmentation chromatograms already processed by xcms and clustered by RamClustR, a feature table (targetTable.csv format) listing features to annotate, and you need rank-1 metabolite or lipid identifications with confidence metrics.
-
holobiomicslab Skill Feature Consolidation Across BatchesUse when you have two or more CSV feature tables from separate metabolomic experiments (each containing mass, retention time, intensity, isotope, and adduct columns), and you need to align and merge them into a single feature-by-sample matrix where features from different experiments are matched if.
-
holobiomicslab Skill Isotope Adduct Anchor IdentificationUse when when you have extracted mass tracks (EICs) from individual LC-MS samples and need to establish reliable landmarks for subsequent pairwise or global alignment across a cohort.
-
holobiomicslab Skill Cross Software Feature MatchingUse when you have collected feature lists in CSV format from two or more MS acquisition methods (e.g., LC-MS vs. LC-IMS-MS), processing software packages (e.g., vendor-specific vs. open-source), or instrument platforms (e.
Audited -
holobiomicslab Skill Gap Filling Algorithm SelectionUse when when processing untargeted LC-MS data with SLAW and observing incomplete feature detection across the sample cohort—i.e., features present in some samples but with missing values (zeros or NAs) in others due to signal dropout, retention time drift, or mass calibration drift.
Audited -
holobiomicslab Skill Lc Ms Data Structure ValidationUse when before launching TARDIS peak detection on a new LC–MS dataset or target compound list. Apply this skill when you have raw MS data files in vendor formats (e.g., .raw, .d) and/or a spreadsheet-based target list (.xlsx or .
Audited -
holobiomicslab Skill Peakmap Heatmap Rendering Mz RtUse when when you have mass spectrometry data organized in a Pandas DataFrame with m/z values, retention time (RT), and intensity measurements, and you want to visualize the joint distribution and correlation of these three dimensions to identify peaks, assess separation, and detect patterns across.
Audited -
holobiomicslab Skill Python Sqlite Migration And EtlUse when you have MS/MS spectral library data currently stored in multiple file formats (JSON, CSV, or binary) and need to enable fast, filtered queries by metadata (e.g., precursor m/z, retention time, molecular class) without loading entire libraries into memory.
Audited -
holobiomicslab Skill Metabolic Parameter VisualizationUse when when you have paired NMR metabolite measurements and corresponding processing metadata (pre-centrifugation delay, post-centrifugation delay, sample type, cohort) for a blood sample cohort and need to determine which metabolites remain stable across the expected or observed delay range, or.
Audited -
holobiomicslab Skill Lexical Analysis TokenizationUse when you have a mass-spectrometry query string written in MassQL (or similar domain-specific SQL-inspired syntax) that must be converted into structured form for execution. The input is raw, unparsed text containing SQL keywords, MS-specific operators (e.
Audited -
holobiomicslab Skill Bubble Plot Aesthetic MappingUse when you have parsed GO enrichment results (GO term identifiers, p-values or adjusted p-values, gene ratios, gene counts) into CSV format and need to render them as interactive bubble plots.
Audited -
holobiomicslab Skill Chemical Structure ValidationUse when after compound database dereplication with SIRIUS or MetFrag has produced candidate annotations (CSV or JSON format), and you need to filter implausible structures, compute standardized molecular descriptors, and rank candidates by confidence before reporting final metabolite.
Audited -
holobiomicslab Skill Database Repository RetrievalUse when when you need to obtain a specific curated database (e.g., DNA adduct compounds) that is published in a GitLab or GitHub repository and available in structured formats (SDF, Excel, Word).
Audited -
holobiomicslab Skill JSON Spectral Data ProcessingUse when you have raw or semi-curated mass spectrometry spectral data in JSON, CSV, MSP, or MGF format from multiple open mass spectra libraries (OMSLs) and need to standardize field names, validate chemical identifiers (SMILES, InChI, InChIKey), remove duplicates, filter by quality criteria.
Audited -
holobiomicslab Skill Metabolon Excel Format HandlingUse when you have raw Metabolon Excel workbooks (metabolon_v1.1_example.xlsx or metabolon_v1.2_example.
-
holobiomicslab Skill Numeric Variable Range AnalysisUse when you have loaded a numeric column (e.g., H/C ratio, O/C ratio, m/z value, or intensity) from a CSV file into Punc'data and need to render a histogram with appropriate bar spacing. The skill is triggered when the range of the column is small enough that default bin widths (1.
-
holobiomicslab Skill Technical Heterogeneity RemovalUse when your metabolomics matrix (log2-scaled, samples × features in .csv format) exhibits spatial or distributional separation by batch/analytical run, visible in PCA or RLA plots.
-
holobiomicslab Skill File Format Parsing And ValidationUse when you have peak/feature tables from one or more of MZmine, XCMS, MS-DIAL, or Compound Discoverer and need to ingest them into LipidMatch for lipid identification. The input files are in tabular format (CSV, TSV, or Excel) and their upstream tool origin may be unknown or mixed.
-
holobiomicslab Skill Mass Grid Construction And MappingUse when after mass track extraction from individual LC-MS samples, when you need to align mass tracks across a cohort to produce a unified feature matrix. Specifically: when study size is ≤10 samples, use pairwise anchor-prioritized alignment;
-
holobiomicslab Skill Msexperiment Backend ConfigurationUse when you have multiple centroided .mzML LC-MS files that need to be loaded into a unified object for targeted peak integration, and you need to distinguish QC runs from sample runs to compute per-group quality metrics (e.g., average SNR, peak correlation, area under curve per QC cohort).
-
holobiomicslab Skill Multiformat Data Export To PDF CSVUse when you have raw MS data in vendor formats (Agilent .d, Thermo .raw, Bruker .
-
holobiomicslab Skill Neural Network Architecture DesignUse when you have raw mzML files and feature tables (CSV from mzMine or XCMS) for LCMS data, have generated training/validation/test batches with known class imbalance, and need to train a CNN model from scratch to achieve AUC ROC > 0.9 for distinguishing true from false positive MS1 peaks.
-
holobiomicslab Skill Peak Property Preparation From CSVUse when you have a CSV file containing nucleoside or peptide molecular data (formulas, identifiers, retention times, intensities) that you want to simulate as LC-MS/MS runs. Use this skill as the mandatory first step before selecting a fragmentation model and noise injector in SMITER.
-
holobiomicslab Skill Quality Control Metric ComputationUse when after feature integration and imputation when you have QC-annotated LC-MS feature intensity data (CSV or data frame format) with replicate QC samples.
-
holobiomicslab Skill Peak Detection And Mass AlignmentUse when when you have raw LC-MS/MS data files (.mzML, .raw, or vendor formats) from multiple samples and need to identify reproducible molecular features across the cohort before annotation or statistical analysis.
-
holobiomicslab Skill Query Result Serialization To CSVUse when after executing a MassQL query against mzML mass spectrometry files and obtaining a tabulated result DataFrame in memory.
-
holobiomicslab Skill Interactive Data Exploration DesignUse when you have NMR metabolomics measurements paired with pre-analytical metadata (processing delay times, centrifugation timing, sample type such as plasma vs. serum, cohort identifiers) and need to interactively explore how variation in processing conditions drives changes in metabolic.
-
holobiomicslab Skill Nightingale 1h Nmr Data IntegrationUse when you have newly assayed 1H-NMR metabolomics data from Nightingale Health (CSV or TSV format) and need to apply one or more published metabolic risk scores (Deelen et al. all-cause mortality, van den Akker MetaboAge, Würtz cardiovascular event risk, etc.).
-
holobiomicslab Skill Pre Analytical Delay StratificationUse when when you have NMR metabolite measurements paired with documented pre-centrifugation and post-centrifugation delay times, and need to assess how processing delays affect metabolic parameter stability within a plasma or serum sample cohort.
-
holobiomicslab Skill Classifier Prediction InferenceUse when you have a CSV or Excel file containing chemical structure descriptors for one or more molecules, and you want to obtain binary bitter/not-bitter predictions for each molecule using the BitterPredict classifier.
-
holobiomicslab Skill Mass Spectrometry Feature ClusteringUse when after XCMS feature detection and alignment when you have a CSV-formatted feature table with m/z and retention time annotations and want to deduplicate isotopic peaks, adducts, and in-source fragments into compound-level clusters before molecular weight inference or spectral matching.
-
holobiomicslab Skill Massgrid Construction And ValidationUse when after individual mass tracks (EICs) have been extracted from each sample's mzML file and you need to create a unified, cross-sample m/z reference structure. Triggered when: (1) you have ≥2 samples in a cohort; (2) mass tracks have been binned at 0.
-
holobiomicslab Skill Multi Dataset Integration Mzrt SpaceUse when you have multiple CSV feature tables from independent metabolomic experiments, each with RT and m/z annotations, and you need to produce a single consolidated feature matrix for comparative analysis across all samples.
Frequently asked questions
What are Data & Analytics agent skills?
Data agent skills make AI agents useful for data work: writing SQL, cleaning datasets, building pipelines, working with spreadsheets, and producing analyses. Each skill is a reviewed SKILL.md file that teaches the agent one workflow well, ready to install in seconds.
Which Data & Analytics skills are most installed?
Popular Data & Analytics skills on SkillMD right now include metabolomics-parameter-extraction, polarity-based-compound-filtering, regulatory-relationship-inference. Rankings shift as installs change; sort this page by "Most installs" for the live list.
Do Data & Analytics skills work with Claude Code and Cursor?
Yes. Every skill here ships as a SKILL.md file, an open format that works in Claude Code, Claude.ai, Cursor, Codex, Windsurf, and 60+ other agents. Install one with npx skillmds@latest add <owner>/<name>, or copy the file into your agent's skills directory.