Data & Analytics
Data agent skills make AI agents useful for data work: writing SQL, cleaning datasets, building pipelines, working with spreadsheets, and producing analyses. Each skill is a reviewed SKILL.md file that teaches the agent one workflow well, ready to install in seconds.
-
holobiomicslab Skill Ms Data Format Parsing And ConversionUse when you have raw breath HRMS data in mzML or mzXML format and need to extract volatile organic compound (VOC) features as a standardized CSV table indexed by m/z value, with columns for scan time or sample identifiers and corresponding intensity measurements.
-
holobiomicslab Skill Pandas Dataframe Column SpecificationUse when when you have mass spectrometry data in a Pandas DataFrame with column names that do not match pyOpenMS-viz's default expectations (e.g., 'm/z' vs 'mz' or 'retention_time' vs 'rt'), or when your data uses domain-specific column labels (e.g., 'mass_to_charge', 'scan_time', 'peak_intensity').
-
holobiomicslab Skill S7 Object Construction And ValidationUse when after successfully parsing vendor-specific metabolomic data files (Metabolon Excel, Nightingale, Olink, SomaLogic) into separate data, samples, and features tables, and before applying quality control or batch normalization pipelines.
-
holobiomicslab Skill Targeted Compound Metadata FormattingUse when you have a raw list of target compounds (in .xlsx, CSV, or database form) with heterogeneous column names and layouts, and you need to prepare it for targeted peak detection, EIC extraction, or quality metric calculation in TARDIS or similar LC–MS metabolomics tools.
-
holobiomicslab Skill Database Schema Design For Spectral DataUse when you have MS/MS spectral library data currently stored in file-based formats (JSON, CSV, binary, MGF, MSP) and need to migrate to database-backed storage to support fast queries by metadata filters (precursor m/z, ion mode, retention time) and computed similarity scores against query.
-
holobiomicslab Skill Lc Ms Feature Grouping By Retention TimeUse when immediately after chromatographic peak detection (findChromPeaks) when you have detected peaks across multiple samples and need to identify which peaks represent the same feature across the sample cohort.
-
holobiomicslab Skill Library Spectrum Database Format ParsingUse when you have experimental MS/MS spectra (from mzML, mzXML, or raw instrument formats) and wish to match them against a reference spectral library provided in MSP or CSV format. The skill is required as the first step before similarity scoring and candidate ranking.
-
holobiomicslab Skill Qc Replicate Identification And GroupingUse when you have a QC-annotated LC-MS feature table (CSV or data frame format with sample metadata) and need to isolate QC replicate measurements prior to computing quality metrics such as D-Ratio or performing signal drift correction.
-
holobiomicslab Skill Linear Score Calculation From CoefficientsUse when when you have 1H-NMR metabolite measurements from Nightingale Health assayed on a new cohort and wish to compute risk scores (e.g., all-cause mortality, cardiovascular event, type 2 diabetes) using published metabolic biomarker weights from a reference study.
-
holobiomicslab Skill Pre Analytical Delay Effect QuantificationUse when you have uploaded a pre-analytical data table containing sample metadata, processing delay annotations (pre- and post-centrifugation timestamps or duration), and paired NMR metabolomic measurements for a sample cohort, and you need to quantify how delays at different time-points affect.
-
holobiomicslab Skill Published Metabolic Profile ImplementationUse when you have Nightingale Health 1H-NMR metabolomics measurements for a new cohort and wish to compute one or more established metabolic risk scores (mortality, MetaboAge, cardiovascular event, type-2 diabetes, COVID-19 severity) without recalibration.
-
holobiomicslab Skill Spectral Quality Filtering Signal To NoiseUse when you have raw spectroscopic datasets (NMR, HSQC, COSY, IR) in standardized array or DataFrame format and need to curate them for multimodal transformer training.
-
holobiomicslab Skill CSV Data Aggregation And DeduplicationUse when when you have downloaded a multi-file .csv library repository (e.g., LipidMatch) and need to verify that it meets minimum thresholds for species diversity (e.g., 500,000+ distinct lipid species) and category breadth (e.g., 60+ lipid-type categories).
-
holobiomicslab Skill Metadata Field Extraction From HeadersUse when you have tabular data (CSV or Excel) with column headers annotated using MESSES tagging syntax (#<table_name>.id for record identifiers, #.
-
holobiomicslab Skill Chemical Structure To Taste PredictionUse when you have a CSV or EXCEL file containing molecular descriptors (pre-computed structural features) for one or more chemical compounds, and you need binary bitterness predictions (bitter vs. non-bitter) for each molecule.
-
holobiomicslab Skill Foundation Model Prediction GenerationUse when you have a pre-trained foundation model checkpoint (e.g., NaFM.ckpt) and new molecules represented as SMILES strings or a CSV file, and you need to generate predictions (e.g., bioactivity scores, taxonomy class, screening rankings) or embeddings for downstream analysis.
-
holobiomicslab Skill Gene Ontology Enrichment VisualizationUse when you have completed a GO enrichment analysis (e.g., via hypergeometric test or similar) and need to visualize the results as a dotplot. Input is a CSV table with columns for GO term identifiers, p-values or adjusted p-values, gene ratios (or counts), and gene count;
-
holobiomicslab Skill Variance Estimation Within And Between GroupsUse when you have a QC-annotated LC-MS feature intensity table (CSV or data frame) with replicate QC samples and biological samples from multiple batches or run orders, and you need to assess which features maintain consistent signal intensity across technical replicates (within-group) relative to.
-
holobiomicslab Skill Binary Classification Output InterpretationUse when you have executed a binary classifier (such as BitterPredict.m) on a set of molecules with chemical structure descriptors and need to translate the raw predictions into a structured CSV output file that maps molecule identifiers to their predicted class labels (bitter or not-bitter).
-
holobiomicslab Skill Dataframe Construction From Backend SourcesUse when when implementing a new MsBackend subclass that stores only a subset of core spectra variables (e.
-
holobiomicslab Skill Mass Spectrometry Data Visualization PandasUse when when you have mass spectrometry data (mzML, Bruker .d, or CSV) loaded into a Pandas DataFrame with columns for m/z, retention time, ion mobility, or intensity values, and you need to render spectrum plots, chromatograms, mobilograms, or 2D peak maps.
-
holobiomicslab Skill CSV Table Validation And Integrity CheckingUse when you have received a raw peak table CSV (in either standardized format or output from metabolomic software tools like XCMS, MZmine, etc.) and a corresponding label file, and you need to confirm both files are well-formed and mutually consistent before passing them to NOREVA preprocessing.
-
holobiomicslab Skill Enrichment Metric Scaling And NormalizationUse when you have parsed GO enrichment results (CSV with GO term identifiers, p-values or adjusted p-values, gene ratios, and gene counts) and need to prepare these metrics for dotplot visualization.
-
holobiomicslab Skill Mass Spectrometry Quantification ExtractionUse when you have raw LipidSearch or LIQUID output files (CSV or TSV format) containing lipid identifiers and per-sample quantification measurements, and you need to convert them into a machine-readable data matrix for downstream statistical analysis, normalization, or differential abundance.
-
holobiomicslab Skill Molecular Identifier Completeness VerificationUse when during MSP, MGF, JSON, or CSV file parsing when standardizing mass spectra from heterogeneous open mass spectral libraries (OMSLs).
-
holobiomicslab Skill Multi Assay Data Integration And HarmonizationUse when you have independent LC-MS assays (e.g., positive and negative ionization modes, different lipid profiling assays, or different chromatographic methods) analyzed on the same sample cohort and want to integrate them into a single discriminant or regression model without losing assay-level.
-
holobiomicslab Skill Cohort Stratified Metabolic Performance AnalysisUse when when you have uploaded a pre-analytical data table containing sample metadata, processing delay annotations (pre- and post-centrifugation times), and paired NMR metabolomic measurements for a plasma or serum cohort, and you need to determine how processing delays impact metabolite.
-
holobiomicslab Skill Python Deep Learning Model Loading And ExecutionUse when when you have a pre-trained deep learning model checkpoint (saved in PyTorch format) and new 1D 1H NMR spectral data in CSV and peak-list TXT formats, and you need to generate peak-to-metabolite assignments or other structured outputs from that model without modification of model weights.
-
holobiomicslab Skill Experiment Design Specification For ClusteringUse when you have a CSV feature table from XCMS or other MS feature detection tools and are about to run RAMClustR clustering, but need to encode which samples are QC replicates, which batch they belong to, and their run order.
-
holobiomicslab Skill Statistical Significance Threshold ApplicationUse when you have generated omu_summary or anova_function output (a dataframe with padj values for all tested metabolites) and need to subset compounds for class-specific frequency counting, fold-change analysis, or visualization.
-
holobiomicslab Skill Missing Value Imputation In Quantification TablesUse when after feature alignment across multiple LC-MS/MS runs, when the unified feature list contains zeros or nulls for specific feature–sample pairs because peaks were not detected in those individual runs, but the feature was detected in other samples in the cohort.
-
holobiomicslab Skill Reference Based Vs Global Coordinate RegistrationUse when you have detected feature tables from multiple LC-IMS-MS/MS samples (each with mz, drift_time, retention_time, and intensity columns) and need to match corresponding features across datasets to enable cross-sample quantitation or cohort analysis.
-
holobiomicslab Skill Mass Spectrometry Data Structure InterpretationUse when you have received an mzPeak archive (a ZIP file containing Parquet tables) and need to understand its internal structure, validate that spectrum metadata aligns with signal data, reconstruct m/z and intensity arrays (especially when null marking or zero-run stripping is present), or verify.
-
holobiomicslab Skill Python Binary Io And SerializationUse when when you have mzPeak files (Parquet-based archives in uncompressed ZIP containers) or other PyArrow-compatible columnar formats containing mass spectrometry spectra, and you need to extract and decode spectral data arrays (m/z values, intensities) into Python memory for downstream.
-
holobiomicslab Skill Mass Spectrometry Peak ClassificationUse when you have raw mzML files and feature tables (CSV format from mzMine or XCMS) from untargeted LCMS experiments and need to distinguish true metabolite peaks from false positives introduced by the peak-picking algorithm.
-
holobiomicslab Skill Metabolomic Data Structure FormattingUse when after peak detection in MZmine2 has produced an MGF file (containing MS1 and MS2 spectra) and a feature abundance table (CSV or BIOM), but before running q2-qemistree tree construction or any QIIME 2-based metabolomic analysis.
Frequently asked questions
What are Data & Analytics agent skills?
Data agent skills make AI agents useful for data work: writing SQL, cleaning datasets, building pipelines, working with spreadsheets, and producing analyses. Each skill is a reviewed SKILL.md file that teaches the agent one workflow well, ready to install in seconds.
Which Data & Analytics skills are most installed?
Popular Data & Analytics skills on SkillMD right now include ms-data-format-parsing-and-conversion, pandas-dataframe-column-specification, s7-object-construction-and-validation. Rankings shift as installs change; sort this page by "Most installs" for the live list.
Do Data & Analytics skills work with Claude Code and Cursor?
Yes. Every skill here ships as a SKILL.md file, an open format that works in Claude Code, Claude.ai, Cursor, Codex, Windsurf, and 60+ other agents. Install one with npx skillmds@latest add <owner>/<name>, or copy the file into your agent's skills directory.