Data & Analytics
Data agent skills make AI agents useful for data work: writing SQL, cleaning datasets, building pipelines, working with spreadsheets, and producing analyses. Each skill is a reviewed SKILL.md file that teaches the agent one workflow well, ready to install in seconds.
-
holobiomicslab Skill Feature Count Verification Across Adducts And IsotopologuesUse when after mzRAPP has exported a benchmark CSV file from centroided mzML files and you need to confirm the benchmark was constructed correctly before using it to evaluate NPP tool performance. Specifically, when you have a reference expectation (e.
-
holobiomicslab Skill Thermodynamic Molecular Index Calculation From Elemental ComUse when when you have peak-abundance .csv files with assigned molecular formulas (elemental composition: C, H, O, N, P, S) from FT-ICR MS or high-resolution MS and need to characterize the redox and structural properties of the molecular pool—e.
-
diegosouzapw Bundle PolarsPolars workflow skill. Use this skill when the user needs Fast in-memory DataFrame library for datasets that fit in RAM. Use when pandas is too slow but data still fits in memory. Lazy evaluation, parallel execution, Apache Arrow backend. Best for 1-100GB datasets, ETL pipelines, faster pandas replacement. For larger-than-RAM data use dask or vaex and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
54 -
diegosouzapw Bundle Calc V2LibreOffice Calc workflow skill. Use this skill when the user needs Spreadsheet creation, format conversion (ODS/XLSX/CSV), formulas, data automation with LibreOffice Calc and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
54 -
holobiomicslab Skill SQL Query Optimization For Similarity SearchUse when when migrating spectral library data from file-based formats (JSON, CSV, binary) into a persistent store and need to support fast filtered queries on metadata and similarity computations against query spectra.
-
holobiomicslab Skill Confounder Adjustment Epidemiological AnalysisUse when when testing associations between metabolic features (from NMR or MS) and a phenotype of interest (e.g., BMI, disease status) in a cohort where age, gender, or clinical confounders are known to correlate with both the metabolite and phenotype.
-
holobiomicslab Skill Metabolite Stability Assessment Across CohortsUse when you have uploaded a pre-analytical data table containing sample metadata, processing timestamps (pre- and post-centrifugation), and NMR metabolomic measurements for a cohort of peripheral blood samples (plasma/serum), and you need to quantify the magnitude and direction of metabolite.
-
holobiomicslab Skill Mass Spectrometry Data Parsing Mzml BrukerUse when you have raw mass spectrometry data in mzML or Bruker .d format and need to ingest it into a tabular format (pandas DataFrame) for visualization, statistical analysis, or integration with other Python-based mass spectrometry tools.
-
holobiomicslab Skill Mass Spectrometry Plot Type SpecializationUse when you have a Pandas DataFrame containing mass spectrometry data (e.
-
holobiomicslab Skill Metabolite Set File Parsing And ValidationUse when when a user has prepared a custom collection of metabolite sets (e.g., from spectral fragmentation clustering, literature curation, or domain-specific grouping) in CSV or JSON format and wants to score their activity levels using PALS without modifying the core PALS codebase.
-
holobiomicslab Skill Data Lineage Preservation In Etl PipelinesUse when when consolidating entries from multiple heterogeneous source databases into a unified table, and you need to maintain auditable connections between final curated records and their original source entries.
-
holobiomicslab Skill Metabolomic Workflow Ranking VisualizationUse when after running NOREVA's multi-class or time-course assessment functions (normulticlassqcall, normulticlassnoall, normulticlassisall, nortimecourseqcall, or nortimecoursenoall) that produce an overall ranking CSV file of preprocessed workflows.
-
holobiomicslab Skill Raw Spectral Data Import And PreprocessingUse when you have raw metabolomics data in mzML or mzXML format and need to convert it into a normalized feature table (CSV or mzTab) via automated batch processing.
-
holobiomicslab Skill Retention Order Prediction Model ExecutionUse when you have access to the aalto-ics-kepaco/retention_order_prediction repository, have installed all Python (scipy, numpy, sklearn, joblib, pandas, networkx) and R dependencies, possess molecular feature data (MACCS fingerprints or equivalent), and need to execute a specific evaluation.
-
holobiomicslab Skill Spectrum Chromatogram Mobilogram RenderingUse when you have mass spectrometry data in a Pandas DataFrame with columns for m/z and intensity (spectrum), retention time and intensity (chromatogram), or drift time and intensity (mobilogram), and need to render 1D traces as static or interactive plots for exploratory analysis, quality control.
-
holobiomicslab Skill Metabolite Quantification Accuracy AssessmentUse when you have executed mzExacto() on a preprocessed GC-MS dataset and need to verify that the returned dataframe correctly matches query chemicals to their m/z peaks, retention times, and quantitative measurements (area values).
-
holobiomicslab Skill Feature Table Annotation With Sample MetadataUse when your input is a feature intensity table (CSV or R data frame) with features as columns and samples as rows, and you have accompanying sample metadata (batch identifiers, QC/study sample labels, run order, sample phenotypes, collection dates).
-
holobiomicslab Skill Feature Table Generation From Aligned SpectraUse when after retention-time and m/z-based peak alignment has been completed across a cohort of LC-MS samples, and you need to create a unified quantitative matrix for statistical testing, multivariate analysis, or annotation workflows.
-
holobiomicslab Skill Memo Ms API Usage And Parameter ConfigurationUse when you have aligned feature tables (CSV format) with corresponding MS2 spectra data (MGF or mzML files), and need to construct a sample-level vectorization matrix where each row represents a sample and columns encode the occurrence counts of MS2 peaks and neutral losses observed in that.
-
holobiomicslab Skill Pairwise Alignment With Anchor PrioritizationUse when when processing LC-MS metabolomics datasets with 10 or fewer samples and requiring reproducible mass track alignment across the cohort.
-
diegosouzapw Bundle Web ScraperWeb Scraper workflow skill. Use this skill when the user needs Web scraping inteligente multi-estrategia. Extrai dados estruturados de paginas web (tabelas, listas, precos). Paginacao, monitoramento e export CSV/JSON and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
54 -
diegosouzapw Bundle Polars V2Polars workflow skill. Use this skill when the user needs Fast in-memory DataFrame library for datasets that fit in RAM. Use when pandas is too slow but data still fits in memory. Lazy evaluation, parallel execution, Apache Arrow backend. Best for 1-100GB datasets, ETL pipelines, faster pandas replacement. For larger-than-RAM data use dask or vaex and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
54 -
diegosouzapw Bundle Database V2Database Workflow Bundle workflow skill. Use this skill when the user needs Database development and operations workflow covering SQL, NoSQL, database design, migrations, optimization, and data engineering and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
54 -
holobiomicslab Skill Machine Learning Regulatory PredictionUse when you have matched multiomics data (CNV, mutations, DNA methylation, histone PTMs, transcriptomics, miRNA, lncRNA, proteomics, phosphoproteomics) and metabolomics measurements across a cell line panel or cohort, and you want to infer which molecular features (genes, regulatory marks, splice.
-
holobiomicslab Skill Metabolite Distance Metric CalculationUse when you have normalized peak intensity tables from FT-ICR MS data (or MetaboDirect pre-processed .csv output) with samples grouped by experimental treatments (e.
-
holobiomicslab Skill Pytorch Distributed Training ExecutionUse when when you have a pre-trained GNN model checkpoint and need to apply transfer learning to a new chromatography or molecular property prediction dataset (e.g., Eawag_XBridgeC18_364.xlsx) by fine-tuning the model weights on domain-specific examples without retraining from scratch.
-
holobiomicslab Skill Quality Control Threshold OptimizationUse when when you have extracted a peak feature table (CSV or tabular format) with mass-to-charge ratios, retention times, and intensity values across multiple samples, and you need to distinguish genuine differential metabolic signals from instrumental noise or low-abundance background before.
-
holobiomicslab Skill Data Visualization From Tabulated ResultsUse when after executing a MassQL query that returns a tabulated results DataFrame (e.g., MS1 or MS2 scan metadata, peak intensities, retention times), and you need to produce visual summaries suitable for publication, presentation, or exploratory analysis.
-
holobiomicslab Skill Feature List Harmonization Across MethodsUse when you have feature lists in CSV format originating from different acquisition methods (e.
-
holobiomicslab Skill Metadata Normalization And ReconciliationUse when you have multiple CSV feature lists from different acquisition methods (e.g., LC-MS vs LC-IMS-MS) or processing software, each using different naming conventions, retention time scales, or m/z precision;
-
holobiomicslab Skill Retention Time Mass Tolerance CalibrationUse when you have multiple feature tables (CSV files) from different LC-MS analytical experiments, each containing mass, retention time, intensity, isotope, and adduct annotations, and you need to merge them into a single aligned feature matrix.
-
holobiomicslab Skill Hierarchical JSON Structure ConstructionUse when your input is a tabular file (CSV or Excel) with column headers annotated using MESSES tagging syntax (#<table_name>.id for record identifiers and #.
-
holobiomicslab Skill Pandas Dataframe Manipulation Ms ColumnsUse when when you have raw mass spectrometry data (from mzML, Bruker .d, or CSV format) loaded into a Pandas DataFrame and need to ensure it has the correct column structure (m/z, retention time, intensity) before invoking pyOpenMS-Viz plotting functions like .plot(kind='spectrum'), .
-
holobiomicslab Skill Classification Performance VisualizationUse when you have a CSV file with predicted probabilities and true binary labels from a classification model, and you need to evaluate classification performance across decision thresholds and communicate it via a standard diagnostic plot suitable for publication or presentation.
-
holobiomicslab Skill Molecular Formula Parsing And ValidationUse when you have received a formula-assigned FT-ICR MS dataset (CSV or tab-delimited table containing molecular formulas and mass values) and need to convert those formula strings into quantified elemental compositions before computing molecular descriptors, diversity indices, or transformation.
-
holobiomicslab Skill Molecular Structure File Format HandlingUse when when you have molecular structures in one format (e.g., SMILES strings in a spreadsheet or text file) but need to feed them to a tool that accepts a different format (e.g., CypReact requires .sdf or .csv with SMILES).
Frequently asked questions
What are Data & Analytics agent skills?
Data agent skills make AI agents useful for data work: writing SQL, cleaning datasets, building pipelines, working with spreadsheets, and producing analyses. Each skill is a reviewed SKILL.md file that teaches the agent one workflow well, ready to install in seconds.
Which Data & Analytics skills are most installed?
Popular Data & Analytics skills on SkillMD right now include feature-count-verification-across-adducts-and-isotopologues, thermodynamic-molecular-index-calculation-from-elemental-composition, polars. Rankings shift as installs change; sort this page by "Most installs" for the live list.
Do Data & Analytics skills work with Claude Code and Cursor?
Yes. Every skill here ships as a SKILL.md file, an open format that works in Claude Code, Claude.ai, Cursor, Codex, Windsurf, and 60+ other agents. Install one with npx skillmds@latest add <owner>/<name>, or copy the file into your agent's skills directory.