HolobiomicsLab
- 7.4k skills
- 0 followers
- 1 day ago last updated
- ▌ Compound Class Annotation Parsing · holobiomicslabUse when after submitting a fingerprint or spectrum query to the CANOPUS web service and receiving a structured response.
- ▌ Compound Structure Representation · holobiomicslabUse when when you have experimental MS/MS data (peak lists, precursor m/z, charge state, adduct type) paired with a chemical structure (SMILES or structural identifier), and need to create a unified Compound object for spectral alignment, modification-site prediction, or comparative fragmentation.
- ▌ Computational Variation Detection · holobiomicslabUse when after peak detection and feature alignment in a metabolomic LC–MS/MS or GC–MS workflow, when you have a feature table (rows=metabolic features, columns=samples) split into separate .csv files for peak height and peak area.
- ▌ Cross Tool Performance Comparison · holobiomicslabUse when when you have raw tandem MS metabolomics data in vendor formats (.
- ▌ Cross Workflow Format Translation · holobiomicslabUse when you have downloaded a GNPS archive from either GNPS1 (https://gnps.ucsd.edu) or GNPS2 (https://gnps2.org) and need to parse spectra (spectra.mgf), molecular family networks (molecular_families.tsv), spectral library annotations (annotations.
- ▌ Data Summarization And Tabulation · holobiomicslabUse when after obtaining structural clusters from the MAMSI framework using different parameter configurations (e.
- ▌ De Novo Precursor Mass Annotation · holobiomicslabUse when when you have unknown MS/MS spectra with observed precursor m/z values and want to infer the molecular formula and adduct type (e.g., [M+H]+, [M+Na]+, [M+K]+) in a de novo setting without access to spectral libraries.
- ▌ Dependency Requirement Validation · holobiomicslabUse when before launching the DaDIA pipeline or any multi-package R workflow that has strict version constraints. Use this skill when you have access to an R environment and need to confirm that R ≥4.0, XCMS ≥3.11.4, and metaMS ≥1.25.
- ▌ Differential Metabolite Detection · holobiomicslabUse when you have normalized and aligned lipidomic and metabolomic spectral features from the Multi-ABLE method across multiple biological samples grouped by phenotype (e.
- ▌ Disease Classification Prediction · holobiomicslabUse when you have raw LC-MS metabolomics data from multiple disease groups (in .npy or .
- ▌ Environment Dependency Management · holobiomicslabUse when when deploying Galaxy-M or similar multi-component metabolomics platforms that depend on heterogeneous runtime environments (Python, R, MATLAB, WINE) across multiple operating systems (Ubuntu 14.
- ▌ Feature Grouping By Molecular Ion · holobiomicslabUse when after peak picking and sample alignment have produced an aligned feature table with m/z and retention time coordinates.
- ▌ Feature Table Parsing And Loading · holobiomicslabUse when you have extracted volatile organic compound (VOC) features from individual breath samples (mzML or mzXML files) and wish to consolidate multiple per-sample feature tables into a single aligned feature table, or you need to programmatically access feature metadata (m/z, intensity, scan.
- ▌ File Format Export And Validation · holobiomicslabUse when after executing MassQL queries on mass spectrometry data that produce tabulated results (e.g., MS1 or MS2 scan metadata, peak intensities, retention times), and you need to persist those results for archival, sharing, or downstream statistical analysis.
- ▌ Fragment Peak Chemical Annotation · holobiomicslabUse when you have MS/MS spectra with assigned precursor formulas and need to annotate the chemical composition of individual fragment peaks for metabolite structure elucidation or fragmentation pathway analysis. Apply this skill when you want to avoid external fragmentation tree computation (e.
- ▌ Fragmentation Spectrum Extraction · holobiomicslabUse when you have raw or peak-detected mass spectrometry data (mzXML, mzML, or netCDF format) from untargeted metabolomics or exposomics studies and need to separate composite fragmentation spectra into individual constituent spectra for annotation.
- ▌ Fuzzy Analog Search Fragmentation · holobiomicslabUse when when you have experimental MS/MS spectra and want to discover structurally similar compounds beyond exact spectral library matches—particularly useful for identifying chemical analogs, homologs, or isomers that share fragmentation logic but differ in molecular structure.
- ▌ In Source Fragment Identification · holobiomicslabUse when you have an LCMS feature table (from XCMS, MS-DIAL, MZmine2, or other feature extraction software) with m/z, retention time, and intensity columns, plus MS2 spectral annotations from DDA data, and you need to identify which features are in-source fragments rather than distinct metabolites.
- ▌ Intensity Distribution Simulation · holobiomicslabUse when when you need to create synthetic noisy MS/MS spectra from clean baseline spectra to validate denoising algorithms, compare denoising performance across noise levels, or generate ground-truth test datasets where the true signal and noise composition are known and controllable.
- ▌ Ion Image Quantification Workflow · holobiomicslabUse when when you have imzML mass spectrometry imaging data files and need to convert raw ion image intensities into quantitative lipid abundance (pmol/mm²) using known internal standards.
- ▌ Ion Species Annotation Assignment · holobiomicslabUse when after temporal correlation has identified feature pairs with matching intensity profiles across time-resolved MS scans, but before final candidate validation. Apply this skill when you need to distinguish between competing ion-species hypotheses (e.
- ▌ Isotope Pattern Spectral Matching · holobiomicslabUse when when you have LC/MS feature data with observed m/z and intensity values across multiple peaks (monoisotopic and isotopologues) and need to narrow candidate annotations from a metabolite database.
- ▌ Isotopologue Abundance Correction · holobiomicslabUse when you have LC-MS data from isotope labeling experiments where FAM measurements must be transformed to MDV values. Specifically, apply this when: (1) you have measured fractional abundances of isotopologues (FAM) in XLSX format from a high-resolution instrument (e.g., Orbitrap);
- ▌ Keras Regularizer API Integration · holobiomicslabUse when when extending an existing neural network class (e.g., SiameseModel) that lacks user-configurable regularization, and you need to prevent overfitting on moderate-sized training datasets (e.
- ▌ Peak Classification Validation · holobiomicslabUse when after training or loading a NeatMS neural network model, before applying it to filter false positive MS1 peaks in a new dataset.
- ▌ Percentile Threshold Filtering · holobiomicslabUse when you have computed multiple independent scoring functions (e.g., standardised strain correlation and IOKR) for a large set of potential genomic–metabolomic links and wish to identify subsets enriched for validated links.
- ▌ Pipeline Prerequisite Checking · holobiomicslabUse when before launching the DaDIA metabolomics pipeline or any multi-package workflow, when you have an R environment with potentially mixed or unknown package versions and need to confirm that R ≥4.0, XCMS ≥3.11.4, metaMS ≥1.25.
- ▌ Plot Customization And Styling · holobiomicslabUse when after generating a numerical visualization (e.g., confusion matrix, heatmap, or similarity array) using matplotlib, when you need to add axis labels, class names, colormaps, normalization annotations, colorbars, titles, and export the figure in a publication-ready format (PNG or PDF).
- ▌ Precursor Fragment Ion Pairing · holobiomicslabUse when you have raw LC-MS/MS data files (mzML/mzXML format from Thermo, Waters, or Bruker instruments) and a list of target compounds defined by precursor m/z values (and optionally retention time windows).
- ▌ Precursor Product Mass Pairing · holobiomicslabUse when you have centroided MS2 spectra from data-dependent acquisition (ddMS2) in mzML format and seek to prioritize potential PFAS features by detecting diagnostic fragment masses.
- ▌ Publication Figure Preparation · holobiomicslabUse when you have a mass spectrometry spectrum from a supported repository (GNPS, MassBank, MetaboLights, Metabolomics Workbench, ProteoXchange, MS2LDA, or MassIVE) and need to include it in a publication or supplementary material.
- ▌ Python Function Implementation · holobiomicslabUse when when you have classification predictions and ground-truth labels and need to generate a confusion matrix visualization with flexible normalization (by row, column, or all elements) and styling options for publication or diagnostic review.
- ▌ R Data Structure Serialization · holobiomicslabUse when after completing Part 4 (Identification of ISF Features) in the ISFrag workflow, when you have a feature table with identified ISF features and their hierarchical fragmentation relationships, and you need to export this relationship structure for interpretation, integration with external.
- ▌ Ranking Performance Evaluation · holobiomicslabUse when after running retention-order prediction experiments on a test or held-out evaluation dataset.
- ▌ Reaction Pathway Interpolation · holobiomicslabUse when after CREST (version >= 3.0.2) has identified an ensemble of low-energy conformers and stationary points (minima and transition states), and before submitting interpolated geometries to ORCA (version >= 6.0.
- ▌ Rescore Column Standardization · holobiomicslabUse when after running FIDDLE v2.0.0 inference on MS/MS spectra and obtaining ranked formula candidates with confidence scores, apply this skill when the rescore model outputs columns named Rescore (0), Rescore (1), ...
- ▌ Retention Time Drift Detection · holobiomicslabUse when you have LC-MS data processed through XCMS grouping that shows signs of RT drift (e.g., data acquired over extended periods or across many samples) and you suspect misalignment of feature groups.
- ▌ Signal Intensity Normalization · holobiomicslabUse when after loading raw LC-MS data from multiple disease groups when you need to compute correlations between metabolite signals and disease classes, or before training a deep learning model on metabolomics profiles.
- ▌ Signal Smoothing Preprocessing · holobiomicslabUse when when working with raw LC-HRMS profile-mode data containing noisy chromatographic signals, apply this skill before peak detection. Smoothing is particularly needed when the signal-to-noise ratio is low or when gradient-based peak detection would be compromised by high-frequency noise.
- ▌ Similarity Scoring For Spectra · holobiomicslabUse when you have LC-MS/MS query spectra in mgf format that you need to match against a custom database (e.g., prepared with CFM-id) to identify compounds. Apply this skill when you want to rank candidate compounds by spectral similarity and return scored match results for downstream interpretation.
- ▌ Sparse Feature Vector Handling · holobiomicslabUse when when you have tandem mass spectra (mz/intensity pairs with precursor m/z) and need to train an interpretable model (e.
- ▌ Spectral Corpus Representation · holobiomicslabUse when you have a collection of tandem mass spectrometry spectra in mzML or similar format and need to prepare them for LDA-based motif discovery.
- ▌ Spectral Feature Consolidation · holobiomicslabUse when when you have generated separate MemoMatrix objects from independent sample sets (e.g., sample set A and sample set B) and need to align and combine their MS2 fingerprint data into a single matrix for comparative analysis.
- ▌ Spectral Feature Normalization · holobiomicslabUse when when you have raw LC-MS metabolomics data in .mzML or .npy format from multiple disease groups with varying ionization efficiencies or detector sensitivities, and you need to train a deep learning model for disease classification.
- ▌ Spectral Library Data Modeling · holobiomicslabUse when when migrating an existing file-based spectral library (stored as JSON, CSV, or binary formats) into a production system that requires frequent subset queries by metadata filters, similarity scoring across large spectral collections, or integration into downstream tools like MS2Query that.
- ▌ Spectral Library Matching Ripp · holobiomicslabUse when you have tandem mass spectrometry data (LC-MS/MS in MGF, mzXML, mzML, or mzData format) and genomic data from a target organism, and you want to identify RiPPs by matching experimental spectra against a database of predicted post-translationally modified RiPP structures derived from.
- ▌ Substructural Motif Annotation · holobiomicslabUse when you have created a GNPS molecular network (classical or feature-based workflow) and separately run an MS2LDA experiment on the corresponding MGF file, and you want to associate each network node with its constituent substructural motifs and visualize which motifs are shared between.
- ▌ Target List Coordinate Mapping · holobiomicslabUse when you have a CSV-formatted target list with m/z, retention time, or ion mobility identifiers and need to locate and extract peak abundances from raw MS data files (Agilent .d, Thermo .raw, Bruker .d, mzML) acquired across LC-MS, LC-IMS-MS, DDA, DIA, or direct infusion modes.
- ▌ Targeted Metabolite Extraction · holobiomicslabUse when you have centroided LC-MS data (.mzML format) and a curated list of targeted metabolites or lipids (with m/z, retention time, and polarity) that you want to quantify and quality-assess across multiple analytical runs, and you need both per-run AUC values and averaged QC metrics for each.
- ▌ Theoretical Mz Grid Generation · holobiomicslabUse when you have a feature table from untargeted LC-MS (m/z, retention time, intensity) and need to annotate which observed m/z values correspond to isotopologues and adducts of the same neutral compound.
- ▌ Thermo Raw File Format Parsing · holobiomicslabUse when you have acquired multidimensional mass spectrometry data (MS1, MS/MS, or data-independent acquisition) from a Thermo instrument saved in the proprietary '.
- ▌ Annotation Confidence Assessment · holobiomicslabUse when after MS-FINDER in silico annotation has been executed on exported LC-MS features and multiple database matches (with HRR scores) have been returned.
- ▌ Clustered Peak Output Formatting · holobiomicslabUse when after peak clustering has been completed in pyINETA (i.
- ▌ Docker Environment Configuration · holobiomicslabUse when you need to deploy CloMet for the first time on a new system, or when you want to ensure reproducible execution of metabolomics data harmonization tasks without manual dependency management.
- ▌ File API Endpoint Implementation · holobiomicslabUse when when you need to construct a POST endpoint that ingests raw spectral data files from multiple vendor formats (jcamp, RAW, mzML) and must standardize them for downstream processing.
- ▌ Metabolite Dataset Preprocessing · holobiomicslabUse when you have raw NMR metabolomics measurements paired with pre-analytical metadata (e.g., processing delay times, sample type designations [plasma vs. serum], cohort identifiers) and need to investigate how delays affect measured metabolic parameters.
- ▌ Metabolite Identifier Annotation · holobiomicslabUse when you have observed compounds (from LC-MS/MS, GC-MS, NMR, or other analytical techniques) with unknown identity and you want to assign candidate metabolite structures by comparing them to computationally predicted metabolism pathways.
- ▌ Metabolomics Data Representation · holobiomicslabUse when you have a raw MGF file containing fragmented LC-MS-MS metabolomics spectra and want to apply Latent Dirichlet Allocation (LDA) to discover hidden topics (molecular families, biochemical patterns) across your sample set.
- ▌ Metabolomics File Format Parsing · holobiomicslabUse when you have mwTab-formatted files from the Metabolomics Workbench containing MS or NMR experimental metadata and tabular data sections (e.g., METABOLITES, DATA blocks), and need to load them into memory for downstream conversion, validation, or analysis rather than manual text parsing.
- ▌ Model Metadata Schema Inspection · holobiomicslabUse when when preparing to send peak data (1H and 13C NMR measurements) to a machine learning classification endpoint and you need to verify the current model's input/output names and schema, especially before implementing or updating code that constructs JSON payloads for the /api/smart3/search.
- ▌ Molecular Connectivity Inference · holobiomicslabUse when you have 1D ¹H and/or ¹³C NMR spectra (as preprocessed numerical arrays or peak lists) from an unknown organic molecule with ≤19 heavy atoms, and you need to recover its molecular formula and connectivity graph.
- ▌ Nmr Spectrum Object Construction · holobiomicslabUse when you have Bruker NMR spectral files (raw instrumental output) and need to prepare them for automated metabolite identification and quantification.
- ▌ Replicate Consistency Assessment · holobiomicslabUse when after NMR or MS data acquisition and preprocessing (phasing, baseline correction) when you have a SummarizedExperiment object containing assay intensity matrix with QC sample columns designated.
- ▌ Signal To Noise Ratio Assessment · holobiomicslabUse when when preparing a 1D 1H NMR spectral peak list for input to the NMRformer metabolite identification model, and you have access to peak intensity measurements and noise level estimates.
- ▌ Spatial Overlap Analysis Imaging · holobiomicslabUse when when annotating matrix-related peaks in MSI datasets where candidate peaks have identical or near-identical m/z values (isobaric ions), or when multiple peaks exhibit overlapping spatial distributions across the tissue image that could confound downstream annotation filtering.
- ▌ Spearman Correlation Computation · holobiomicslabUse when : (1) you have metabolomic data (NMR or MS-derived) and a continuous phenotype variable; (2) you need to quantify associations while controlling for known confounders (age, gender, disease status);
- ▌ Spectra To Structure Elucidation · holobiomicslabUse when you have one or more spectroscopic datasets (IR, Raman, UV-Vis, mass spectra, NMR) from an unknown compound and need to generate candidate molecular structures ranked by likelihood. Use this when retrieval-based approaches are infeasible (e.
- ▌ Spectral Roi Boundary Definition · holobiomicslabUse when when you have identified a spectral window of interest in a 1H NMR spectrum from a complex biological sample (serum, urine, CSF, tissue, saliva, or sweat) and need to systematically retrieve all metabolites from HMDB whose reference proton NMR chemical shifts fall within that window.
- ▌ Spectroscopic Data Preprocessing · holobiomicslabUse when when you have raw spectroscopic measurements in heterogeneous formats (IR, Raman, UV-Vis, mass spectra, or NMR) and need to feed them into a spectrum-conditioned diffusion model for de novo molecular structure elucidation.
- ▌ Summarized Experiment Subsetting · holobiomicslabUse when when you have a SummarizedExperiment containing metabolomic abundances and a corresponding vector of quality metrics (e.g., coefficient of variation computed across QC samples), and you need to filter to retain only features meeting a reproducibility threshold (e.g., CV ≤ 0.
- ▌ Tool Initialization Verification · holobiomicslabUse when after completing Docker installation and container build steps for CloMet, before attempting substantive data analysis or pipeline execution.
- ▌ Wasserstein Distance Computation · holobiomicslabUse when when you have both an observed NMR mixture spectrum and a candidate reconstructed spectrum (each represented as intensity distributions across chemical shift bins), and you need a scalar similarity metric to evaluate how closely the reconstruction matches the observed data.
- ▌ Atac Seq Signal Normalization · holobiomicslabUse when you have aligned ATAC-seq BAM files and need to detect transcription factor binding sites via footprint analysis. The skill is essential because raw Tn5 insertion signal contains systematic bias toward certain DNA sequences;
- ▌ Encode Hic Pipeline Execution · holobiomicslabUse when you have raw Hi-C FASTQ files from a public repository (NCBI SRA, GEO, or ENCODE-deposited accession) and need to reproduce or validate Hi-C map generation following the ENCODE uniform processing standard, or you need to verify that your pipeline output conforms to reference format and.
- ▌ Genomic Coordinate Conversion · holobiomicslabUse when you need to map computed per-bin metrics (insulation scores, boundary calls, contact frequencies) back to genomic coordinates for export to BED/GFF format, cross-reference with external annotations, or validate that computed features fall within expected genomic ranges.
- ▌ Hi C Fastq Read Preprocessing · holobiomicslabUse when you have raw Hi-C FASTQ files from a Hi-C wet-lab protocol and need to convert them into processed Hi-C contact maps (.hic files) for loop detection, TAD identification, or 3D structure inference. Use when starting from deposited public Hi-C datasets (e.
- ▌ Kmer Motif Synergy Assessment · holobiomicslabUse when after computing deviations for both motif and kmer annotations on the same chromVAR dataset, when you need to determine whether kmers and motifs are redundant predictors of chromatin accessibility variability or provide complementary information for downstream clustering, annotation, or.
- ▌ Monocle3 Trajectory Embedding · holobiomicslabUse when you have an ArchR project object with dimensionality reduction results (LSI or combined dimensions from scATAC-seq ± scRNA-seq) and want to infer pseudotime trajectories and cell-state transitions.
- ▌ Python Environment Management · holobiomicslabUse when you are preparing to run Hi-C data normalization or read alignment filtering steps that depend on Python modules (iced, pysam, numpy, scipy) and you need to ensure consistent module versions across multiple runs or compute nodes.
- ▌ Scatac Seq Peak Matrix Export · holobiomicslabUse when after completing peak calling and cell annotation in an ArchR project, when you intend to perform trajectory analysis using STREAM rather than ArchR's native monocle3 or Slingshot integrations, or when you need to share peak-by-cell matrices with collaborators using STREAM pipelines.
- ▌ Sparse Matrix Subset Indexing · holobiomicslabUse when when you have a chromVARDeviations object with multiple annotation sets (e.
- ▌ Tn5 Insertion Bias Correction · holobiomicslabUse when you have aligned ATAC-seq BAM files from Tn5-based chromatin accessibility assays and need to perform footprinting analysis.
- ▌ API Specification Extraction · holobiomicslabUse when you have access to the source code of a webservice component (Python, configuration files, route definitions) and need to produce machine-readable API documentation (OpenAPI 3.
- ▌ Backend Routing And Dispatch · holobiomicslabUse when you need to support multiple plotting backends for the same data visualization task, and you want to centralize backend selection logic so that users can specify their preferred rendering engine (matplotlib, bokeh, or plotly) at call time without modifying the core plotting logic.
- ▌ Command Line Tool Invocation · holobiomicslabUse when you have an existing mass spectrometry data file in a vendor or standard format (mzML, NetCDF, etc.) and need to convert it to mzPeak format for downstream analysis, visualization, or archival.
- ▌ Configuration Object Pattern · holobiomicslabUse when when designing a library that needs to support multiple plotting backends (e.g., matplotlib, bokeh, plotly) and you want to avoid reimplementing parameter validation, storage, and dispatch logic for each backend.
- ▌ Conversion Directive Parsing · holobiomicslabUse when you have validated intermediate JSON data (conforming to the Experiment Description Specification) and need to configure how it should be converted to a supported output format (e.g., mwTab for Metabolomics Workbench submission).
- ▌ Dataset Integrity Assessment · holobiomicslabUse when when you have downloaded a released version of a structured dataset (e.g., LOTUS from Zenodo) and need to confirm it matches the documented headline statistics before downstream analysis, or when auditing data integrity after ingestion into a processing pipeline.
- ▌ Enrichment Ratio Calculation · holobiomicslabUse when after generating combined or alternative scores for a set of BGC-metabolite (GCF-MF) link candidates, you need to evaluate whether a scoring function preferentially ranks true validated links higher than spurious ones.
- ▌ Export Tag Syntax Validation · holobiomicslabUse when you have a tabular file (CSV or Excel) that has been manually or semi-automatically tagged with export tags, and you need to verify tag correctness before running the extract command to convert the tagged table into intermediate JSON.
- ▌ Github Repository Operations · holobiomicslabUse when when you need to verify that a software package (such as MassQL) passes its periodic integration test suite as indicated by CI workflow badges in the project documentation, or when you must reproduce pass/fail results for package-testing workflows distinct from unit tests to establish.
- ▌ HTTP API Integration Testing · holobiomicslabUse when you have deployed a microservice (e.g., TensorFlow Serving, REST API) and need to verify that specific endpoints (e.g., /model/metadata, /classify) return responses with correct schema, field names, and data types before consuming them in production workflows or downstream applications.
- ▌ HTTP Endpoint Identification · holobiomicslabUse when when you have source code access to a webservice component (such as the MAGMa joblauncher) and need to enumerate all exposed HTTP endpoints, their methods (GET, POST, etc.), URL patterns, parameter names, request/response payload structures, and authentication requirements in order to.
- ▌ Mandatory Field Verification · holobiomicslabUse when after generating mzPeak files from prototype implementations (Rust, Python, R, or .NET) or after format conversion, and before integrating files into a mass spectrometry data repository or sharing them with collaborators. Use it when specification compliance is a hard requirement (e.
- ▌ Metabolomics Ora Methodology · holobiomicslabUse when you have a metabolomics dataset and want to perform pathway enrichment analysis using ORA, but need to first understand its behavior, limitations, and correct application through reproducible simulation.
- ▌ Model Deployment Preparation · holobiomicslabUse when you have a pre-trained Keras model and need to deploy it via a Docker-based TensorFlow Serving API (e.g., for molecular classification via SMILES), but the model's layer naming or format does not yet match the target runtime's expectations (e.
- ▌ Msp File Parsing And Writing · holobiomicslabUse when you have one or more .msp spectral library files (NIST format) that need to be ingested for metadata curation, enrichment via web services, or export after transformation. Use this skill as the entry and exit point for any .msp-based annotation or analysis pipeline.
- ▌ Mzml Data Access Abstraction · holobiomicslabUse when when building or extending a mass spectrometry data parser that must support multiple mzML storage formats (plain .mzML, indexed .mzML.gz, standard-compressed .mzML.
- ▌ Pathway Database Integration · holobiomicslabUse when you have intensity measurements (peak features, protein intensities, or gene expression values) with compound or gene annotations (KEGG IDs, ChEBI IDs, UniProt IDs, or ENSEMBL IDs), and you need to aggregate them into biologically meaningful pathway groups for differential analysis.
- ▌ Peak Abundance Normalization · holobiomicslabUse when after peak detection and before any comparative analysis (e.g., diversity indices, ordination, or statistical testing) when working with direct injection FT-ICR MS data where raw peak intensities vary across samples due to instrumental factors. The task_id=task_005 example applies it to S.