HolobiomicsLab
- 7.4k skills
- 0 followers
- 1 day ago last updated
- ▌ Deconvolved Spectrum Comparison · holobiomicslabUse when after auto-deconvolution of GC-MS data has produced a table of deconvolved mass spectra, and you need to organize these spectra into clusters or detect which compounds co-elute or share similar fragmentation patterns.
- ▌ Drug Database Record Extraction · holobiomicslabUse when when you have obtained a DrugBank release file (requiring access credentials) and need to integrate drug chemical structure, name, and identifier information into a metadata cleanup or chemical enrichment pipeline.
- ▌ Dysregulation Score Computation · holobiomicslabUse when you have per-sample metabolite abundance data (e.g., from LC-MS or GC-MS) and a metabolite-to-pathway assignment table, and need to generate a sample-by-pathway dysregulation matrix for downstream classification, prognosis prediction, or pathway-level phenotype association.
- ▌ Gc Ms Spectral Library Matching · holobiomicslabUse when when you have GC-MS data with detected peaks that require structural annotation, retention index calibration has been applied (typically using FAMES standards), and you need to assign compound identities with confidence scores.
- ▌ Mass Spectral Data Augmentation · holobiomicslabUse when when you have limited real GC-MS overlapped peak data but need thousands of labeled examples to train a deep learning model for mass spectral deconvolution.
- ▌ Mass Spectral Matrix Prediction · holobiomicslabUse when you have raw GC-MS data with overlapped peaks in a specific retention time region and need to resolve the individual pure mass spectra of all components present in that region.
- ▌ Molecular Network Graph Parsing · holobiomicslabUse when after GNPS_GC molecular networking job completion, when you have retrieved raw network output files and need to extract, validate, and structure the network topology for further metabolite assignment, comparative network analysis, or visualization.
- ▌ Ms Peak Table Format Validation · holobiomicslabUse when immediately after loading a raw GC-MS CSV file and before executing the spreadOut() function. Use it when you have received peak table data from an instrument vendor (e.
- ▌ Peak Detection Cwt Optimization · holobiomicslabUse when when you have a pre-aligned GCIMSDataset and need to systematically identify and annotate chromatographic peaks across multiple samples.
- ▌ Qc Type Color Marker Assignment · holobiomicslabUse when configuring a new injection-plate design template in InjectionDesign if you have multiple QC types to position on a plate and need to visually distinguish them in the final worksheet.
- ▌ Raw Chromatography Data Parsing · holobiomicslabUse when you have raw GC-MS output files (vendor formats or netCDF) from a chromatography instrument and need to prepare them for automated peak deconvolution and spectral analysis.
- ▌ Spectral Peak Clustering Gc Ims · holobiomicslabUse when after peak detection on individual GC-IMS samples, when you need to assign consistent cluster IDs to peaks detected across multiple samples to enable cross-sample comparison and quantification.
- ▌ Xcms Parameter Optimization Msw · holobiomicslabUse when when you have direct-injection or low-complexity mass spectrometry data (mzML files) and need to detect chromatographic peaks using wavelet-based methods instead of centWave, especially when standard retention-time-dependent peak detection is not suitable or when you need to tune.
- ▌ Algorithm Interface Abstraction · holobiomicslabUse when you have multiple independent peak-picking algorithms available and need to allow end-users to select among them for the same analytical task (peak detection in untargeted LC-MS data) without coupling the rest of your pipeline to each algorithm's API.
- ▌ Background Ion Blank Comparison · holobiomicslabUse when after generating an initial LC-MS feature table (from Asari or similar peak detection) and before normalization or statistical analysis, especially when blank samples (e.g., solvent-only or buffer-only runs) were acquired alongside study samples.
- ▌ Batch Generation And Validation · holobiomicslabUse when you have raw mzML files and a feature table (CSV) from LCMS data processed by tools like mzMine, and you need to create train/test/validation batches with specific matrix dimensions (120 × 2) and verified margin/peak signal separation before training or evaluating a neural network.
- ▌ Blank Intensity Ratio Filtering · holobiomicslabUse when apply this filter after feature detection and before downstream statistical analysis when your experimental design includes blank samples (e.
- ▌ Candidate Metabolite Prediction · holobiomicslabUse when you have an unknown compound's mass spectrum (m/z peaks and intensities) in positive or negative ion mode and need to identify candidate metabolites from a structure database.
- ▌ Chemical Annotation Integration · holobiomicslabUse when after training a tandem mass spectrometry embedding model (such as MSBERT) on a reference spectral library (e.g., GNPS), apply this skill to confirm that the high-dimensional embedding vectors project into interpretable chemical space.
- ▌ Chemical Formula Representation · holobiomicslabUse when you need to feed chemical formulas into a neural network-based formula scorer (such as MIST-CF) that must learn data-dependent representations of formula structure and composition.
- ▌ Chemical Metadata Harmonization · holobiomicslabUse when when aggregating MS/MS spectra from multiple public repositories (GNPS, MassBank, Mona) or in-house sources with inconsistent metadata naming conventions, missing or malformed adduct annotations, or incomplete chemical structure annotations (SMILES/InChI/InChIKey).
- ▌ Compound Identification Scoring · holobiomicslabUse when you have preprocessed MS/MS spectra (noise-filtered, normalized) and need to compute pairwise similarity or distance scores for compound library matching, when your goal is to rank candidate compounds by spectral match quality and maximize correct identification rate above dot-product.
- ▌ Configuration File Modification · holobiomicslabUse when when you need to expand or contract a mass spectrometry dataset by adding or removing allowed instrument types (e.
- ▌ Contrastive Loss Implementation · holobiomicslabUse when training embeddings from MS/MS spectra and you need to simultaneously enforce: (1) discrimination between spectra with different structural properties via contrastive learning on peak information and metadata embeddings, and (2) accurate reconstruction of embeddings from peak features via.
- ▌ Metabolite To Gene Mapping · holobiomicslabUse when you have metabolomic data (e.g., from LC-MS or GC-MS comparing patient to controls) showing differential abundant metabolites (DAMs), candidate genes from exome sequencing or variant calling, and access to a protein–protein or gene–gene interaction network (e.g., STRING).
- ▌ R S4 Object Accessor Usage · holobiomicslabUse when you have constructed or received a SummarizedExperiment object (or similar S4 class) containing MS feature tables, counts matrices, or sample-level metadata, and need to retrieve specific slots (e.
- ▌ Sample Metadata Extraction · holobiomicslabUse when when you have an Excel file uploaded by a user following the InjectionDesign template schema and need to convert it into a modifiable, structured sample list that preserves up to three classification dimensions and QC type assignments for LC/GC-MS multi-omics experiments.
- ▌ Transformer Model Training · holobiomicslabUse when you have a dataset of augmented simulated overlapped GC-MS peaks and need to train a Transformer model to automatically deconvolve them into pure component mass spectra.
- ▌
- ▌ Candidate Ranking By Score · holobiomicslabUse when you have a query mass spectrum and a set of candidate molecular structures, and you need to prioritize candidates by their likelihood of matching the query. Typical triggers include: (1) you have computed or extracted spectral features (e.
- ▌ Cluster Assignment Mapping · holobiomicslabUse when after density-based clustering (e.g., DBSCAN) has been performed on a sparse pairwise distance matrix derived from MS/MS spectra nearest neighbor indexes.
- ▌ Composite Spectra Analysis · holobiomicslabUse when your untargeted LC/HRMS dataset contains Data-Independent Acquisition data (MS^E, AIF, or SWATH-MS) or MS1-only composite spectra where multiple precursor ions fragment simultaneously, and you need to deconvolve overlapping fragmentation spectra to enable accurate chemical structure.
- ▌ Compound Candidate Ranking · holobiomicslabUse when after compound database dereplication has produced candidate annotations (from SIRIUS or MetFrag) and you need to select the most reliable candidates for final annotation.
- ▌ Compound Database Matching · holobiomicslabUse when you have MS2 .mzML spectral data from untargeted metabolomics and need to assign chemical identities to detected precursor ions.
- ▌ Compound Identifier Lookup · holobiomicslabUse when you have an experimental MS/MS spectrum (m/z and intensity pairs in mzML/mzXML format from DDA or targeted acquisition on Thermo, Waters, or Bruker instruments) and need to identify the unknown compound by comparing it against a reference database.
- ▌ Container Image Management · holobiomicslabUse when your LC-HRMS metabolomics data (.mzML or .abf files) must be processed reproducibly across multiple machines (local workstations, HPCs, cloud) without manual tool installation, or when you need to enforce identical computational environments for peer review and long-term archival.
- ▌ Cross Omics Strain Mapping · holobiomicslabUse when when you have paired genomics (AntiSMASH BGC annotations) and metabolomics (GNPS spectra and molecular families) data from the same microbial strains and need to identify which biosynthetic pathways produce which observed natural products.
- ▌ Custom Data Schema Mapping · holobiomicslabUse when a practitioner has pre-computed features from an external feature-finding procedure (e.g., vendor software, alternative open-source tools) and wishes to incorporate them into PFΔScreen's PFAS prioritization pipeline without re-detecting features from raw mzML data.
- ▌ De Novo Peptide Sequencing · holobiomicslabUse when you have annotated MS/MS spectra in MGF format and need to identify peptide sequences that may not exist in reference protein databases—such as in immunopeptidomics, metaproteomics, paleoproteomics, venomics, or monoclonal antibody assembly workflows.
- ▌ Dependency Version Parsing · holobiomicslabUse when before launching a multi-tool computational workflow (e.g., QCxMS2 mass spectra calculations) that depends on external programs with version-sensitive APIs or features. Apply this skill when: (1) the workflow has explicit minimum version requirements for one or more dependencies;
- ▌ Distance Matrix Clustering · holobiomicslabUse when you have a sparse pairwise distance matrix derived from nearest neighbor indexing of MS/MS spectra (or similar high-dimensional objects) and need to partition spectra into groups based on local density and neighborhood connectivity.
- ▌ Docker Container Execution · holobiomicslabUse when you have GNPS-style MGF spectral files as input and need to run Mass2SMILES MS/MS-to-structure inference without installing TensorFlow, CUDA, or Python dependencies locally.
- ▌ Dual Branch Feature Fusion · holobiomicslabUse when when you have molecular input data available in two or more distinct formats (e.g., RDKit-extracted fingerprints AND torch_geometric Graph objects representing molecular topology) and your prediction target (e.
- ▌ Eic File Format Generation · holobiomicslabUse when after completing MS2 annotation in the JPA metabolomics workflow, when you have aligned feature data (feature matrix with m/z, retention time, intensity, and sample assignments) and need to extract and export EIC traces for individual features or feature subsets for external validation.
- ▌ Email Delivery Integration · holobiomicslabUse when when a QC check fails during an active LC-MS run and configured email notification targets exist in the system. Use this skill to ensure that QC failures are communicated to stakeholders immediately, complementing Slack-based alerts for users who prefer or require email notification.
- ▌ Entropy Similarity Scoring · holobiomicslabUse when you need to quantify the degree of match between two MS/MS spectra—either to validate that a denoised spectrum remains faithful to a reference ground-truth spectrum, or to rank candidate library matches for a query spectrum.
- ▌ Executable Path Resolution · holobiomicslabUse when when preparing to run QCxMS2 or similar multi-tool orchestration software that depends on five or more external programs with strict version floors.
- ▌ External Registry Querying · holobiomicslabUse when your project JSON document contains genome identifiers but lacks organism name or taxonomic annotations. The platform needs to auto-populate these fields to enable browsing and cross-linking with public genomic databases. Trigger this skill when you have genome IDs (e.
- ▌ Feature Annotation Mapping · holobiomicslabUse when you have a trained decision tree model on ChemEcho sparse feature vectors and need to convert a specific decision path (root to leaf) into a deployable query.
- ▌ Fragment Ion Mass Matching · holobiomicslabUse when you have a tandem mass spectrum (MSMS) of a known or hypothesized peptide, along with its ProForma 2.
- ▌ Gnps Network File Handling · holobiomicslabUse when you have a GNPS molecular network job and need to programmatically load the network structure, merge external annotations (chemical class, MS2LDA substructural motifs), and export a unified annotated network for visualization in Cytoscape or other graph analysis tools.
- ▌ Gradient Flow Verification · holobiomicslabUse when after implementing a composite loss function that combines multiple loss terms (e.g., InfoNCE contrastive loss and MSE reconstruction loss) in a PyTorch module, and before running full-scale training on MS/MS spectra data.
- ▌ Graph Attribute Annotation · holobiomicslabUse when after dereplication and cosine similarity clustering have been completed on merged LC-MS/MS data, when you need to construct the final molecular network output with predicted molecules as nodes and their parent ions as a second node class, connected by edges that preserve the fragmentation.
- ▌ Hmdb Metabolite Extraction · holobiomicslabUse when you have downloaded raw HMDB data (hmdb_metabolites.zip or pickle file) and need to generate a representative set of chemical objects for simulating LC-MS/MS acquisition strategies.
- ▌ Interactive Plot Embedding · holobiomicslabUse when you have resolved USI (Unified Spectrum Identifier) spectrum data from a supported repository (GNPS, MassBank, MetaboLights, Metabolomics Workbench, ProteoXchange, or MS2LDA) and need to create a figure suitable for journal publication or supplementary materials that retains a link to the.
- ▌ Ionization Mode Annotation · holobiomicslabUse when when converting MS/MS spectra from .msp format library files (e.g., MassBank) into a custom fragment library for metabolite annotation, and the source spectra are tagged with ionization mode information (positive or negative).
- ▌ Large Scale Data Retrieval · holobiomicslabUse when you have a query mass spectrum (or a metabolite reference spectrum from public data) and need to search it against a large-scale spectral repository (≥billions of spectra, e.g., GNPS library) where execution time and resource efficiency are critical.
- ▌ Lc Ms Spectral Data Import · holobiomicslabUse when you have raw LC-MS/MS spectral data in .mgf format (or vendor-specific raw data that can be converted to .mgf via MZmine or similar tools) and need to prepare it for interactive exploration using the specXplore dashboard.
- ▌ Lcms Feature Table Parsing · holobiomicslabUse when you have raw nontargeted LCMS feature tables from one or more analytical methods in tabular format (with m/z, RT, and intensity columns) that need to be aligned or clustered, or when integrating multiple feature tables into a shared BMXP processing pipeline that requires standardized.
- ▌ Lipid Maps Database Lookup · holobiomicslabUse when when you have a parsed lipid species table (output from LipidSearch or LIQUID) containing lipid names or identifiers, and you need to annotate each entry with its standardized LIPID MAPS category (e.g., Glycerophospholipids, Sphingolipids) and subcategory (e.
- ▌ Lipid Nomenclature Mapping · holobiomicslabUse when you need to generate a comprehensive, non-redundant inventory of lipid species that span a defined lipid class (e.g., phosphatidylcholine, triacylglycerol) and a range of fatty acid compositions (e.g., C14:0 to C22:6).
- ▌ Lipid Spectral Data Export · holobiomicslabUse when after generating a complete lipid spectral library with adduct-specific fragmentation patterns and retention time metadata, when you need to deploy the library for targeted or data-dependent acquisition on specific mass spectrometry instruments—either Excalibur-controlled orbitrap.
- ▌ M Z Based Feature Grouping · holobiomicslabUse when after retention-time clustering has grouped features from multiple LC-MS samples, when you need to refine feature assignments by enforcing m/z consistency and eliminate duplicate or near-duplicate features with the same mass but potentially misaligned retention times.
- ▌ Mass Spectrum Scan Parsing · holobiomicslabUse when you have raw mass spectrometry data files from a Thermo instrument (e.
- ▌ Masst Output Visualization · holobiomicslabUse when you have completed one or more domain-specific MASST searches (microbeMASST, plantMASST, tissueMASST, microbiomeMASST, foodMASST) and have aggregated search outputs (matches.tsv, library.tsv, datasets.
- ▌ Memomatrix Object Handling · holobiomicslabUse when you have generated one or more MemoMatrix objects (MS2 fingerprint matrices from separate sample sets) and need to combine them for cross-cohort alignment, validate structural consistency after merging, or prepare merged matrices for downstream filtering and visualization.
- ▌ Model Artifact Persistence · holobiomicslabUse when after a deep neural network model has completed training on LC-MS spectral peak classification data and you need to preserve the learned weights and architecture for downstream inference, validation on held-out test sets, or sharing with collaborators.
- ▌ Ms2 Spectrum Consolidation · holobiomicslabUse when after sample alignment and feature grouping steps in untargeted LC-MS workflows, when you have DDA-mode raw files with both MS1 and MS2 scans and need to link tandem mass spectra to quantified features for annotation and structural characterization.
- ▌ Mzml Feature Table Parsing · holobiomicslabUse when you have raw LCMS data in mzML format and a feature table (CSV) from a peak detection pipeline (e.g., MZmine) and need to prepare these inputs for NeatMS preprocessing, batch creation, or peak classification. This skill is the mandatory entry point for any NeatMS workflow.
- ▌ Network File Format Export · holobiomicslabUse when after completing dereplication and cosine similarity clustering in the MolNotator pipeline, when you have finalized molecular network data with molecule–ion relationships and need to visualize, analyze, or share the network in external software.
- ▌ Pathway Annotation Mapping · holobiomicslabUse when you have a metabolomics dataset with metabolite identifiers in mixed formats (e.g., common names, KEGG accessions, HMDB IDs) and need to assign each metabolite to its canonical pathway(s) before performing pathway-level classification, feature selection, or prognosis modeling.
- ▌ Peak Boundary Localization · holobiomicslabUse when after you have (1) extracted regions of interest (ROIs) around candidate peaks in LC-MS data and (2) run those ROIs through a trained CNN-Transformer peak detection model that outputs both binary peak classifications and bounding box coordinates.
- ▌ Peak Data Integrity Checks · holobiomicslabUse when implementing or modifying a writable MsBackend subclass (e.g., MsBackendMemory, MsBackendDataFrame) and need to replace peak data (m/z values, intensity values, or peaksData).
- ▌ Precursor Mass Calculation · holobiomicslabUse when when you have a compound's SMILES string or molecular formula and need to determine the expected precursor ion m/z for comparison against observed spectra, particularly before applying formula-based denoising, entropy similarity scoring, or denoising search against reference libraries.
- ▌ Python Library Integration · holobiomicslabUse when when you have Thermo Fisher RAW mass spectrometry files and need to extract mass-to-charge ratios, intensities, scan metadata, and peak lists within a Python script or notebook for downstream computational analysis, and you require programmatic control over extraction parameters rather.
- ▌ Python Package Integration · holobiomicslabUse when when you have LC-MS/MS data in MZmine-generated MGF and CSV files (for positive and/or negative ionization modes) and need to apply a sequence of deduplication, annotation, and dereplication steps defined in a MolNotator YAML configuration file to predict actual molecules and build.
- ▌ Pytorch Module Development · holobiomicslabUse when when constructing a composite loss function for contrastive learning on structured data (e.
- ▌ Qc Summary Data Extraction · holobiomicslabUse when after applying one or more mpactr filters (filter_mispicked_ions, filter_group, filter_cv, filter_insource_ions) to a feature table, use qc_summary() to extract the pass/fail status of each ion across all applied filters.
- ▌ Qiime2 Artifact Inspection · holobiomicslabUse when you need to verify that a QIIME 2 artifact (e.g., a Chemical Feature Tree from q2-qemistree, a FeatureTable[Frequency], or a Phylogeny[Rooted] object) has been correctly produced, before using it as input to downstream analyses.
- ▌ Qvalue Threshold Filtering · holobiomicslabUse when after loading search result files (e.g., from DIA-NN or OpenSwath) containing feature identification results with associated Q-value scores, apply this filter when you need to select a subset of high-confidence identifications before generating comparison plots or summary statistics across.
- ▌ Raw File Header Extraction · holobiomicslabUse when beginning an LC-MS data analysis pipeline and you need to rapidly inspect instrument metadata, acquisition parameters, or scan statistics from proprietary Thermo .raw files without the I/O overhead of loading full spectra.
- ▌ Retention Time Mz Indexing · holobiomicslabUse when when you have parsed .mzML or Thermo .raw LC-MS data and need to support interactive or programmatic queries by retention time (RT) and mass-to-charge ratio (m/z) without re-scanning the entire file.
- ▌ Sample Membership Tracking · holobiomicslabUse when when aligning detected features across multiple LC-IMS-MS/MS samples and you need to identify which input samples contributed to each consensus feature cluster, especially to filter out spurious or low-confidence alignments, validate clustering completeness, or perform sample-specific.
- ▌ Schema Constraint Checking · holobiomicslabUse when a user uploads a JSON project file to the platform and you need to verify it matches the required format defined in app/public/schema.json before accepting it into the database.
- ▌ Spatial Coordinate Mapping · holobiomicslabUse when after LC-MS feature detection, alignment, quantification, and optional filtering/normalization are complete, and you have intensity values for molecular features that correspond to spatial positions (e.g., tissue coordinates, imaging pixel locations).
- ▌ Spectra Mgf Format Loading · holobiomicslabUse when when you have downloaded a GNPS molecular networking archive (GNPS1 or GNPS2 workflow output) and need to reconstruct spectral records for integration with genomic data (BGCs, antiSMASH results) or for computing molecular family links and spectral similarity scores.
- ▌ Spectral Alignment Scoring · holobiomicslabUse when you have paired MS/MS spectra (known compound and its structural analog) with assigned precursor m/z, charge, and SMILES; you want to quantify which parts of the molecular structure could have undergone modification by scoring peak alignment quality.
- ▌ Spectral Format Conversion · holobiomicslabUse when when raw spectral data exists in one mass spectrometry file format but downstream analysis requires a different format; when integrating spectra from multiple sources or instruments that produce heterogeneous file formats;
- ▌ Spectral Output Formatting · holobiomicslabUse when after generating tandem mass spectrum predictions from a neural model (ICEBERG, SCARF, or baseline), and before attempting retrieval ranking, metric computation, or validation against experimental spectra.
- ▌ Spectral Window Extraction · holobiomicslabUse when you have loaded multidimensional MS data (from MZA HDF5 files or other formats) and need to examine a specific m/z region—for example, to visualize a known lipid or metabolite mass range, perform peak detection within a narrow window, or reduce computational overhead by working on a subset.
- ▌ Spectrum Alignment Scoring · holobiomicslabUse when after generating probability predictions for potential modification sites (via ModiFinder.
- ▌ Spectrum Metadata Handling · holobiomicslabUse when when you have a USI (Universal Spectrum Identifier) string referencing a spectrum in a public metabolomics repository (GNPS Molecular Networking, GNPS Spectral Libraries, MassBank, MetaboLights, Metabolomics Workbench, MS2LDA, or ProteoXchange) and need to extract its raw spectral data.
- ▌ Spectrum Subset Extraction · holobiomicslabUse when after duplicate filtering of MZmine-exported MGF and CSV files, when you have combined spectra from multiple samples in a single MGF and need to segregate them by sample identifier before fragment annotation or adduct assignment.
- ▌ Taxonomy Database Querying · holobiomicslabUse when a paired omics project record contains a genome identifier (e.g., from GenBank or NCBI) but lacks the corresponding organism scientific name.
- ▌ User Interface Integration · holobiomicslabUse when when you have a multi-step computational workflow (e.g., peak detection, filtering, manual review) implemented in R and need to expose it to end-users who lack R expertise.
- ▌ Bond Connectivity Prediction · holobiomicslabUse when you have predicted or partially assembled molecular fragments (as token sequences or substructure embeddings) and need to determine which atoms are bonded to which—that is, when the formula (atom inventory) is known or predicted but the connectivity graph is uncertain.
- ▌ Cohort Performance Reporting · holobiomicslabUse when you have NMR metabolite measurements from peripheral blood samples (plasma/serum) paired with processing delay metadata (pre-centrifugation and post-centrifugation times) and need to benchmark metabolic parameter stability across delay windows.
- ▌ Column Name Variant Matching · holobiomicslabUse when when processing mwTab metabolomics data files with variable column naming conventions (e.
- ▌ Hybrid Model Fusion Strategy · holobiomicslabUse when you have 1H NMR spectral data from complex mixtures and need to identify component compounds, but a single architecture (CNN or Transformer alone) fails to capture both fine local patterns in peak structures and long-range dependencies across the full spectral range.
- ▌ Isotopic Impurity Accounting · holobiomicslabUse when when analyzing LC-MS data from stable isotope labeling experiments where measured isotopologue abundances are contaminated by naturally occurring isotopes and tracer isotopic impurity, and you have access to unlabeled sample reference measurements to empirically model these confounding.