HolobiomicsLab
- 7.4k skills
- 0 followers
- 1 day ago last updated
- ▌ Precursor Mz Window Filtering · holobiomicslabUse when when preparing augmented training data for Siamese or contrastive learning architectures in mass spectrometry, specifically when you need to generate hard negative examples that are spectrally distinct but mass-similar to positive examples.
- ▌ Quantitation Table Generation · holobiomicslabUse when you have raw MS data in a supported instrument format (Agilent .d, Thermo .raw, Bruker .d, mzML) and a defined list of m/z, retention time, or other identifiers for which you need to extract and quantify peak abundances across one or more samples.
- ▌ R Data Structure Manipulation · holobiomicslabUse when when xcms has produced misaligned feature groups and you need to extract raw LC-MS profiles from source files into a structured format acceptable by ncGTW's realignment functions.
- ▌ Raw Ms Data Format Conversion · holobiomicslabUse when you have raw UPLC-HRMS data from ThermoFisher or Agilent instruments and need to feed it into MSThunder for nontargeted pollutant identification. Your input is a vendor binary format (.raw or .d) that MSThunder cannot directly ingest. Environment constraints (e.
- ▌ Rawrr Spectral Data Retrieval · holobiomicslabUse when you have Thermo Orbitrap .raw files and need to access raw spectral data (individual MS1 or MS2 scans, base-peak values, chromatogram traces, retention times, or scan-level metadata) for custom analysis, visualization, or integration into an R-based pipeline.
- ▌ Rdkit Molecular Featurization · holobiomicslabUse when you have molecular structure data (SMILES strings or MOL files) from an in-house chemical database and need to prepare it as input for a graph neural network that predicts liquid chromatography retention times for small molecule identification.
- ▌ Replicate Spectrum Comparison · holobiomicslabUse when you have multiple MS/MS spectra (replicates) for a single metabolic feature (same m/z and RT window) and need to identify which fragments are reproducibly detected across replicates versus noise.
- ▌ Retention Time Mass Alignment · holobiomicslabUse when you have two independent LC-MS untargeted metabolomic feature datasets (each with retention time and m/z values) and need to identify which features in one dataset correspond to features in the other.
- ▌ Run Order Dependency Modeling · holobiomicslabUse when your feature table includes QC (quality control) sample replicates distributed throughout the analytical run sequence, and you observe systematic intensity variation correlated with sample injection order or batch identifier—typical indicators of instrumental drift in untargeted LC-MS.
- ▌ Signal Residual Deconvolution · holobiomicslabUse when analyzing 1D signal arrays (e.g., extracted ion chromatograms, arrival time distributions, or MS1 spectra intensity profiles) where multiple peaks may overlap or where peak shape information (amplitude, position, width) is required beyond simple local-maximum detection.
- ▌ Spectra Object Manipulation R · holobiomicslabUse when when you have extracted and concatenated MS/MS spectra from multiple replicates for a set of metabolomic features (stored in a preprocessed list), and need to apply intensity-based filtering (e.
- ▌ Spectral Data Standardization · holobiomicslabUse when you have completed a GNPS1 (METABOLOMICS-SNETS, METABOLOMICS-SNETS-V2, FEATURE-BASED-MOLECULAR-NETWORKING) or GNPS2 (classical_networking_workflow, feature_based_molecular_networking_workflow) molecular networking job and need to access its output files in a standardized format for.
- ▌ Spectral Embedding Generation · holobiomicslabUse when you have a collection of pre-processed MS/MS spectra (binned, intensity-normalized) and a trained MS2DeepScore base network, and you need to compute structural similarity scores between spectrum pairs or visualize spectra in chemical space via dimensionality reduction (e.g., UMAP).
- ▌ Spectral Match Interpretation · holobiomicslabUse when you have obtained search results from one or more domain-specific MASST web applications (microbeMASST, plantMASST, tissueMASST, microbiomeMASST, foodMASST) for a query mass spectrum and need to consolidate, rank, and visualize those matches to infer the identity and biological source of.
- ▌ Spectral Metadata Integration · holobiomicslabUse when after isotopologue and adduct grouping has been completed and you need to associate MS2 spectra with consolidated feature groups in DDA LC-MS experiments.
- ▌ Spectral Noise Classification · holobiomicslabUse when you have MS/MS spectra contaminated with noise ions and need to improve compound identification accuracy.
- ▌ Spectral Noise Peak Filtering · holobiomicslabUse when when working with raw or partially processed tandem mass spectrometry (MS/MS) spectra that contain low-intensity background noise peaks.
- ▌ Spectrum Embedding Clustering · holobiomicslabUse when after embedding MS/MS spectra into 32-dimensional GLEAMS vectors, when you need to group spectra by the same peptide origin.
- ▌ Spectrum Vector Serialization · holobiomicslabUse when after successfully constructing a nearest neighbor index from hashed spectrum feature vectors and before performing density-based clustering or similarity searches.
- ▌ Structural Similarity Scoring · holobiomicslabUse when when you have a collection of mass spectra with annotated chemical structures (SMILES/InChI) and need to generate structural similarity labels to train or validate a model that predicts molecular similarity from spectral pairs.
- ▌ Structure Selection Filtering · holobiomicslabUse when when you have an unknown metabolite's predicted structural similarity scores (from a deep learning model such as DeepMASS) against all known metabolites in a reference database, and need to identify which known metabolites are most likely structurally related to the unknown to guide.
- ▌ Synthetic Spectrum Generation · holobiomicslabUse when when you need to create benchmark LC-MS/MS datasets with controlled, known composition for testing MS analysis algorithms, validating retention-time or fragmentation predictions, or predicting co-elution and co-fragmentation challenges before conducting real experiments.
- ▌ Tandem Mass Spectrum Decoding · holobiomicslabUse when when you have raw predictions from a trained fragment generation or intensity prediction neural network model and need to convert those predictions into a standard spectrum file format (m/z–intensity pairs) for comparison against experimental spectra or for structural elucidation workflows.
- ▌ Taxonomy String Normalization · holobiomicslabUse when when preparing a metadata table (TSV format with species, genus, and family columns) for natural product metabolomics analysis where the Literature Component score must query known compounds by taxon.
- ▌ Total Ion Current Calculation · holobiomicslabUse when after loading all MS1 scans from a raw or intermediate mass spectrum file (e.
- ▌ Train Test Split Verification · holobiomicslabUse when after applying a configuration fix (e.g., adding an instrument type to an allowlist, updating filtering thresholds) to a dataset preprocessing pipeline, you need to confirm that the change produces the documented training/test split counts.
- ▌ Transfer Learning Fine Tuning · holobiomicslabUse when you have collected liquid chromatography (LC) spectra and retention time labels for your in-house molecular database, and you want to leverage a pretrained GNN-RT model rather than train from scratch.
- ▌ Transition Group Peak Picking · holobiomicslabUse when you have loaded extracted ion chromatogram (XIC) data from DIA mass spectrometry and need to identify peak boundaries for peptide precursor transitions.
- ▌ Bruker Nmr Spectral Data Import · holobiomicslabUse when you have raw Bruker NMR spectral files (from a Bruker instrument) in a directory and need to prepare them for automated metabolite identification and quantification using ASICS.
- ▌ Cnn Spectral Feature Extraction · holobiomicslabUse when you have preprocessed 1D NMR spectra (¹H and/or ¹³C) and need to extract spectral features for molecular structure inference on molecules with up to 19 heavy atoms. The skill is necessary as the first stage before fragment assembly or connectivity prediction;
- ▌ J Coupling Multiplet Generation · holobiomicslabUse when you have parsed metabolite identities with known spin-system coupling constants (J-values) and chemical shifts, and need to generate the theoretical multiplet patterns that will form the basis of a simulated 1D or 2D NMR spectrum.
- ▌ Metabolite Structure Prediction · holobiomicslabUse when you have a parent compound (or set of compounds) in SMILES, MOL, or SDF format and need to predict plausible metabolite structures and pathways in a specific biological context (mammalian Phase I/II metabolism, human gut microbiota, or soil/aquatic microbial degradation).
- ▌ Neural Network Model Deployment · holobiomicslabUse when you have LC-MS feature tables (m/z and retention time columns) and corresponding .mzXML or .mzML files, and you need to automatically classify whether extracted ion chromatograms represent genuine metabolomic features or false positives.
- ▌ Nmr Spectral Feature Extraction · holobiomicslabUse when when you have raw or preprocessed 1H NMR spectral tensors from flavor or chemical mixtures and need to generate high-level feature representations that capture both fine-grained local patterns (e.g., peak multiplet structure, coupling constants) and global spectral context (e.
- ▌ Nmr Workflow Pipeline Execution · holobiomicslabUse when you have raw 1D NMR spectra (FID or processed spectrum files) and need to extract peak parameters (chemical shift, intensity, linewidth) in a tabular format for downstream metabolomic or structural analysis.
- ▌ Restful API Endpoint Invocation · holobiomicslabUse when you have NMR peak data (1H and 13C chemical shift values) and need to obtain SMART 3 classification predictions from the DeepSAT service.
- ▌ Spectral Connectivity Filtering · holobiomicslabUse when you have picked peaks (coordinates and intensities) from an INADEQUATE NMR spectrum and need to cluster them into networks representing individual compounds.
- ▌ Spectral Intensity Thresholding · holobiomicslabUse when you have loaded raw INADEQUATE NMR spectrum data and need to distinguish true molecular peaks from noise and instrumental artifacts.
- ▌ Spectral Modality Preprocessing · holobiomicslabUse when you have downloaded raw spectroscopic data files (NMR, HSQC, COSY, IR modalities) from the Zenodo repositories and need to convert them into the standardized multi-modal input format required by the MultiModalSpectralTransformer before inference or retraining.
- ▌ Structure Similarity Comparison · holobiomicslabUse when after executing a molecular structure prediction model on spectroscopic input data and obtaining predicted molecular structures in a standardized format (e.g., SMILES, MOL, SDF).
- ▌ Structure Similarity Evaluation · holobiomicslabUse when after an NMR-based structure prediction model has generated predicted molecular structures (formula and connectivity) for a test set of molecules with up to 19 heavy atoms.
- ▌ Validation Error Categorization · holobiomicslabUse when when a parsed mwTab file (MS or NMR experimental data) must be assessed for conformance to its corresponding JSON schema specification. Apply this skill after loading the mwTab file using the mwtab parser but before quality assurance sign-off or deposition to the Metabolomics Workbench.
- ▌ Webassembly Local Loading Setup · holobiomicslabUse when you have downloaded a web application (e.g., COLMARvista) that uses WebWorker and WebAssembly components and need to run it locally by opening index.html in a browser, rather than accessing it through a web server.
- ▌ Atac Seq Tn5 Bias Correction · holobiomicslabUse when you have raw ATAC-seq BAM files and need to perform footprinting analysis to detect transcription factor binding through Tn5 insertion patterns.
- ▌ Hi C Map Artifact Validation · holobiomicslabUse when after executing the Juicer pipeline on raw Hi-C FASTQ files, to confirm that the pipeline has generated the expected .hic output artifact and that the contact matrix construction and normalization steps completed without error.
- ▌ Batch Document Verification · holobiomicslabUse when when you have deposited a collection of JSON project documents in a platform and need to verify that all conform to the published schema before public release or after schema updates.
- ▌ Batch Prediction Comparison · holobiomicslabUse when when you have a trained molecular classifier (like BitterPredict) that accepts structured descriptor input, and you need to understand which chemical descriptor subgroups drive prediction outcomes.
- ▌ Chemical Identifier Mapping · holobiomicslabUse when you have .msp spectral library files with compound names but lack standardized chemical identifiers (SMILES, InChI, InChI Key, CAS number, IUPAC names, or molecular formulas).
- ▌ Embedding Vector Validation · holobiomicslabUse when after instantiating and invoking a sinusoidal formula embedding layer (such as SCARF embeddings in MIST-CF) on chemical formula inputs, validate that the output embeddings meet dimensionality and value constraints before using them for downstream transformer or ranking tasks.
- ▌ Error Handling Confirmation · holobiomicslabUse when when implementing or auditing a data replacement method (e.g., `mz<-`, `intensity<-`) in an MsBackend subclass that must enforce ordering or format constraints on peak data.
- ▌ Figure Table Interpretation · holobiomicslabUse when you need to verify claims about algorithm performance, data processing correctness, or workflow outcomes in a scientific article or software repository. Use it when source documents contain figures, tables, or visualization badges (e.
- ▌ Git Repository Tag Checkout · holobiomicslabUse when when you need to reproduce or validate a specific historical release artifact (e.g., a Semantic Release v1.0.
- ▌ Graph Serialization Graphml · holobiomicslabUse when after constructing a network graph where nodes represent Mass2Motifs (or spectra) and edges encode pairwise spectral similarity scores, and you need to export the network for visualization, post-processing, or sharing with collaborators using standard graph software (e.
- ▌ Im Ms Drift Time Correction · holobiomicslabUse when when you have IM-MS lipidomics data acquired on samples spiked with U13C-labeled internal standards (fully labeled yeast extract) and need to quantify systematic CCS bias and apply lipid class-specific bias correction to all measured CCS values, particularly when multiple lipids per lipid.
- ▌ Metadata Field Verification · holobiomicslabUse when you have located a workflow definition file (YAML or JSON) from a versioned release and need to confirm that all mandatory workflow metadata fields (name, version, inputs, outputs, steps) are declared, properly formatted, and cross-references are resolved before validation or execution.
- ▌ Package Integration Testing · holobiomicslabUse when you need to verify that a Python package (or similar installable software) passes its declared integration test suite as a prerequisite to trusting its reliability in production or downstream analysis. Specifically, apply it when you observe a periodic testing CI workflow badge (e.
- ▌ Pathway Coverage Assessment · holobiomicslabUse when when preparing to run Over-representation Analysis (ORA) on a metabolomics study, after constructing the background set but before running the enrichment test.
- ▌ Pytest Test Suite Execution · holobiomicslabUse when after installing a package in development mode (e.g., via `pip install -e .[dev]`) to verify the package functions as intended, or before submitting pull requests to confirm no regressions were introduced by code changes.
- ▌ Python Source Code Analysis · holobiomicslabUse when you have a Python webservice codebase (e.g., a Flask, Django, or FastAPI application) and need to document its HTTP API surface (endpoints, methods, parameters, schemas, authentication) for integration, testing, or specification generation.
- ▌ Reference Library Alignment · holobiomicslabUse when when you have IM-MS lipidomics data with measured CCS values from samples spiked with U13C labeled internal standards, and you need to assess systematic CCS bias or enable CCS correction by comparing measured lipids against known library entries with validated CCS values.
- ▌ Release Artifact Generation · holobiomicslabUse when when a software project has reached a stable milestone (v-tagged commit) and you need to produce official distribution artifacts with verified version metadata, checksums, and release documentation that can be validated against a published GitHub release record.
- ▌ Repository History Analysis · holobiomicslabUse when when you need to understand how a complex feature or architectural pattern was implemented in a codebase, particularly when the current README or documentation does not fully explain the control flow, decision criteria, or parameter passing between subsystems.
- ▌ Rust Build System Execution · holobiomicslabUse when you have obtained a Rust source repository (e.g., mzpeak_prototyping) and need to compile it into a working command-line converter tool or library. Use this skill when the source includes a Cargo.
- ▌ Scientific Task Formulation · holobiomicslabUse when you are starting a new mass spectrometry analysis task or feature request where the problem scope is unclear, the tool chain (e.g., OpenMS + Python + KNIME integration) is not yet selected, or acceptance criteria (code quality, test coverage, documentation) have not been established.
- ▌ Software Plugin Development · holobiomicslabUse when when you have a scientific software tool (e.g., Met-ID) that is architected to support plugins or configuration-driven modules, and you need to register and apply a novel reagent, derivatizing matrix, or analytical method (e.
- ▌ Sparsity Metric Computation · holobiomicslabUse when when you have loaded a dataset of molecular fingerprint vectors (such as biosynfoni fingerprints from a Zenodo deposit) and need to quantify how sparse the bit-representations are—that is, what fraction of bit positions are zero across the fingerprint collection.
- ▌ Tima Entry Point Validation · holobiomicslabUse when when you have obtained or are considering use of the tima Docker image (adafede/tima-r) and need to confirm that the containerized environment is operational before proceeding with metabolite annotation workflows. This is a smoke test to catch environment or registry issues early.
- ▌ Twim Ms Calibration Mapping · holobiomicslabUse when you have raw or processed arrival-time data from a TWIM-MS instrument and need to convert it to CCS values for comparison across experiments or biomolecular classes.
- ▌ Abstract Syntax Tree Design · holobiomicslabUse when when you have a domain-specific query language (such as MassQL for mass spectrometry) and need to convert query strings into structured representations that preserve domain constraints (e.g., mass tolerance, scan type, intensity thresholds).
- ▌ Abundance Matrix Processing · holobiomicslabUse when you have multiple CSV files containing feature-by-sample matrices from different analytical experiments or batches, each with mass, retention time, intensity, isotope, and adduct information across different samples.
- ▌ Adduct Ion Mass Calculation · holobiomicslabUse when when you have derivatized metabolite structures (SMILES or mol format) and need to predict their ionization products in MS imaging, particularly when the derivatizing matrix produces non-standard adducts (e.
- ▌ Anova Pvalue Adjustment Fdr · holobiomicslabUse when you have completed ANOVA or G-test statistical testing across multiple biomolecules in an omics dataset and need to account for multiple testing.
- ▌ API Response Error Handling · holobiomicslabUse when when building asynchronous metadata enrichment workflows that call multiple external APIs (CIR, CTS, PubChem, IDSM, BridgeDb) to annotate mass spectra .
- ▌ Byte Offset Seek Operations · holobiomicslabUse when when you have an indexed gzip-compressed mzML file and need to retrieve individual spectra or chromatograms by index without sequential file reading or full decompression. Typical scenario: you want spectrum[42] from a 10 GB indexed mzML.gz file and need sub-second access time.
- ▌ C Python Interface Wrapping · holobiomicslabUse when you have a mature C++ library (like OpenMS) with stable APIs that you want to make accessible from Python environments, and you need to preserve performance-critical C++ execution while supporting rapid prototyping or integration into Python-based data pipelines (e.
- ▌ Candidate Structure Ranking · holobiomicslabUse when you have (1) a set of candidate structures generated by in silico fragmentation (e.g., MetFrag output), (2) one or more seed nodes with known spectral library matches or identity scores, and (3) a fragmentation relationship graph connecting candidates.
- ▌ Ccs Prediction Model Design · holobiomicslabUse when you have a dataset of molecules with known or reference CCS values, and you need to construct a trainable model that learns the mapping from molecular structure (encoded as SMILES or feature vectors) to scalar CCS predictions.
- ▌ Circular Barplot Generation · holobiomicslabUse when after running a comprehensive preprocessing workflow comparison (e.g., via normulticlassqcall or nortimecoursenoall) that produces an overall ranking CSV file of candidate workflows.
- ▌ Contrastive Pair Generation · holobiomicslabUse when you have raw ion images from MSI data and need to train a contrastive encoder to learn stable, mode-specific representations.
- ▌ Covariance Matrix Inversion · holobiomicslabUse when after generating a covariance matrix from normalized metabolite abundance data in MetaboAnalyst, and before performing network-level metabolomic inference or visualizing metabolite interaction networks.
- ▌ Data Integrity Verification · holobiomicslabUse when after applying matchms metadata cleaning tools to normalize field values and standardize naming conventions on imported spectra (mzML, mzXML, msp, MGF, or JSON formats).
- ▌ Dictionary List Aggregation · holobiomicslabUse when after applying a matrix directive variant (fields_to_headers, headers, or collate) to extract and transform individual records into dictionaries with field filtering and type coercion applied.
- ▌ Distance Geometry Embedding · holobiomicslabUse when when you have ionized adduct structures (in SMILES or MOL format) from an upstream ionization-state determination step and need to produce multiple low-energy 3D conformations for collision cross section prediction, metabolite annotation, or structure-property modeling.
- ▌ Distance Metric Computation · holobiomicslabUse when after obtaining CNN-predicted embeddings for query mass spectra and having loaded pre-computed embeddings from a reference database, you need to identify which reference molecules are most similar to each query.
- ▌ Dna Adduct Characterization · holobiomicslabUse when when you have a collection of DNA adduct compound structures in SDF format that requires validation for structural integrity and completeness, and you need to generate predicted fragment spectra at defined ionization levels and mass ranges for comparison against experimental mass.
- ▌ Docker Container Deployment · holobiomicslabUse when you need to launch a pre-built Docker image of a scientific tool (e.
- ▌ Domain Composition Analysis · holobiomicslabUse when you have a set of Biosynthetic Gene Clusters in GenBank format and need to understand the domain architecture of constituent genes, either to detect statistically significant domain co-occurrence patterns within BGCs, to filter redundant BGCs by domain similarity, or to link detected.
- ▌ Embedding Vector Generation · holobiomicslabUse when when you have tokenized mass spectra (peak-mass and peak-intensity pairs from experimental or in-silico libraries such as NIST 2017 or MassBank) and need to perform rapid similarity searches or spectrum matching at scale.
- ▌ Entry Status Classification · holobiomicslabUse when you need to assess the overall curation coverage and quality of a MIBiG repository snapshot, identify which entries require further review, or track how entry validation status changes over time. Specifically use it when the `data` directory contains JSON files with `cluster.
- ▌ File Format Standardization · holobiomicslabUse when when you have raw outputs from LipidSearch or LIQUID identification software (CSV or TSV format with vendor-specific column naming and lipid identifiers) and need to construct a structured data matrix suitable for batch normalization, statistical testing, or visualization.
- ▌ Filter Parameter Validation · holobiomicslabUse when reconstructing or modifying dashboard filters (e.
- ▌ Format Conversion Chemistry · holobiomicslabUse when when you have retrieved a user database entry (sequence or structure data) from the MassSpecBlocks backend and need to enable mass spectra analysis in the open-source CycloBranch program, or when you need to export NRP sequences and building-block annotations for consumption by external.
- ▌ Ftms Transient Data Loading · holobiomicslabUse when you have a Bruker Solarix FT-ICR transient file in .d format (e.g., ESI_NEG_SRFA.d containing ser and fid files in CompassXtract format) and need to programmatically load it into a Python environment for signal processing, calibration, and mass spectrum generation.
- ▌ Ggplot2 Pie Chart Rendering · holobiomicslabUse when you have a frequency table (generated by count_fold_changes and converted to relative abundance via ra_table) of metabolite subsets stratified by class or direction of change (e.g., increased vs. decreased organic acids at padj ≤ 0.
- ▌ Group Comparison Statistics · holobiomicslabUse when after data integration, batch correction, and sample separation when you have a feature-by-sample abundance matrix (finalData) and corresponding sample group labels (finalLabel), and your research goal is to identify which metabolites significantly differ between experimental groups for.
- ▌ Hpc Job Array Orchestration · holobiomicslabUse when you have a machine learning training workflow (e.g., k-fold cross-validation) where each fold is independent, GPU-accelerated, and can run in parallel.
- ▌ Imms Data Format Conversion · holobiomicslabUse when when you have raw Agilent MassHunter (.d) or UIMF IM-MS data files from drift tube (DT) or structure for lossless ion manipulations (SLIM) instruments and need to ingest them into a preprocessing pipeline that requires standardized in-memory or intermediate representations for.
- ▌ Interactive Plot Generation · holobiomicslabUse when you have loaded m/z and intensity arrays from an MZA file (via mzapy) and need to inspect a mass spectrum, extracted ion chromatogram (XIC), or arrival time distribution (ATD) visually, either for QC purposes, method development, or publication-ready output.
- ▌ Ionization State Prediction · holobiomicslabUse when when you have SMILES strings representing neutral organic molecules and need to enumerate the likely protonated (e.g., [M+H]+) and deprotonated (e.
- ▌ Kaleido Backend Integration · holobiomicslabUse when when a Shiny application currently uses orca for static plot export but requires a lighter-weight, Python-native alternative that avoids Node.js/Electron runtime overhead.