HolobiomicsLab
- 7.4k skills
- 0 followers
- 1 day ago last updated
- ▌ Metabolomics Data Format Validation · holobiomicslabUse when you have a tab-delimited metabolomics data file (raw measurement output from xcms, Sciex OS, or similar acquisition pipelines) and need to load it into mzQuality before building a SummarizedExperiment.
- ▌ Metabolomics Data Output Formatting · holobiomicslabUse when after completing feature annotation with the annotateRC function on LC–MS All-ion fragmentation (AIF) datasets, when you need to persist ranked metabolite candidates, matched ion spectra, and global summary tables to disk for archival, manual review, or integration into downstream.
- ▌ Mgf Metadata Completion From Smiles · holobiomicslabUse when when processing MGF-format MS2 spectral libraries (e.g., GNPS) that contain SMILES but lack the Molecular Formula field, and you need to prepare the library for MS-DIAL import or polarity-based separation workflows.
- ▌ Mispicked Ion Detection And Merging · holobiomicslabUse when immediately after importing raw LC-MS peak tables (e.g., Progenesis format) and before applying group or replicability filters.
- ▌ Model Validation Regression Metrics · holobiomicslabUse when after training a deep-learning model on paired MS/MS spectra with annotated structural similarity labels, use this skill to assess whether predicted similarity scores correlate with ground-truth reference similarities on data the model has never seen.
- ▌ Molecular Family Graph Construction · holobiomicslabUse when you have downloaded and extracted a GNPS archive (from METABOLOMICS-SNETS, METABOLOMICS-SNETS-V2, FEATURE-BASED-MOLECULAR-NETWORKING for GNPS1, or classical_networking_workflow/feature_based_molecular_networking_workflow for GNPS2) and need to construct a queryable molecular family graph.
- ▌ Molecular Property Data Structuring · holobiomicslabUse when you have a CSV file with rows of molecule definitions (chemical formulas, retention times, intensities, or other peak properties) and need to feed them into SMITER's simulation workflow.
- ▌ Ms2 Spectral Similarity Computation · holobiomicslabUse when you have two or more LC-MS/MS datasets (in mzML, mzXML, or MGF format) and need to quantify their overall spectral content similarity—typical scenarios include data quality control, species identification, molecular phylogenetics, or proteome comparison across samples, cell lines, or.
- ▌ Msms Spectrum Metadata Preservation · holobiomicslabUse when removing invalid or malformed entries (e.g., SMILES validation, format errors) from large spectral datasets (GNPS, MoNA, MTBLS1572, MassBank).
- ▌ Multi Criterion Scoring Integration · holobiomicslabUse when you have an LC-HRMS feature table (m/z, retention time, isotope ratios, fragmentation patterns) and a suspect compound database with reference properties (m/z, expected RT windows, theoretical isotope ratios, characteristic neutral losses), and you need to rank which features most likely.
- ▌ Multidimensional Feature Annotation · holobiomicslabUse when you have a peak-picked feature table (HDF5 format) from high-dimensional MS data (m/z, drift_time, retention_time, intensity) and need to identify and label isotopic signatures to distinguish monoisotopic peaks from isotopologues, reduce feature redundancy, and support multi-dimensional.
- ▌ Multidimensional Ms Data Conversion · holobiomicslabUse when you have acquired untargeted MS data with orthogonal separations (LC, ion mobility) and/or data-independent acquisition (DIA) from Thermo, Agilent, or Bruker instruments, and you need to: (1) store multidimensional spectra in a vendor-neutral, platform-agnostic format;
- ▌ Nearest Neighbor Index Construction · holobiomicslabUse when when you have millions of high-resolution MS/MS spectra converted to low-dimensional vectors (via feature hashing) and need to compute a sparse pairwise distance matrix for downstream density-based clustering.
- ▌ Network Based Functional Prediction · holobiomicslabUse when you have an untargeted metabolomics feature table with m/z values, retention times, intensity measurements, and p-values from statistical testing, but lack or wish to bypass metabolite identification.
- ▌ Nextflow Pipeline Configuration Hpc · holobiomicslabUse when your analysis target is an HPC environment (e.g., Slurm-managed cluster, university research computing center) where Singularity is available but Docker is restricted or unavailable; your workflow is already packaged in Docker but needs portability to HPC;
- ▌ Paired Omics Data Structure Mapping · holobiomicslabUse when when you have a paired omics project document (JSON format) that combines MS/MS mass spectrometry data with genome identifiers, biosynthetic gene cluster information, sample preparation, extraction method, and instrumentation metadata, and you need to verify it conforms to the platform's.
- ▌ Peak Area And Intensity Measurement · holobiomicslabUse when when you have vendor-format GC-CI-MS or LC-MS data from stable isotope labeling experiments, a target list of compounds with known monoisotopic m/z, retention time, and elemental formula, and you need per-isotopologue area and intensity values for quantification or downstream statistical.
- ▌ Per Sample Pass Fail Classification · holobiomicslabUse when after LC-MS data acquisition is complete (or during real-time monitoring) and you have loaded processed LC-MS data in mzML or vendor format and defined QC criteria (retention time windows, m/z tolerances, intensity thresholds) for your internal standards and target analytes.
- ▌ Permutation Test P Value Estimation · holobiomicslabUse when after computing Multi-Block Variable Importance in Projection (MB-VIP) scores on a fitted MB-PLS discriminant model, when you need to distinguish signal features from noise by establishing empirical significance thresholds rather than relying on parametric assumptions.
- ▌ Publication Image Asset Preparation · holobiomicslabUse when you have resolved a USI string pointing to a spectrum in a supported repository (GNPS, MassBank, MassIVE, MetaboLights, Metabolomics Workbench, ProteoXchange, or MS2LDA) and need to embed a publication-ready visualization or machine-readable reference that preserves the spectrum's.
- ▌ Retention Time Alignment Clustering · holobiomicslabUse when you have two or more nontargeted LCMS feature tables from the same analytical method and need to establish matched feature correspondences across datasets, then consolidate redundant features.
- ▌ Retention Time Alignment Correction · holobiomicslabUse when when processing a batch of centroided mzML or mzXML LC-MS raw data files where chromatographic retention times drift between sample acquisitions (common in large-scale metabolomics studies).
- ▌ Retention Time Alignment Validation · holobiomicslabUse when after sample alignment step in untargeted LC-MS workflows, particularly when processing multi-sample cohorts with QC samples interspersed throughout the sequence.
- ▌ Retention Time And Mz Weight Tuning · holobiomicslabUse when after anchor selection and RT mapping spline construction, when you have a fitted metabCombiner object with pre-aligned feature pair candidates and need to tune the scoring metric that combines retention time, m/z, and cosine similarity components.
- ▌ Retention Time Clustering Alignment · holobiomicslabUse when you have extracted feature tables (via MS1 peak picking, MS2 recognition, or targeted list extraction) from two or more individual LC-MS samples and need to identify which features represent the same metabolite across samples before generating a unified, sample-aligned feature table for.
- ▌ Retention Time Co Elution Detection · holobiomicslabUse when after matching mass-to-charge ratios to a compound database (e.g., KEGG) and assigning adduct/fragment types, when you have an annotated feature table with retention times, m/z values, and intensity profiles across samples.
- ▌ Retention Time Intensity Tabulation · holobiomicslabUse when when you have a resolved mzML or mzXML spectrum file and need to visualize or analyze the temporal intensity profile of a specific analyte (defined by its m/z value).
- ▌ Retention Time Mz Intensity Mapping · holobiomicslabUse when you have processed LC-MS run data (feature table or peak detection output) containing internal standard identifications with retention times, m/z values, and intensity measurements across multiple samples, and you need to rapidly detect instrumental drift, retention time shifts, or.
- ▌ Root Mean Squared Error Calculation · holobiomicslabUse when when you have paired predictions and ground-truth structural similarity labels (e.g., predicted Tanimoto scores from a neural network and reference Tanimoto scores from RDKit Daylight fingerprints) and need to report a single scalar metric of model prediction error across all pairs.
- ▌ S4 Class Definition And Inheritance · holobiomicslabUse when you are designing a new backend or data container that must integrate seamlessly with an existing Spectra-based workflow. You have identified a virtual parent class (e.
- ▌ Sample Metadata Annotation For Lcms · holobiomicslabUse when when you have loaded centroided .mzML files into a Spectra object and plan to use TARDIS (tardisPeaks) with an MsExperiment object rather than file paths, and you need TARDIS to distinguish QC runs from sample runs for separate quality metric calculation, polarity filtering, and.
- ▌ Set Activity Prioritization Ranking · holobiomicslabUse when after computing activity scores for a collection of metabolite sets (pathways, GNPS Molecular Families, or MS2LDA Mass2Motifs) from intensity and annotation data.
- ▌ Similarity Threshold Interpretation · holobiomicslabUse when when you have computed Spec2Vec similarity scores (typically cosine similarity in [0, 1] range) between discovered Mass2Motifs and a spectral library, and need to decide which matches are sufficiently confident to include in per-motif annotation output.
- ▌ Singly Charged Ion Mass Calculation · holobiomicslabUse when you have detected monoisotopic features (m/z, drift_time, retention_time, intensity) from LC-IMS-MS/MS data and need to identify and cluster their C13 isotopologues for charge state z=+1.
- ▌ Singularity Container Backend Setup · holobiomicslabUse when your LC-HRMS metabolomics analysis must run on a high-performance computing cluster (e.g., HiPerGator, SLURM-managed systems) that lacks Docker support or prefers Singularity for security and portability. You have .mzML or .
- ▌ Skyline Import Format Specification · holobiomicslabUse when you have computationally generated precursor m/z values, fragment m/z values, collision energies, and retention time predictions for a set of lipid targets (e.
- ▌ Sparse Distance Matrix Construction · holobiomicslabUse when you have a large collection of MS/MS spectra (hundreds of thousands to millions) that need to be clustered, you have already constructed nearest neighbor indexes on low-dimensional spectrum vectors (via feature hashing), and you need to compute only the relevant pairwise distances between.
- ▌ Spectral Entropy Quality Assessment · holobiomicslabUse when after feature detection and alignment in untargeted MS data processing, when you need to filter or rank candidate metabolite annotations by spectral quality before committing to xenobiotic metabolite assignments. Use when combining fragmentation similarity scores (e.
- ▌ Spectral Feature Vector Aggregation · holobiomicslabUse when you have generated per-sample MS2 fingerprints (as spec2vec document representations counting MS2 peaks and neutral losses to precursor in each sample) and need to align them into a single matrix for downstream cross-sample comparison, filtering, or visualization (e.
- ▌ Spectral Library Record Structuring · holobiomicslabUse when you have extracted MS1 and MS2 scans (in mzML/mzXML format) from raw chromatogram files and possess user-provided metadata (retention time, m/z, compound name, molecular weight, annotation fields) that you need to bind together into a queryable spectral library record for local compound.
- ▌ Spectral Match Result Consolidation · holobiomicslabUse when you have executed batch searches of MS/MS spectra against multiple domain-specific MASST indices and need to integrate the resulting match outputs into a single coherent view.
- ▌ Spectral Peak Binning Preprocessing · holobiomicslabUse when when you have raw high-resolution tandem mass spectra (mzML, mzXML, or MGF format) that you intend to cluster or compare at scale, and you need to convert continuous m/z and intensity measurements into discrete bins suitable for feature hashing or similarity searching.
- ▌ Spectral Peak Intensity Aggregation · holobiomicslabUse when you have multiple replicate MS/MS spectra for the same metabolic feature (e.g., 66 top-TIC spectra for feature 1982) and need to identify robust peaks by merging nearby m/z values and pooling their signal strength.
- ▌ Spectral Quality Control Assessment · holobiomicslabUse when after MS1 extraction (coarse/fine error correction, EIC window extraction) and retention time windowing on a set of mzML files tagged with ionization mode and compound adducts.
- ▌ Spectral Scan Generation And Export · holobiomicslabUse when you have real LC-MS/MS data (mzML) from a complex sample (e.g., beer, metabolomics extract) and need to prototype or validate a new data-dependent acquisition (DDA) strategy—such as Top-N fragmentation—before deploying it on physical instrumentation.
- ▌ Spectral Vectorization For Indexing · holobiomicslabUse when when you have a collection of MS/MS reference spectra and unknown query spectra that must be rapidly matched against a large spectral library, and you need to enable approximate nearest neighbor indexing to reduce computational cost from exhaustive pairwise comparison to K-nearest neighbor.
- ▌ Spectrum Pair Retrieval And Ranking · holobiomicslabUse when you have a test set of annotated MS/MS spectra with known structural similarity (via molecular fingerprints or InChIKey), and you want to assess how well a spectral similarity measure (learned or classical) retrieves structurally related compound pairs across a full range of thresholds.
- ▌ Substructure Annotation Integration · holobiomicslabUse when you have (1) a GNPS molecular network (classical or feature-based) with cluster/feature identifiers, (2) MS2LDA output containing Mass2Motif assignments with probability and overlap scores for those same clusters/features, and (3) a goal to annotate network nodes with substructural and.
- ▌ Table Join Alignment On Identifiers · holobiomicslabUse when you have a feature quantification table output from MZmine3 feature detection (rows = features, columns = sample abundance values) and a separate sample metadata file (rows = samples, columns = sample attributes) that need to be unified before downstream statistical analysis, data cleanup.
- ▌ Tandem Mass Spectra Standardization · holobiomicslabUse when you have acquired raw tandem MS data from ProteomeXchange or vendor instruments in proprietary formats (e.g., Thermo .raw files) and need to perform comparative clustering benchmarks or quality assessments across multiple clustering tools.
- ▌ Targeted Transition List Generation · holobiomicslabUse when you have a set of lipid targets defined by species name, acyl chain composition, and expected adducts, and you need to configure a targeted mass spectrometry workflow (PRM or MRM) in Skyline.
- ▌ Technical Replicate Signal Modeling · holobiomicslabUse when your LCMS metabolomics dataset exhibits run-order-dependent intensity drift (signal decay or gain over the course of a sample batch), you have pooled technical replicates (identical biospecimen injected multiple times across the run sequence) and/or known internal standard compounds, and.
- ▌ Theoretical Fragment Ion Generation · holobiomicslabUse when when you have a peptide sequence and need to predict its fragment ion spectrum for stable isotope labeling validation, particularly when comparing against observed mass spectrometry data with natural or enriched isotopic abundance (e.g., 1% or 50% 13C incorporation).
- ▌ Throughput Optimization Ms Querying · holobiomicslabUse when you have experimental MS/MS spectra that must be matched against a large hierarchical fragmentation library (e.g., 168.
- ▌ Time Series Intensity Normalization · holobiomicslabUse when when LCMS metabolomics abundance tables show systematic intensity drift across the injection sequence (e.g., instrument signal decay or gain over hours), and you have pooled technical replicate injections and/or internal standard compounds distributed throughout the run.
- ▌ Tracer Impurity Correction Modeling · holobiomicslabUse when when processing LC-MS data from isotope labeling experiments where the tracer (13C, 2H, 15N, 18O, or 34S) has known isotopic impurity and you observe discrepancies between measured isotopologue abundances (FAM) and expected labeling patterns.
- ▌ Transition Group Feature Extraction · holobiomicslabUse when when you have loaded a TransitionGroup (extracted ion chromatogram or mobilogram from DIA mass spectrometry data) and need to identify precise peak boundaries, apex retention/drift time, and intensity values for quantitative feature detection.
- ▌ Untargeted Lc Ms Data Preprocessing · holobiomicslabUse when when you have raw untargeted LC-MS metabolomics data and need to detect low-quality or mis-integrated peaks in an XCMS-processed xcmsSet object before performing metabolite annotation, statistical analysis, or biomarker discovery.
- ▌ Virtual Chemical Mixture Generation · holobiomicslabUse when you need to create a synthetic chemical population for testing data-dependent acquisition (DDA) strategies in a simulation environment before committing to real mass spectrometry analysis.
- ▌ Workflow Architecture Documentation · holobiomicslabUse when you need to understand how a complex MS/MS spectral search system routes query spectra through multiple parallel processing pipelines with different objectives (e.g., reliable exact matching vs. fast approximate matching).
- ▌ Workflow Reproducibility Validation · holobiomicslabUse when after implementing or deploying a containerized Nextflow workflow that processes LC-HRMS metabolomics .
- ▌ Xcms Grouping Result Interpretation · holobiomicslabUse when xCMS grouping has been performed on LC-MS data from studies with hundreds of samples or data acquisition periods longer than a week, where retention time drift structures are complex and the single-warping-function assumption is likely violated.
- ▌ Browser Security Policy Configuration · holobiomicslabUse when you need to run a web application locally (by opening index.html directly in the browser) and the application uses WebWorker or WebAssembly modules that fail to load with cross-origin policy or file-access errors.
- ▌ Cumulative Distribution Visualization · holobiomicslabUse when when you have per-feature quality metrics (such as CV values from NMR or MS reproducibility analysis) and need to: (1) confirm that a specified proportion of features meet regulatory thresholds (e.g., 99% < 0.30, 92% < 0.15 for CV);
- ▌ Deep Learning Spectral Language Model · holobiomicslabUse when when you have an unknown compound's mass spectrum (m/z peaks and intensities in .mgf or equivalent format with mandatory PRECURSOR_MZ and IONMODE tags) and need to identify structurally related metabolites from a reference database.
- ▌ Memory Mapped File Selection Strategy · holobiomicslabUse when initializing a Dataset object in NMRFx and must decide which storage backend to use for in-memory or memory-mapped file access.
- ▌ Peak Network To Metabolite Assignment · holobiomicslabUse when after peak clustering has produced peak network groups (ideally from the same compound) and you need to assign chemical identities.
- ▌ Repository Cloning And Initialization · holobiomicslabUse when you have a GitHub repository URL, a documented Python version requirement, and a list of pinned package versions, and you need to verify that the application will initialize without import or runtime errors before proceeding to data analysis or method replication.
- ▌ Spectral Similarity Scoring Parent Tp · holobiomicslabUse when after generating TP candidates (via in-silico prediction or library lookup) and extracting MS/MS peak lists for both parent features and TP feature candidates, use spectral similarity scoring to quantify fragmentation pattern overlap.
- ▌ Summarized Experiment Object Handling · holobiomicslabUse when you have cross-validated, filtered metabolomic NMR or MS data in a SummarizedExperiment container and need to prepare it for metabolome-wide association studies (MWAS) with epidemiological confounders.
- ▌ Top K Accuracy Ranking And Evaluation · holobiomicslabUse when a machine learning model produces multiple ranked predictions (each with an associated confidence score) for a single input, and you need to quantify how often the correct answer appears in the top-k predictions.
- ▌ Transformer Cnn Hybrid Model Training · holobiomicslabUse when you have preprocessed 1H NMR spectral data with compound labels and need to identify multiple compounds in a flavor mixture where both local spectral patterns (handled by CNN) and long-range spectral dependencies (handled by Transformer) are diagnostic.
- ▌ Anndata Backed Object Manipulation · holobiomicslabUse when when working with large single-cell ATAC-seq or multi-omics datasets where in-memory storage is infeasible (>1M cells), and you need to iteratively add or modify count matrices (tile-based, peak-based, or gene-based) while preserving fragment-level data for reproducibility and re-analysis.
- ▌ Differential Tf Occupancy Analysis · holobiomicslabUse when you have aligned ATAC-seq BAM files and peak annotations from two or more experimental conditions (e.
- ▌ Genomic Feature Annotation Overlap · holobiomicslabUse when after you have identified a set of differentially methylated bases or regions (e.
- ▌ Hi C Fastq To Contact Map Pipeline · holobiomicslabUse when you have raw Hi-C FASTQ files from a sequencing experiment and need to generate kilobase-resolution Hi-C contact maps conforming to ENCODE reference standards.
- ▌ Hic Contact Map Feature Annotation · holobiomicslabUse when you have completed Hi-C map generation (producing .hic files from aligned reads) and need to detect and annotate topological features such as chromatin loops, topologically associating domains (TADs), or interaction peaks.
- ▌ Illumina 450k Epic Dataset Loading · holobiomicslabUse when you have raw .idat files or a beta-valued matrix from an Illumina HumanMethylation450 or EPIC array experiment and need to import the full probe set into R for downstream quality control, normalization, and differential methylation analysis.
- ▌ Methylation Batch Effect Detection · holobiomicslabUse when you have loaded a normalized beta-valued methylation matrix (e.
- ▌ Methylation Probe Count Validation · holobiomicslabUse when immediately after loading raw methylation array data using champ.load() or champ.import() to verify data integrity.
- ▌ Module Import And API Verification · holobiomicslabUse when before invoking any Python module in a multi-step Hi-C processing pipeline, or when a dependency has been freshly installed or reinstalled.
- ▌ Paired Insertion Counting Strategy · holobiomicslabUse when you have loaded fragment data from single-cell ATAC-seq experiments into a backed AnnData object (with fragments stored in .obsm['fragment_paired'] or .
- ▌ Peak Calling Output Interpretation · holobiomicslabUse when you have run a peak-calling algorithm on sparse CUT&RUN bedGraph data and received a BED-format output file;
- ▌ System Dependency Version Checking · holobiomicslabUse when you are preparing to run a complex multi-tool bioinformatics pipeline (such as HiC-Pro) on a new system or cluster, and need to confirm that all required binaries exist in the execution environment and meet minimum version thresholds (e.g., samtools >=1.9, Python >3.
- ▌ Assembly Version String Retrieval · holobiomicslabUse when when you need to validate that a .NET assembly (such as ThermoFisher.CommonCore.RawFileReader) is correctly installed and accessible before attempting data reading operations.
- ▌ Atom Feature Extraction Chemistry · holobiomicslabUse when you have canonicalized SMILES strings from a chemical database (e.
- ▌ Bayesian Meta Learning Projection · holobiomicslabUse when you need to transfer retention time predictions from one chromatographic method to another, but have access to only a small number (≥10) of molecules with known retention times in both methods.
- ▌ C Sharp Wrapper Method Validation · holobiomicslabUse when when integrating an R package that wraps a compiled .NET assembly (such as rawrr), you need to verify that the internal dispatch mechanism between the R layer and the C# layer is operational before attempting to read actual raw data files.
- ▌ Calibration Curve Fitting Ms Data · holobiomicslabUse when you have raw mass spectrometry intensity data from targeted analytes and a set of calibration standard measurements with known concentrations.
- ▌ Chemodiversity Metric Calculation · holobiomicslabUse when when you have sum-normalized peak-abundance matrices from FT-ICR MS data with assigned molecular formulas and need to compare metabolite diversity between treatment groups (e.g., inoculated vs. control samples).
- ▌ Conditional Dependency Resolution · holobiomicslabUse when a Python library exposes functionality that depends on external packages (like sqlalchemy, pandas, or lxml) that are not required for core operations.
- ▌ Container Image Size Verification · holobiomicslabUse when after completing a multi-stage Docker build targeting a compiled runtime environment (e.g., airdpro:cli produced from a Wine + .NET Framework 4.8 + Ubuntu 22.
- ▌ Count Verification And Validation · holobiomicslabUse when you have grouped unique 2D chemical structures by organism prevalence and need to confirm that the counts in each frequency bin (singleton, low-diversity, medium-diversity, high-diversity) match published or curated reference values.
- ▌ Cross Domain Masst Output Parsing · holobiomicslabUse when you have executed batch searches against one or more domain-specific MASST tools and received multiple output files (_microbe.html, _plant.html, _tissue.html, _microbiome.html, _food.html, _matches.tsv, _library.tsv, _datasets.tsv, _count_*.
- ▌ Custom Metabolite Set Integration · holobiomicslabUse when you have a user-supplied metabolite set file (CSV or JSON) defining custom groupings of metabolites (e.
- ▌ Dataframe Lazy Loading Comparison · holobiomicslabUse when when designing or optimizing an MsBackend implementation (or similar columnar data structure) you must decide whether to pre-allocate all known columns in the backing DataFrame at initialization or defer column creation until first access.
- ▌ Docker Multistage Build Execution · holobiomicslabUse when you need to containerize a C#-based Windows application (like AirdPro CLI) for Linux deployment, require Wine and .NET Framework 4.
- ▌ Dotnet Assembly Path Verification · holobiomicslabUse when when you have just loaded the rawrr R package and need to confirm that the bundled .NET 8.0 assembly (rawrr.exe) is present and functional before performing any mass spectrometry data extraction operations.
- ▌ Environment Variable Provisioning · holobiomicslabUse when when extending a multi-service project (like MAGMa with its four subproject components) to container orchestration, and you need to ensure each microservice (magmaweb, joblauncher, job, pubchem) receives the correct configuration—such as port mappings, service URLs, and data paths—without.
- ▌ File Format Compliance Validation · holobiomicslabUse when you have generated or received mzPeak files from a Rust, Python, R, or other implementation and need to verify they comply with the published HUPO-PSI specification before integration into a production workflow, data repository, or downstream analysis pipeline.