HolobiomicsLab
- 7.4k skills
- 0 followers
- 2 days ago last updated
- ▌ Lipid Fragmentation Pattern Recognition · holobiomicslabUse when you have experimental tandem MS (MS/MS) spectra from lipid samples (in mzML format) and need to assign molecular identities and lipid classes.
- ▌ Mass Action Kinetics Propensity Scoring · holobiomicslabUse when you have intracellular metabolomics measurements (absolute metabolite abundances) for multiple biological samples and want to identify which metabolic reactions are controlled by substrate availability rather than gene expression.
- ▌ Mass Spectrometry Data Table Formatting · holobiomicslabUse when when you have raw or processed TWIM-MS data (arrival time and m/z values) from a mass spectrometry instrument and need to organize it into a feature table before biomolecular class assignment or CCS calculations.
- ▌ Mass Spectrometry Feature Deconvolution · holobiomicslabUse when you have a peak table from LC-MS peak picking software (e.
- ▌ Mass Spectrometry Feature Deduplication · holobiomicslabUse when immediately after MZmine feature detection when you have both MGF (MS/MS spectra) and CSV (metadata) output files for one or both ionization modes and wish to construct a deduplicated molecular network.
- ▌ Mass Spectrometry Structural Annotation · holobiomicslabUse when after identifying statistically significant features (e.
- ▌ Mass Spectrometry Tolerance Calibration · holobiomicslabUse when after generating a feature table from mzML data (via Asari) and before performing MS1 or MS2 annotation.
- ▌ Mass Spectrum Binning And Vectorization · holobiomicslabUse when when preparing MS/MS spectra for neural network training or inference, particularly when you need to feed variable-length spectra into a Siamese network or embedding model that requires fixed-dimensional input.
- ▌ Mass Spectrum Peak Manipulation Merging · holobiomicslabUse when after generating electronic noise (uniformly sampled m/z with Poisson-distributed intensities) and chemical noise (formula database-sampled m/z with Poisson intensities) and you need to combine both noise types with a clean baseline spectrum into a single unified peak array.
- ▌ Metabolite Feature Table Interpretation · holobiomicslabUse when immediately after executing the MetaboAnalystR 4.0 unified LC-MS workflow (feature detection and quantification module) on raw mzML or netCDF data.
- ▌ Metabolite Identifier Mapping To Lipids · holobiomicslabUse when you have (1) peak-picked LC-MS AIF features in a feature table with m/z and retention time, (2) corresponding xcmsSet and RAMClustR pseudo-MS/MS spectral objects from centroid-mode raw data, and (3) a research goal to identify which features are lipids rather than other metabolite classes.
- ▌ Metabolite Protein Network Construction · holobiomicslabUse when after generating metabolite-disease correlation data and protein association predictions from a deep learning metabolomics module (e.g., DeepMSProfiler's feature extraction step).
- ▌ Metabolomic Data Format Standardization · holobiomicslabUse when you have raw peak table data from liquid chromatography–mass spectrometry (LC-MS) or related metabolomic instruments, generated by one of 12 supported software tools (or already in NOREVA's standardized format), and need to prepare it for preprocessing method evaluation or biomarker.
- ▌ Metabolomics Data Integration With Xcms · holobiomicslabUse when you have untargeted LC-MS metabolomics data preprocessed with XCMS and need to filter out low-quality peak integrations that could introduce false positives or noise into metabolite quantification.
- ▌ Metabolomics Experiment Object Handling · holobiomicslabUse when when converting raw metabolomics data from external formats (tab-delimited text, Sciex OS exports) into a unified R analysis environment, or when you have an existing SummarizedExperiment from another pipeline (e.
- ▌ Metabolomics Feature Quality Assessment · holobiomicslabUse when after drift correction and before missing value imputation when your LC-MS peak table contains features with variable detection rates across samples.
- ▌ Metabolomics Npp Reliability Assessment · holobiomicslabUse when you have completed NPP runs from one or more metabolomics tools on a set of LC-HRMS mzML files AND you have reference information (target molecule list with molecular formula, main adduct, and RT boundaries) available for a subset of expected compounds in those files.
- ▌ Molecular Graph Representation Encoding · holobiomicslabUse when when you have a collection of molecular structures (as InChI strings, SMILES, or RDKit Mol objects) and need to feed them into a pretrained or transfer-learning neural network that expects both molecular graph topology and structural fingerprints as inputs.
- ▌ Molecular Graph Representation Learning · holobiomicslabUse when when you have molecular structures (SMILES or chemical graphs) from a database like PubChem and need to predict molecular properties (e.
- ▌ Molecular Network Metadata Organization · holobiomicslabUse when when you have a GNPS molecular networking task ID (from GNPS1 or GNPS2 workflows: METABOLOMICS-SNETS, METABOLOMICS-SNETS-V2, FEATURE-BASED-MOLECULAR-NETWORKING, classical_networking_workflow, or feature_based_molecular_networking_workflow) and need to prepare the job archive for NPLinker.
- ▌ Motif Similarity Ranking Interpretation · holobiomicslabUse when after executing MassQL queries against a MotifDB reference database and retrieving ranked motif matches, when you need to determine which database entries represent true structural correspondence versus spurious matches, and to decide whether a motif's top-ranking hit is sufficiently.
- ▌ Ms Dial Feature Detection And Alignment · holobiomicslabUse when when you have raw LC-HRMS data in .mzML or .abf format and need to detect metabolite features (peaks) across multiple samples, align them temporally and by mass-to-charge ratio, and generate a reproducible feature matrix.
- ▌ Ms Ms Spectrum Annotation Preprocessing · holobiomicslabUse when you have raw LC-MS/MS data acquired in Data-Dependent Acquisition (DDA) mode and need to create a labeled training dataset for customized purification model development.
- ▌ Ms Spectrum Filtering And Normalization · holobiomicslabUse when you have raw or parsed tandem MS spectra (MGF, mzML, or in-memory Spectrum objects) and need to remove artifacts and normalize intensities prior to spectral matching, library searching, or downstream analysis.
- ▌ Transformation Product Parent Linkage · holobiomicslabUse when after suspect screening has identified both parent features (from before-treatment or reference samples) and TP candidate features (from after-treatment or exposed samples) in the same analysis set, and you have MS/MS spectral data or formula annotations available.
- ▌ Transformer Encoder Decoder Inference · holobiomicslabUse when when you have preprocessed MS/MS spectra (normalized intensities, filtered for quality, with top peaks retained) that have been encoded using a spectral representation method (e.
- ▌ Transient Window Function Application · holobiomicslabUse when when processing raw FT-ICR transient files (Bruker Solarix .d format or equivalent) intended for high-resolution mass spectral analysis.
- ▌ Unit Test Design For Model Parameters · holobiomicslabUse when you have extended a neural network model class (e.
- ▌ Algorithm Parameter Comparison Analysis · holobiomicslabUse when when you need to evaluate how a specific algorithm parameter (such as SearchMolecularFormulas first_hit mode) affects the quantity and quality of molecular formula assignments on a given spectrum or dataset.
- ▌ Candidate Formula Ranking By Mass Error · holobiomicslabUse when after querying a formula database (KEGG, PubChem, or user-supplied) with neutral mass values derived from observed m/z peaks and adduct transformations, when multiple candidate formulae fall within the configured mass tolerance window (ppm or Da) and you need to rank them by likelihood.
- ▌ Data Normalization In Mass Spectrometry · holobiomicslabUse when you have raw or partially processed metabolomics data (mzML/mzXML format) from LC-MS or GC-MS runs and need to apply standardized feature detection, alignment, and intensity normalization as part of a reproducible workflow.
- ▌ Fda Repeatability Compliance Assessment · holobiomicslabUse when you have completed NMR data quality control analysis and possess per-feature CV values, and you need to formally assess whether the metabolomic dataset meets FDA regulatory standards for downstream biomarker discovery or quantitative assays.
- ▌ Metabolite Prediction Pathway Selection · holobiomicslabUse when when you have a small-molecule structure (SMILES, MOL, or SDF format) and need to predict its metabolic fate across one or more biological systems.
- ▌ Nmr Peak Prediction Via Database Lookup · holobiomicslabUse when you have extracted peak data (chemical shift values in ppm, multiplicities, integration) from a processed NMR spectrum (JCAMP, RAW, or mzML format) and need to match these peaks against known NMR signals to propose or confirm structural assignments.
- ▌ Peak List Filtering For Quality Control · holobiomicslabUse when you have extracted a raw peak list (chemical shifts in a TXT file, one per row) from a 1D 1H NMR spectrum and intend to pass it to NMRformer or a similar deep learning model for metabolite identification.
- ▌ Qc Coefficient Of Variation Calculation · holobiomicslabUse when after extracting NMR spectra and designating replicate QC samples (typically 10 samples run throughout the study), calculate CV for each metabolite feature to assess which signals are reproducible enough for downstream metabolite-phenotype association testing.
- ▌ Tabular Data Transformation With Pandas · holobiomicslabUse when when converting mwTab-formatted metabolomics files (containing MS/NMR tabular data blocks) to JSON, or when you need to extract, manipulate, and re-serialize tabular sections from mwTab files while maintaining column structure and type information.
- ▌ Tensorflow Serving Endpoint Integration · holobiomicslabUse when you have nuclear magnetic resonance (NMR) peak data (1H and 13C measurements) that you need to classify using a deployed SMART 3 model, and you want to submit peaks programmatically rather than through a web UI.
- ▌ Atac Seq Feature Matrix Construction · holobiomicslabUse when when you have processed scATAC-seq data (peak calling complete, cell-barcode matrix generated) and need to register it into ArchR for downstream multiome analysis alongside scRNA-seq gene expression data.
- ▌ Bed Format Generation From Dataframe · holobiomicslabUse when you have extracted quantitative genomic features (e.g., insulation scores, boundary annotations) as a pandas DataFrame with bin coordinates and boolean or numeric columns, and need to export them as BED format for visualization in genome browsers (e.
- ▌ Binary Path Detection And Validation · holobiomicslabUse when when setting up a bioinformatics pipeline (particularly Hi-C data processing) that depends on multiple external binaries with version constraints, and you need to configure the environment in a way that is both portable across systems and reproducible across runs.
- ▌ Contact Distance Binning Logarithmic · holobiomicslabUse when you have a precomputed expected contact frequency table (TSV format with columns: dist_bp, contact_frequency, n_valid) derived from cooler files and need to compress distance-dependent contact probabilities into log-spaced bins.
- ▌ Cosine Similarity Metric Application · holobiomicslabUse when you are performing dimensionality reduction on a sparse single-cell count matrix (in CSR format) and need to compute pairwise cell similarities before spectral decomposition.
- ▌ Cross Modality Embedding Integration · holobiomicslabUse when you have paired scATAC-seq peak matrices and scRNA-seq gene expression matrices from the same cells (multiome data) and need to perform joint clustering, visualization, or correlation analysis across both chromatin accessibility and gene expression in a single coordinate system.
- ▌ Dependency Resolution And Validation · holobiomicslabUse when before running HiC-Pro or similar multi-stage pipelines on a new system or environment, especially when dependency installation is not automated (e.g., not in a conda environment or container).
- ▌ Dna Methylation File Format Handling · holobiomicslabUse when you have CpG methylation call files (from Bismark or MethylDackel) and need to load them into R for differential methylation analysis, but anticipate memory constraints or want to avoid loading the entire dataset into memory.
- ▌ Methylation Object Merging And Union · holobiomicslabUse when you have loaded individual methylation call files as methylRawList objects from bisulfite sequencing experiments (via methRead()) and need to perform base-level comparative analysis across two or more samples.
- ▌ Motif Enrichment Statistical Testing · holobiomicslabUse when after identifying a set of differentially accessible peaks (via tl.diff_test or equivalent), when you need to infer which transcription factors may regulate the observed chromatin state changes.
- ▌ Peak Calling Pseudo Bulk Aggregation · holobiomicslabUse when after clustering single-cell ATAC-seq data (e.g., via Leiden clustering on spectral embeddings), use this skill to identify peaks within each cluster. Triggering conditions: (1) you have sparse, per-cell insertion counts organized in a tile matrix;
- ▌ Python Dependency Version Resolution · holobiomicslabUse when when setting up a new conda environment for a Python-based bioinformatics pipeline and you need to confirm that all declared dependencies (e.g., pysam >=0.15.4, bx-python >=0.8.8, numpy >=1.18.1, scipy >=1.4.
- ▌ Adduct Mass Calculation From Smiles · holobiomicslabUse when when you have a metabolite SMILES structure and need to predict which adduct ions will appear in a mass spectrum acquired with a chemical derivatizing matrix.
- ▌ API Response Parsing And Validation · holobiomicslabUse when after submitting a POST request to the /api/smart3/search endpoint with peak data as a JSON payload, you receive an HTTP response and need to extract classification predictions and confidence scores.
- ▌ Bgc Mf Link Scoring Standardisation · holobiomicslabUse when you have computed raw strain correlation scores and IOKR scores for the same set of GCF–MF (gene cluster family–molecular feature) pairs, and you want to compare or combine them fairly without one score dominating due to scale differences.
- ▌ Biochemical Transformation Matching · holobiomicslabUse when you have a filtered FT-ICR MS peak list (m/z values and assigned molecular formulas per sample) and wish to reconstruct biochemical transformation networks ab initio to characterize how microbial or environmental metabolic pathways differ across conditions.
- ▌ Byte Level Serialization Validation · holobiomicslabUse when when you have implemented a binary file format encoder (such as igzip header construction) and need to verify that the binary output is correct before deploying it to read or write real files.
- ▌ Chemical Structure Feature Encoding · holobiomicslabUse when when you have a set of molecules with known chemical structures and need to prepare them for classification or prediction tasks.
- ▌ Code Quality Metrics Interpretation · holobiomicslabUse when after a GitHub Actions CI workflow has executed static analysis (e.g., via Sonarcloud) and generated a quality report.
- ▌ Collision Cross Section Calculation · holobiomicslabUse when you have raw or processed arrival-time data from a traveling-wave ion mobility mass spectrometry (TWIM-MS) platform and need to convert it into standardized collision cross section (CCS) values for comparative analysis across samples or datasets.
- ▌ Cross Reference Publication Linking · holobiomicslabUse when when cataloging a suite of related bioinformatics tools or web applications (particularly in domains like metabolomics, microbiology, or systems biology) and you need to establish the authoritative peer-reviewed or preprint publication for each tool, verify publication URLs are live, and.
- ▌ Cytoscape Edge Node File Formatting · holobiomicslabUse when after computing pairwise mass-difference transformations between FT-ICR MS peaks and matching them to a reference biochemical transformation key, you have putative edge data (source peak, target peak, transformation type, mass error) and node data (peaks with m/z, molecular formula.
- ▌ Dataframe Plotting Interface Design · holobiomicslabUse when when building a scientific visualization library that must support multiple plotting backends and needs to avoid backend-specific code duplication. Specifically: (1) your domain (e.
- ▌ Feasible Flux Distribution Sampling · holobiomicslabUse when after integrating transcriptomics-derived (RAS), metabolomics-derived (RPS), and extracellular flux constraints into cell-relative metabolic models, sample the feasible flux region when you need to: (1) visualize and compare the metabolic phenotype distributions across biological samples.
- ▌ Functional Trait Diversity Analysis · holobiomicslabUse when when you have abundance-normalized FT-ICR MS peak data with assigned molecular formulas and need to distinguish between richness (total number of distinct metabolites) and functional diversity (diversity in metabolic potential).
- ▌ Growth Yield Computation From Omics · holobiomicslabUse when you have constraint-based metabolic models with integrated transcriptomics (gene expression), intracellular metabolomics (substrate concentrations), and extracellular flux measurements (glucose uptake, lactate production, etc.), and you need to test whether differential expression of.
- ▌ Ion Mobility Feature Classification · holobiomicslabUse when you have raw or processed TWIM-MS data (arrival time and m/z pairs) from multiple lipid, protein, or metabolite classes and need to classify features by biomolecular type before—or instead of—performing feature identification.
- ▌ Java Application Build Verification · holobiomicslabUse when when you need to confirm that a Java project's GitHub Actions workflow (e.g., 'dev_build_release.yml') has completed successfully and generated usable build artifacts;
- ▌ JSON Data Serialization And Parsing · holobiomicslabUse when you have discovered Mass2Motifs via LDA and need to (1) load a pre-computed motifset JSON file (e.g., motifset_optimized.
- ▌ Mass Spectrometry Adduct Annotation · holobiomicslabUse when you have preprocessed MS/MS spectra with measured precursor m/z values but the ionization adduct type is unknown or ambiguous—particularly in metabolomics workflows where multiple adducts co-occur or in de novo formula annotation without access to curated spectral databases.
- ▌ Metabolite Abundance Stratification · holobiomicslabUse when you have a peak-abundance matrix from FT-ICR MS (peaks as rows, samples as columns with raw peak intensities) and need to compute abundance-based diversity indices or functional diversity metrics that are sensitive to relative vs. absolute peak heights.
- ▌ Metabolite Generation Logic Mapping · holobiomicslabUse when when you have access to the MAGMa source code and need to understand or audit how in silico metabolite candidates are enumerated from parent structures.
- ▌ Motif Database Lookup And Retrieval · holobiomicslabUse when after completing the MS2LDA LDA modeling step when you have a JSON-serialized inferred motifset (Mass2Motifs with fragment and neutral-loss patterns) and need to annotate those motifs by comparing them against a curated MotifDB reference database to identify known structural subpatterns.
- ▌ Multi Dimensional Feature Alignment · holobiomicslabUse when you have two or more feature tables in HDF5 format with detected peaks characterized across multiple dimensions (mz, drift_time, retention_time, intensity) and need to harmonize feature coordinates across samples to account for instrument variation, enabling downstream cross-sample.
- ▌ Multi Service Dependency Management · holobiomicslabUse when your research software comprises multiple independent subprojects or microservices (calculation engines, web services, data processors, websites) that must be deployed and initialized in a coordinated sequence, with explicit dependency declarations and network communication paths between.
- ▌ Mzml To Mzpeak Binary Serialization · holobiomicslabUse when you have one or more mzML files (XML-based mass spectrometry data) and need to convert them into mzPeak format for downstream analysis, archival, or integration with tools that consume Parquet-based spectra.
- ▌ Natural Product Database Validation · holobiomicslabUse when after curating and integrating structure-organism pairs from multiple source databases, and before publishing or using the dataset for computational research. Apply this skill when you have aggregated organism counts binned by structural diversity (e.
- ▌ Peak Data Extraction And Formatting · holobiomicslabUse when when you have initialized an MsBackend subclass (e.
- ▌ Poisson Noise Injection For Imaging · holobiomicslabUse when augmenting mass spectrometry ion images for contrastive learning, particularly when the model must generalize across different detector conditions or signal-to-noise ratios.
- ▌ Reaction Activity Score Computation · holobiomicslabUse when you have RNA-seq read count data and a metabolic model with GPR rules, and you need to assess how differential gene expression translates into differential metabolic reaction capacity across multiple biological conditions or cell lines.
- ▌ Response Status Code Interpretation · holobiomicslabUse when when you need to confirm that a documented web service endpoint is deployed and accessible before using it for analysis, or when troubleshooting tool availability in a bioinformatics pipeline. Apply this skill after obtaining a service URL (e.
- ▌ Sinusoidal Embedding Implementation · holobiomicslabUse when building a transformer-based neural network for chemical formula ranking or classification from mass spectrometry spectra, and you need to encode categorical chemical formulas (e.
- ▌ Spectral Library Parallel Ingestion · holobiomicslabUse when you have multiple MSP or spectral library files (e.g., one per batch of analytical standards, or organized in a directory structure) that need to be read and merged into a single library object for downstream enrichment (SMILES assignment, RI annotation, write operations).
- ▌ Toml Configuration File Preparation · holobiomicslabUse when you have exported lipid identifications from MS-DIAL (version 4 or 5) and need to run LipoCLEAN quality filtering on that output.
- ▌ Volume Mount And Persistence Design · holobiomicslabUse when you are composing multiple containerized services that have data dependencies—e.g., a web application that launches calculation jobs, a calculation engine that consumes a pre-built lookup database, and a data-processing service that generates that lookup.
- ▌ Aligned Feature Matrix Construction · holobiomicslabUse when when you have extracted feature tables from multiple breath samples (mzML/mzXML files) using feature extraction, and you need to identify which features are the same across samples to enable downstream statistical or comparative analysis.
- ▌ Analyte Degradation Risk Assessment · holobiomicslabUse when you have measured metabolites or lipids from biobanked or processed blood samples (EDTA plasma or serum) and know the pre-analytical conditions (time delay before/after centrifugation in hours, processing temperature in °C, sample matrix).
- ▌ Approximate Nearest Neighbor Search · holobiomicslabUse when when you have pre-computed spectrum embeddings (e.g., Word2vec vectors) and need to rapidly retrieve the most similar spectra from a large in-silico or experimental library (thousands to millions of spectra).
- ▌ Arrival Time Based Class Assignment · holobiomicslabUse when you have raw or processed TWIM-MS data with arrival time and m/z dimensions, and you need to label each experimental feature by biomolecular class before feature identification or peak detection steps are complete.
- ▌ Asynchronous Workflow Orchestration · holobiomicslabUse when when you have a Streamlit-based scientific application (e.g. OpenMS workflows) that must support both online Docker deployments with multiple worker processes and offline local deployments, and you need to prevent long-running analyses from blocking the web UI thread.
- ▌ Batch Effect Correction Application · holobiomicslabUse when you have a SummarizedExperiment object containing metabolomics assay data (compound areas, internal standard areas) organized by batch and sample type (including pooled SQC samples), and you need to correct systematic variation across batches before calculating Relative Standard Deviation.
- ▌ Batch File Processing Orchestration · holobiomicslabUse when when you have multiple CDF imaging files (e.g., from mass spectrometry imaging scans of biological samples) that need to be read into a single Matlab workspace with consistent structure and metadata (spectral intensity, m/z arrays, spatial coordinates).
- ▌ Bipartite Graph Layout Optimization · holobiomicslabUse when after constructing a bipartite network graph with metabolites and proteins as nodes weighted by association strength and disease class, when you need to render the network visually for interpretation and publication.
- ▌ Blockwise Data Parsing And Indexing · holobiomicslabUse when when you need random access into a large text or XML file that you want to keep compressed, where the file has natural logical divisions (chapters, spectra, records) that can be written independently.
- ▌ Breath Volatile Peak Classification · holobiomicslabUse when after feature extraction and alignment when you have a numerical feature table (CSV or dataframe) with intensity values across retention time or m/z dimensions, and you need to identify and rank peaks by signal quality and prominence rather than relying on all extracted features equally.
- ▌ Candidate Peak Filtering Annotation · holobiomicslabUse when you have a peak list extracted from MSI data that includes candidate peaks with potential m/z overlap or spatial co-localization patterns across tissue images.
- ▌ Cascade Search Strategy Fdr Control · holobiomicslabUse when when searching high-resolution mass spectra against spectral libraries and you need to identify both unmodified and post-translationally modified peptides while maintaining strict control over false positive identifications.
- ▌ Ccs Value Assignment From Standards · holobiomicslabUse when you have TWIM-MS experimental data with arrival/drift times and m/z values, and you possess calibrant reference standards with known CCS values.
- ▌ Chemical Classification Aggregation · holobiomicslabUse when you have a collection of standardized molecular structures (SMILES or MOL format) that have been submitted to ClassyFire and you need to collect and organize their classification results into a single CSV or JSON table for downstream analysis, validation, or retention time modeling.
- ▌ Chemical Formula Annotation Mapping · holobiomicslabUse when you have processed MSI peak data (in rMSIproc format) and need to distinguish matrix-related ions from analyte signals. Specifically: (1) you have a peak matrix with m/z values and spatial intensity maps; (2) you have a reference matrix identity (e.g., 'Ag1' for silver);
- ▌ Chromatographic Baseline Estimation · holobiomicslabUse when you have extracted ion chromatogram (EIC) candidate data from untargeted LC/HRMS files (mzXML, mzML, or netCDF format) and need to identify genuine peaks within each EIC.
- ▌ Chromatographic Peak Classification · holobiomicslabUse when you have (1) a benchmark dataset of reference peaks with validated m/z, retention time boundaries, and isotopologue assignments, and (2) NPP output feature tables (unaligned and aligned) from tools like XCMS, MZmine 2, or MS-DIAL that you wish to evaluate.