HolobiomicsLab
- 7.4k skills
- 0 followers
- 19 hours ago last updated
- ▌ Entry Status Tabulation · holobiomicslabUse when you need to assess the curation completeness and status distribution of a MIBiG JSON dataset—for instance, to identify how many entries are in 'active', 'retired', or 'pending' states, or to audit changes in entry status over time.
- ▌ Feature Matrix Assembly · holobiomicslabUse when when you have a set of chemical structures (SMILES, SDF, mol, mol2, or hin files) and need to convert them into a tabular feature representation for retention time prediction, metabolite annotation, or other quantitative structure–property modeling tasks.
- ▌ Fold Change Calculation · holobiomicslabUse when after completing statistical tests (t-test, Mann–Whitney U, or ANOVA) on matched lipid abundances grouped by experimental condition or sample category, and before compiling final results tables.
- ▌
- ▌ Gwas Result Integration · holobiomicslabUse when you have independent metabolomic GWAS results (metabolite p-values, effect sizes, metabolite IDs) and separate meta-genome GWAS results (variant p-values, effect sizes, genomic variant IDs) from similar diseases or phenotypes, and you want to discover novel metabolite–gene associations and.
- ▌ Hpc Job Parallelization · holobiomicslabUse when you have a filtered set of conformers (100s–1000s) from ASE-ANI that each require independent quantum calculations via QUICK, and you have access to HPC resources with multiple cores or nodes. Parallelization is necessary when serial execution would exceed practical time budgets (e.
- ▌ Lipid Database Querying · holobiomicslabUse when you have acquired full-scan mass spectrometry imaging data (e.g., from a mouse bladder or tissue section) with detected m/z features and want to assign chemical identities to those features by querying a structured lipid database.
- ▌ Matlab Script Execution · holobiomicslabUse when you have located MATLAB scripts in a Codes-Explained folder with accompanying Read-Me.
- ▌ Model Residual Analysis · holobiomicslabUse when after fitting a linear model to normalized metabolomics featuredata using LinearModelFit(), to validate model assumptions (homogeneity of variance, normality), detect influential observations, and confirm that the model's coefficient and p-value estimates are reliable for downstream.
- ▌ Mz Binning And Indexing · holobiomicslabUse when immediately after parsing mzML files into (m/z, scan_number, intensity) tuples when you need to build mass tracks from raw MS1 spectra. Use it when working with high-resolution instruments (e.
- ▌ Peak Collision Flagging · holobiomicslabUse when you have extracted a peak list from MSI data and need to annotate matrix-related signals, but overlapping peaks or isobaric ions (ions with identical or near-identical m/z values) risk being misclassified.
- ▌ Psm Record Augmentation · holobiomicslabUse when your PSM input file (e.g., from MaxQuant or other search engines that do not report fixed modifications) lacks modification annotations for residues that were chemically modified during sample preparation (e.g., carbamidomethylation of cysteines, TMT labeling of lysines and N-termini).
- ▌ Pubchem API Integration · holobiomicslabUse when when you have raw chemical structures in diverse input formats (SMILES, SDF, or other molecular representations) from multiple sources and need to produce a uniform, canonicalized representation before molecular descriptor calculation, fingerprinting, or retention time modeling.
- ▌ Qc Signal Normalization · holobiomicslabUse when your metabolomic dataset contains dedicated QC samples (pooled or standard reference material injected at intervals throughout the analytical sequence) and you observe intensity drift, batch effects, or systematic signal variation correlated with run order or acquisition time rather than.
- ▌ Sbml Model Manipulation · holobiomicslabUse when when you have consensus metabolic reconstructions in SBML format for individual community members and need to resolve metabolic gaps by leveraging cross-member dependencies and community-level constraints before phenotypic validation or flux analysis.
- ▌ Smiles Mol File Parsing · holobiomicslabUse when when you have molecular structures encoded as SMILES strings or MOL files from an in-house database or spectroscopy repository, and need to convert them into node-edge graph representations with explicit atom features (atomic number, degree, formal charge, hybridization) and bond features.
- ▌ Tabular Data Validation · holobiomicslabUse when after curating and integrating structure-organism pairs from multiple sources when you need to produce a high-confidence subset suitable for publication or computational validation.
- ▌ Tensor Embedding Design · holobiomicslabUse when when you need to represent discrete chemical formulae (e.
- ▌ Type Coercion To String · holobiomicslabUse when when converting intermediate JSON records to output dictionaries via matrix directives and all field values must be serialized as strings for downstream format compatibility (e.g., mwTab format or other repository-specific formats that expect string-typed metadata fields).
- ▌ Unit Test Design Pytest · holobiomicslabUse when after implementing a custom Filter subclass (e.g., Tanimoto threshold filter) in minedatabase/filters.py and before integrating it into a pickaxe_run.py workflow.
- ▌ Volcano Plot Generation · holobiomicslabUse when you have CSV files containing fold-change estimates and statistical p-values from a differential expression analysis (e.
- ▌
- ▌ Qc Failure Event Detection · holobiomicslabUse when rapid QC-MS is actively monitoring LC-MS data acquisition and a QC check result (e.g., internal standard retention time drift, m/z deviation, or intensity threshold breach) returns a fail status.
- ▌ Breath Biomarker Discovery · holobiomicslabUse when you have GC–MS data from human breath samples and need to identify marker metabolites for disease diagnosis, phenotyping, or biomarker discovery without a predefined target list. Your data is noisy or conventional peak picking has produced high false-positive rates.
- ▌ Metabolomics Nmr Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — NMR — search this unit's 276 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Metabolomics Ce Ms Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — CE-MS — search this unit's 114 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Metabolomics Gc Ms Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — GC-MS — search this unit's 367 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Metabolomics Lc Ms Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — LC-MS — search this unit's 2,621 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Masster · holobiomicslabUse when you need to run the Zamboni-lab Masster (MASSter) workflow for untargeted LC-MS metabolomics data analysis. NONCOMMERCIAL tool — confirm permitted use before applying (see License notice).
- ▌ Metabolomics Collection Router · holobiomicslabUse when an agent needs to find and apply a computational-metabolomics / LC-MS-MS skill from this collection, and optionally ground it against the source paper via Perspicacité before acting.
- ▌ Metabolomics Ms Generic Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — mass-spectrometry — search this unit's 804 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Metabolomics Ms Imaging Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — MS-imaging — search this unit's 292 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Metabolomics Ion Mobility Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — ion-mobility-MS — search this unit's 390 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Msp File Parsing · holobiomicslabUse when you have a Mass Spectrum Point (MSP) file containing electron ionization mass spectral records with header fields and peak intensity pairs, and you need to load it into an R data structure for library searching, format validation, or round-trip conversion.
- ▌ Mgf File Parsing · holobiomicslabUse when you have raw or MZmine-processed MGF files (containing MS/MS spectra with m/z values, intensities, and precursor masses) that need to be segmented by sample, deduplicated, or prepared for fragment annotation and adduct analysis in the MolNotator pipeline.
- ▌ Jcamp File Parsing · holobiomicslabUse when you receive uploaded spectral data in JCAMP format (jcamp) as input to the /api/v1/chemspectra/file/convert endpoint, or when you need to extract and validate metadata and peak information from an existing JCAMP file before converting to another format or performing spectral analysis.
- ▌ Mwtab File Parsing · holobiomicslabUse when you have mwTab format files (Mass Spectrometry or Nuclear Magnetic Resonance experimental data from Metabolomics Workbench) that need to be loaded into Python for downstream analysis, validation, conversion to JSON, or programmatic manipulation of metadata and data sections.
- ▌ Asb Contribute · holobiomicslabUse when an ASB skill proved wrong, stale, missing or wasteful in practice — its steps failed, no skill covered the task, the leaves existed but nothing composed them, or the tool has changed. Turns that friction into a redacted, dedupable report the user approves before anything is filed.
- ▌ Metabolomics Direct Infusion Router · holobiomicslabUse when a task needs a skill from ASB Metabolomics — direct-infusion-MS — search this unit's 97 evidence-grounded skills, then apply and optionally ground the one that fits.
- ▌ Ms2 Dereplication · holobiomicslabUse when you have MS2 tandem mass spectrometry data in .mzML format and need to match unknown spectra against a reference library (GNPS, HMDB, MassBank) to identify which known compounds are present in your sample.
- ▌ Mzml File Parsing · holobiomicslabUse when you have raw LC- or GC-HRMS data from vendor instruments (ESI or APCI ionization) that needs to be converted to a vendor-neutral format for non-target screening, or you already have mzML files that require loading into a Python environment for downstream feature detection and MS2 spectral.
- ▌ Peak M Z Indexing · holobiomicslabUse when when you have LC-MS-MS metabolomics data in MGF format and need to prepare it for unsupervised analysis (e.g., topic modeling with LDA).
- ▌ Mnar Data Handling · holobiomicslabUse when you have metabolomics data (targeted LC/MS or untargeted GC/MS) with left-censored missing values below the limit of quantification (LOQ) or limit of detection (LOD), and you need to impute these values while preserving the underlying distributional structure and avoiding bias from.
- ▌ Msp To CSV Parsing · holobiomicslabUse when you have a .msp format MS/MS spectrum library (e.g., from MassBank or similar public databases) and need to convert it into individual CSV entries indexed by positive or negative ionisation mode for use as a custom fragment library in MetaboAnnotatoR annotation pipelines.
- ▌ Mzml Mzxml Parsing · holobiomicslabUse when you have raw LC-MS/MS data in mzML or mzXML format and need to isolate specific MS1/MS2 scan pairs for a targeted compound list or for building a local spectral library.
- ▌ Peak Count Capping · holobiomicslabUse when apply peak-count capping when preprocessing tandem mass spectrometry (MS/MS) spectra for peptide identification or spectral library matching, particularly when working with high-resolution spectra that may retain numerous low-intensity noise peaks after intensity filtering.
- ▌ Topic Modeling Lda · holobiomicslabUse when after generating a corpus/features JSON file from raw MS2 fragmentation spectra or similar document-term data, when the goal is unsupervised discovery of latent topics (e.g., metabolite families or spectral motifs) without ground-truth labels.
- ▌ Usi String Parsing · holobiomicslabUse when when you have a USI string (e.g., mzspec:GNPS:TASK-d93bdbb5cdda40e48975e6e18a45c3ce-f.mwang87/data/Yao_Streptomyces/roseosporus/0518_s_BuOH.
- ▌ Evidence Linking · holobiomicslabUse when you have tandem MS/MS spectra annotated by at least two of GNPS (FBMN), ISDB-LOTUS (CFM-ID 4.0 spectral matching), and Sirius 6, and you need to rank features by annotation agreement rather than trust a single tool's output.
- ▌ Output Grounding · holobiomicslabUse when when developing or extending mass spectrometry data processing workflows (e.
- ▌ Gcf Table Export · holobiomicslabUse when you have completed a BiG-SLiCE v2 clustering run and need to extract the pre-calculated BGC and GCF cluster assignments in TSV format for postprocessing, integration with SQL pipelines, or sharing with collaborators who require flat tabular output rather than the SQLite database or.
- ▌ Asb Metabolomics · holobiomicslabUse when starting any task with the ASB Metabolomics skill collection — read this meta-skill first. It explains good practice (search -> apply -> ground), enforces the license-tier acknowledgment for non-open tools, then hands off to the _router skill for actual skill selection.
- ▌ Bgc Mf Link Scoring · holobiomicslabUse when you have preprocessed GCF-MF link pairs from paired genomics–metabolomics datasets (antiSMASH-detected BGCs clustered into GCFs, and GNPS spectra grouped into MFs) and need to rank them by likelihood of representing true natural product–biosynthetic gene associations.
- ▌ File I O Automation · holobiomicslabUse when you have MZmine-exported LC-MS/MS data (MGF spectra and CSV metadata files) in both positive and negative ionization modes and need to execute the full MolNotator pipeline from duplicate filtering through dereplication and network generation.
- ▌ Gcf Mf Link Scoring · holobiomicslabUse when after BGC detection and clustering (producing GCFs) and metabolomics profiling (producing MFs with MS/MS spectra), when you have paired genomic and metabolomic data from the same microbial strains and need to rank which GCF–MF pairs are likely to represent true biosynthetic relationships.
- ▌ Pyteomics API Usage · holobiomicslabUse when when you have polypeptide sequences and need to compute their monoisotopic or average mass, isotopic distribution patterns, or other physico-chemical properties; or when you need to parse and manipulate MS/LC-MS data, FASTA databases, or search engine output in a Python workflow.
- ▌ R Package API Usage · holobiomicslabUse when when you have metabolomics data (tab-delimited text or SummarizedExperiment object) and need to apply batch correction, outlier detection, internal standard recommendation, and quality filtering at scale or in non-interactive workflows.
- ▌ Tabular Data Export · holobiomicslabUse when you have extracted MS1 or MS2 peak lists and scan headers from Thermo Fisher RAW files using MetaXtract and need to load them into pandas, NumPy, or external analysis tools.
- ▌ Peak Table Extraction · holobiomicslabUse when you have raw or converted spectral data (jcamp, RAW, or mzML format) from NMR, IR, or MS instruments and need to identify individual peaks, extract their properties (chemical shift, m/z, intensity, width), and generate a structured peak table for annotation, comparison, or publication.
- ▌ Tabular Data Cleaning · holobiomicslabUse when you have a CSV or table-format spectral peak list (with chemical shift, intensity, and metadata columns) destined for NMRformer or similar peak-to-metabolite assignment models, and you need to exclude low-quality peaks that would otherwise harm prediction accuracy.
- ▌ JSON File Parsing · holobiomicslabUse when when you have JSON-formatted curation data organized in a directory structure (e.g., the MIBiG `data` directory) and need to extract a specific nested field (e.g., `cluster.status`) across all entries to build an index, summary table, or quality report.
- ▌ Kegg API Querying · holobiomicslabUse when after metabolite KEGG identifiers and hierarchy metadata have been assigned (via assign_hierarchy with 'KEGG' identifier) and you need to enrich your metabolomics count data frame with functional orthology and associated gene information from KEGG.
- ▌ 2d Tic Preprocessing · holobiomicslabUse when when you have raw GCxGC-MS data imported from NetCDF into a 2D-TIC chromatogram object and need to remove chemical and instrumental noise (column bleeding, baseline drift, detector contamination) to reveal metabolite differences between sample groups for downstream multiway PCA or.
- ▌ Drift Time Filtering · holobiomicslabUse when you have loaded a raw GCIMS dataset and need to isolate the region of interest in drift time (typically 5–16 ms for small organic molecules) to exclude low-drift-time chemical noise, high-drift-time tail artifacts, or off-scale ion signals that would degrade subsequent alignment and peak.
- ▌ Lcms Raw Data Import · holobiomicslabUse when you have raw LC/MS data in mzML or mzXML format and need to initiate untargeted metabolomics analysis.
- ▌ Ms Scan File Parsing · holobiomicslabUse when you have acquired a Thermo mass spectrometry RAW file (or other proprietary instrument format) and need to extract MS1 and/or MS2 scans in an open, interoperable format (mzML, MGF, or Raxport-processed FT1/FT2 files) for TIC visualization, PSM scoring, or stable isotope labeling analysis.
- ▌ Ms2lda Motif Mapping · holobiomicslabUse when you have a GNPS-generated molecular network (either classical or feature-based) and corresponding MS2LDA experiment output containing Mass2Motif-to-spectrum assignments, and you want to visualize which substructural motifs are shared across clusters or features and how they distribute.
- ▌ Peak List Formatting · holobiomicslabUse when after successfully resolving a USI string to a specific mass spectrum scan, and before performing spectral matching, library search, or comparative analysis.
- ▌ Sample Label Mapping · holobiomicslabUse when you have a raw peak table (CSV format, from any of 12 supported LC-MS software tools or standardized format) and a separate label file that assigns each sample to an experimental class (e.
- ▌ Xcms Profile Parsing · holobiomicslabUse when you have xcms-processed LC-MS data with detected feature groups (from xcms grouping), suspect retention time misalignment across samples due to long acquisition periods or large sample cohorts, and need to feed raw profiles into ncGTW's realignment algorithm.
- ▌ JSON Schema Validation · holobiomicslabUse when you have loaded an mwTab file into a structured MWTabFile object and need to verify it conforms to MS or NMR schema specifications before deposition, curation, or downstream analysis.
- ▌ Nmr Peak Deconvolution · holobiomicslabUse when when you have raw or processed 1D NMR spectra (FID or frequency-domain format) and need to extract individual peak identities and quantitative parameters in tabular form.
- ▌ API Data Retrieval · holobiomicslabUse when you need to obtain all project JSON documents currently deposited in a data platform (such as the Paired Omics Data Platform) to validate their structure against a JSON Schema specification, or when you require a complete snapshot of published records for quality assurance, data migration.
- ▌ Bgc Cluster Export · holobiomicslabUse when you have completed a BiG-SLiCE v2 clustering analysis and need to convert the internal SQLite3 database results into human-readable, tabular TSV files for import into spreadsheet applications, statistical tools, or custom analysis pipelines that do not support SQLite3 directly.
- ▌ Data Deduplication · holobiomicslabUse when after parsing and validating a .csv file containing comma-separated SMILES strings, and before formatting the molecule list for CypReact input.
- ▌ Jndi Binding Setup · holobiomicslabUse when deploying a Java web application (such as CEU Mass Mediator) that requires access to an internal or external database and the application server provides JNDI as the connection pooling and naming mechanism.
- ▌ R Function Calling · holobiomicslabUse when after AutoTuner has completed peak identification (TIC analysis), peak isolation, and EIC parameter extraction on at least 3 raw mass spectrometry samples (qTOF, orbitrap, or FTICR formats converted to mzML/mzXML/CDF), and you need to export the tuned parameters in a format ready for XCMS.
- ▌ R Script Execution · holobiomicslabUse when you have located example R scripts in a version-controlled repository (e.g., Codes-Explained folder), a Read-Me.txt file documents the purpose and parameters of each script, and you need to reconstruct or validate a simulation procedure for a specific data sub-sample scenario.
- ▌ S3 Method Dispatch · holobiomicslabUse when your metabolomics data frame has inherited or assigned class metadata (e.
- ▌ Smiles Sdf Parsing · holobiomicslabUse when you have a natural product molecule or compound library provided as SMILES strings, InChI strings, or SDF files, and you need to convert them into an in-memory molecular representation (RDKit Mol object) suitable for fingerprinting, property prediction, or other cheminformatic operations.
- ▌ Ms Data Preprocessing · holobiomicslabUse when you have received raw CE-MS or LC-MS output files in vendor-specific formats from a mass spectrometry instrument and need to process them through an untargeted metabolomics workflow (e.g., AriumMS) that requires standardized, interoperable file formats.
- ▌ Xic Marker Annotation · holobiomicslabUse when when you have a resolved spectrum file (mzML, mzXML) and need to visualize where MS2 precursor scans occur on an XIC display.
- ▌ Mzml File Import Xcms · holobiomicslabUse when you have raw mzML files from a mass spectrometry instrument and need to begin a preprocessing workflow in xcms. This is the essential first step before any peak detection (centWave, MSWParam) or feature grouping can occur.
- ▌ Fragment Ion Matching · holobiomicslabUse when you have two tandem mass spectra (query and reference) with precursor m/z and fragment ion peaks, and you need to identify which fragment ions correspond between them to assess spectral similarity, detect structural variants, or validate compound identifications.
- ▌ Hmdb Formula Sampling · holobiomicslabUse when when you need to create in silico LC-MS/MS experiments with diverse chemical backgrounds for testing fragmentation strategies or acquisition controllers, and you want the chemical diversity to reflect real metabolomic samples. Use this when you have a target m/z range (e.
- ▌ Mass Defect Filtering · holobiomicslabUse when after MS-Dial peak picking and feature table construction, when you observe a high proportion of features with anomalous m/z decimal values that are inconsistent with known metabolite ionization patterns.
- ▌ Mass Track Clustering · holobiomicslabUse when after constructing initial data bins from mzTree (indexed by int(mz × 1000)), determine whether a single bin contains one or multiple mass tracks. Apply clustering when the m/z range of points in a bin exceeds 2 × ppm tolerance (e.
- ▌ Motif To Node Mapping · holobiomicslabUse when you have a GNPS-generated classical or feature-based mass spectral molecular network (graphml or JSON format) and a corresponding MS2LDA experiment with Mass2Motif assignments on the same spectra, and you want to visualize and quantify which structural motifs are shared within and across.
- ▌ Mz Decimal Extraction · holobiomicslabUse when after loading a feature table with m/z values from MS-Dial output when you need to identify and remove features with anomalous decimal m/z values.
- ▌ Mzml Spectral Parsing · holobiomicslabUse when when beginning a metabolomics annotation workflow with raw MS2 spectral data in .mzML format. This step is necessary when you have vendor-converted or standard .
- ▌ Newick Format Parsing · holobiomicslabUse when you have a Chemical Feature Tree artifact (phylogeny) output from q2-qemistree's make-hierarchy method and need to verify its structural validity, count nodes (leaves and internal nodes), measure tree depth, and assess branching patterns before using it for alpha- or beta-diversity.
- ▌ Peak Shape Assessment · holobiomicslabUse when after peak detection in a nontargeted LC-MS workflow when you have a feature table with detected peaks and need to filter low-quality features or understand why certain features have inconsistent intensity or poor annotation confidence.
- ▌ Slack API Integration · holobiomicslabUse when rapid QC-MS detects a QC failure (e.g., internal standard retention time drift, m/z deviation, or intensity anomaly) during an active LC-MS instrument run and users have configured Slack as a notification target.
- ▌ Spectral JSON Parsing · holobiomicslabUse when after submitting an LC-MS/MS fragmentation spectrum to the MSNovelist web service and receiving a JSON response containing ranked de-novo structure candidates.
- ▌ Spectrum Data Loading · holobiomicslabUse when when you have a USI (e.g., mzspec:MTBLS1124:QC07.mzML) pointing to a public mzML or related spectrum file in MetaboLights, MassIVE, or GNPS repositories, and need to load the spectrum data for interactive visualization, quality control assessment, or downstream analysis.
- ▌ Upset Plot Generation · holobiomicslabUse when after loading and filtering search results from two or more DIA-MS analysis tools at a specified Q-value cutoff, when you need to summarize which analytes are identified by all tools, by specific subsets, or uniquely by individual tools.
- ▌ Usi Namespace Parsing · holobiomicslabUse when when you need to retrieve mass spectrometry spectrum data from a metabolomics repository but only have a USI string (e.g., 'mzspec:GNPS:TASK-c95481f0c53d42e78a61bf899e9f9adb-spectra/specs_ms.mgf:scan:1943' or 'mzspec:MASSBANK::accession:SM858102').
- ▌ Entrypoint Verification · holobiomicslabUse when when you have obtained a Python package from a repository (e.g., via git clone) and need to confirm that the documented Python version constraint and pinned dependency versions are sufficient to execute the package's main entry point (typically main.py or a console script).
- ▌ Nmr Spectral Processing · holobiomicslabUse when you have raw NMR/IR/MS spectral data in vendor-specific (RAW), open (jcamp), or mass spectrometry (mzML) formats and need to parse, validate, and convert them to a standardized internal representation with extracted metadata and peak tables for visualization or further analysis in.
- ▌ Bedgraph File Export · holobiomicslabUse when after computing per-bin coverage depth using cooltools.coverage() on a loaded cooler object, when you need to (1) share the coverage track with non-Python tools, (2) visualize it in a genome browser, or (3) integrate it with downstream analyses that expect bedGraph or tabular input.