HolobiomicsLab
- 7.4k skills
- 0 followers
- 2 days ago last updated
- ▌ Spectral Clustering Density Based · holobiomicslabUse when you have computed a sparse pairwise distance matrix from MS/MS spectra (via nearest neighbor indexing) and need to partition spectra into homogeneous clusters—typically when clustering bottom-up proteomics data with the goal of grouping spectra from the same peptide sequence or when you.
- ▌ Spectral Data Object Construction · holobiomicslabUse when you have a set of centroided .mzML LC-MS files from a targeted metabolomics or lipidomics experiment and need to represent them as a structured object that links raw spectra to sample-level metadata (e.
- ▌ Spectral Database Query Execution · holobiomicslabUse when when you have an unknown mass spectrum (query spectrum) and need to search it against a reference database of billions of spectra to find matching or structurally related compounds.
- ▌ Spectral Denoising Formula Method · holobiomicslabUse when you have a noisy MS/MS spectrum and need to identify and remove chemical noise ions (as opposed to electronic noise). You have the precursor compound's SMILES string or molecular formula and its adduct type.
- ▌ Spectral Dimensionality Reduction · holobiomicslabUse when you have high-resolution tandem MS spectra (in mzML, mzXML, or MGF format) that need to be clustered or searched at scale (millions of spectra).
- ▌ Spectral Feature Table Generation · holobiomicslabUse when you have raw LC-MS data in mzXML format (or vendor formats convertible via MS-Convert) and need to identify and quantify metabolic features before multi-sample alignment. Use MS1 peak picking for full-scan or DDA data to extract Gaussian and non-Gaussian shaped peaks;
- ▌ Spectral Library Entry Generation · holobiomicslabUse when you have an experimental or public MS/MS spectrum (e.g., from MassBank in msp format, or a raw centroid-mode chromatogram) and need to create a reusable library entry for a known metabolite.
- ▌ Spectral Metadata Standardization · holobiomicslabUse when you have mass spectrometry spectra stored across multiple, disparate metabolomics repositories (GNPS, MassBank, MetaboLights, Metabolomics Workbench, ProteoXchange, MS2LDA) and need to retrieve them using a single identifier scheme, or you are publishing spectrum figures and need.
- ▌ Spectral Molecular Family Linking · holobiomicslabUse when when you have pre-processed genomic data (GCFs from AntiSMASH/BigScape clustering) and metabolomic data (spectra and molecular families from GNPS molecular networking) and need to systematically score and rank putative relationships between biosynthetic gene clusters and their.
- ▌ Spectral Peak Annotation Proforma · holobiomicslabUse when you have an annotated or raw tandem mass spectrometry spectrum and need to identify which observed peaks correspond to expected peptide fragment ions from a known or predicted peptidoform. Use it before spectrum visualization if you want highlighted, labeled fragment matches;
- ▌ Spectrum Peak To Fragment Mapping · holobiomicslabUse when when you have a tandem mass spectrum (MS/MS) and a ProForma 2.0 peptidoform string (e.g., DLTDYLM[Oxidation]K) and need to identify which observed spectrum peaks correspond to expected b-ion and y-ion fragments, in order to validate peptide identification or annotate spectrum quality.
- ▌ Spectrum Relevance Classification · holobiomicslabUse when you have a collection of MS/MS spectra from reference standards representing your compounds of interest (e.g., flavonoids, prenylated chalcones) and a set of MS/MS spectra from non-target or other compounds.
- ▌ Structural Annotation Integration · holobiomicslabUse when you have structural candidates from in silico tools (SIRIUS/CANOPUS) and library spectral matches from GNPS, but need to resolve conflicting or incomplete chemical classifications into a unified consensus.
- ▌ Taxonomic Weighting In Annotation · holobiomicslabUse when you have a feature table with candidate metabolite annotations (m/z, retention time, chemical identifiers) and MS/MS spectra, and you know the organism or taxon of origin for your samples.
- ▌ Thermo Raw Binary Data Extraction · holobiomicslabUse when you have acquired .raw files from a Thermo mass spectrometer (e.g., Q Exactive, Orbitrap) and need to expose their contents—scan numbers, retention times, m/z values, intensities, and precursor information—for downstream nontargeted LCMS feature detection and alignment.
- ▌ Topn Acquisition Parameter Tuning · holobiomicslabUse when when you have real LC-MS/MS data (mzML format) from an untargeted metabolomics experiment and want to test how variations in TopN DDA parameters affect which precursor ions are selected and fragmented, before deploying the optimized strategy on physical instruments.
- ▌ Transformation Product Prediction · holobiomicslabUse when after parent chemical suspects have been identified in a non-target screening workflow, use this skill when you need to screen for downstream products formed by chemical or biological transformation.
- ▌ Usi String Parsing And Resolution · holobiomicslabUse when you have a USI string referencing a spectrum in an online public repository (PRIDE, MassIVE, etc.) and need to load its raw spectral data without downloading the entire dataset file.
- ▌ Workflow Output Validation And QA · holobiomicslabUse when after executing a Nextflow-based MS-DIAL workflow on .mzML LC-HRMS metabolomics data using Docker or Singularity container backends.
- ▌ Xcms Ramclustr Object Integration · holobiomicslabUse when you have centroid-mode LC–MS all-ion fragmentation (AIF) data already processed through xcms for feature detection and retention-time correction, and a corresponding RamClustR object that groups co-eluting fragment ions into putative spectral clusters.
- ▌ Compound Peak Association Inference · holobiomicslabUse when after peak picking on INADEQUATE NMR spectra when you have a set of peak coordinates and intensities and need to determine which peaks belong to the same molecular compound.
- ▌ Deep Learning Metabolite Annotation · holobiomicslabUse when you have UPLC-HRMS data (ThermoFisher, Agilent, or MSConvert-compatible format) from a water sample, a precursor m/z and retention time of interest, and want to annotate an unknown compound by predicting its molecular formula, structure, and name using deep learning scoring rather than.
- ▌ Dom Chemodiversity Characterization · holobiomicslabUse when you have a formula-assigned FT-ICR MS dataset (molecular formulas already assigned to individual mass features) and seek to understand the chemodiversity landscape and transformation relationships within DOM.
- ▌ Feature Table Consensus Aggregation · holobiomicslabUse when you have detected feature tables from multiple LC-IMS-MS/MS samples and need to establish a unified feature catalog in which each row represents a distinct molecular entity observed across one or more samples, with harmonized m/z, drift time, and retention time coordinates.
- ▌ Hmdb Metabolite Query And Retrieval · holobiomicslabUse when you have identified one or more proton NMR spectral regions-of-interest (ROIs)—defined by lower and upper chemical-shift bounds in ppm—from complex biological samples (serum, saliva, urine, tissue, CSF) and need to generate a ranked list of plausible metabolite identities.
- ▌ Interactive Data Exploration Design · holobiomicslabUse when you have NMR metabolomics measurements paired with pre-analytical metadata (processing delay times, centrifugation timing, sample type such as plasma vs. serum, cohort identifiers) and need to interactively explore how variation in processing conditions drives changes in metabolic.
- ▌ Mass Spectrum Structure Elucidation · holobiomicslabUse when you have an experimental tandem mass spectrum (collision-induced dissociation, CID) and a known chemical formula (or narrow set of candidate formulas), and you need to identify the most likely structure(s) by ranking against a large candidate library such as PubChem.
- ▌ Metabolite Metadata Column Matching · holobiomicslabUse when validating mwTab files deposited to the Metabolomics Workbench and you need to verify that metadata columns match standard naming conventions and contain values in the expected format.
- ▌ Metabolite Peak Assignment From Nmr · holobiomicslabUse when you have a 1D ¹H NMR spectrum (as chemical shift vs. intensity) and a corresponding peak list (chemical shift values), and you need to identify which metabolites are responsible for each detected peak.
- ▌ Nightingale 1h Nmr Data Integration · holobiomicslabUse when you have newly assayed 1H-NMR metabolomics data from Nightingale Health (CSV or TSV format) and need to apply one or more published metabolic risk scores (Deelen et al. all-cause mortality, van den Akker MetaboAge, Würtz cardiovascular event risk, etc.).
- ▌ Nmr Spectrum To Structure Inference · holobiomicslabUse when you have 1D NMR spectra (¹H or ¹³C or both) for an unknown organic compound with ≤19 heavy atoms and need to rapidly predict its molecular formula and connectivity graph without manual peak interpretation or exhaustive combinatorial search.
- ▌ Pre Analytical Delay Stratification · holobiomicslabUse when when you have NMR metabolite measurements paired with documented pre-centrifugation and post-centrifugation delay times, and need to assess how processing delays affect metabolic parameter stability within a plasma or serum sample cohort.
- ▌ Spectral Correlation Interpretation · holobiomicslabUse when you have preprocessed 1H NMR spectral data (e.g., from plasma or biological samples acquired on a 600 MHz instrument) and need to identify the chemical composition of a prominent but structurally ambiguous peak.
- ▌ Spectral Peak Detection And Picking · holobiomicslabUse when you have raw INADEQUATE NMR spectra files and need to transition from continuous spectral data to discrete peak coordinates. Use it as the first signal-processing step before clustering peaks into networks or matching against metabolite databases.
- ▌ Transformer Based Fragment Assembly · holobiomicslabUse when when you have CNN-encoded spectral features (¹H and/or ¹³C NMR) and a set of predicted or candidate molecular fragments, and you need to determine which fragments are present and how they connect to form a valid molecular structure.
- ▌ Weighted Loss Function Optimization · holobiomicslabUse when when training a dual-encoder architecture (bi-encoder + cross-encoder) on NMR spectral data where independent encoding and joint pair processing produce competing or imbalanced gradient signals.
- ▌ Chip Seq Signal Pileup Extension · holobiomicslabUse when after duplicate filtering and fragment length prediction (d) in ChIP-Seq analysis, when you need to convert discrete read alignments into continuous coverage signal for comparison against control background.
- ▌ Cooler File Loading And Querying · holobiomicslabUse when your Hi-C data is stored in cooler format (a binary HDF5-based sparse matrix with associated genomic bins and genomic tracks); you need to programmatically access the contact matrix, bin coordinates, or track data (e.g., eigenvectors, GC content) for further analysis;
- ▌ Epigenetic Sample Stratification · holobiomicslabUse when after merging methylation call files into a unified methylBase object (covering all samples at common base positions), apply this skill to assess whether biological replicates cluster together, whether case/control or treatment groups separate as expected, and to identify potential sample.
- ▌ Genomic Loop Call Interpretation · holobiomicslabUse when you have a pre-generated .hic contact map file and need to identify and annotate chromatin loops or topologically associating domains (TADs) at high resolution.
- ▌ Local Background Bias Estimation · holobiomicslabUse when you have paired ChIP and control BED/BEDPE files and need to account for local sequencing bias before peak calling. Use it specifically when control signal varies across genomic regions at multiple spatial scales (e.
- ▌ Poisson Test Statistical Scoring · holobiomicslabUse when after generating ChIP pileup and local lambda (background) BEDGRAPH tracks with matched sequencing depth, use this skill to assign statistical significance scores to each genomic region.
- ▌ Probe Detection Pvalue Filtering · holobiomicslabUse when immediately after loading raw methylation array data (.idat files or beta-valued matrix) from HumanMethylation450 (450k) or EPIC arrays when conducting primary quality control.
- ▌ Random Matrix Theory Application · holobiomicslabUse when when analyzing normalized DNA methylation beta matrices (450K or EPIC arrays) and you need to identify the true number of latent batch or technical factors present in the data.
- ▌ Adduct Mass Difference Matching · holobiomicslabUse when you have computed a histogram of pairwise mass differences from MS imaging data and need to (1) identify which observed mass differences correspond to biologically relevant or chemically known adducts, or (2) rank the most frequently observed mass differences to discover dominant adduct.
- ▌ Batch Corrected Data Extraction · holobiomicslabUse when after batch correction has been applied to metabolomics data using pooled SQC samples, and you need to retrieve the corrected ratios (compound / internal standard) for quality metrics calculation, internal standard recommendation, concentration estimation, or statistical modelling.
- ▌ Biclustering For Omics Features · holobiomicslabUse when you have a normalized matrix of feature attribution scores (microbes × metabolites) derived from a trained neural network, and you want to partition both microbes and metabolites simultaneously into co-clusters that share similar interaction patterns.
- ▌ Biotransformation Rule Encoding · holobiomicslabUse when you have untargeted metabolomics data with unknown metabolite structures and need to generate plausible candidate products by systematically applying known enzymatic or chemical transformation rules.
- ▌ Chromatographic Method Transfer · holobiomicslabUse when you have retention time predictions from a source chromatographic method and need to predict retention times for a target chromatographic method, but have limited calibration data (10–100 molecules) measured on both methods.
- ▌ Classifier Prediction Inference · holobiomicslabUse when you have a CSV or Excel file containing chemical structure descriptors for one or more molecules, and you want to obtain binary bitter/not-bitter predictions for each molecule using the BitterPredict classifier.
- ▌ Confidence Score Interpretation · holobiomicslabUse when after executing forward inference on preprocessed mass spectrometry spectra with a deep learning model (e.g., PS²MS), when you have per-spectrum predictions with associated confidence scores or per-class probabilities.
- ▌ Cross Language Interface Design · holobiomicslabUse when when you have domain-specific functionality (e.g., spectral similarity scoring, peak detection algorithms) implemented in one language (Python) but need to make it callable and composable within an R-based analytical pipeline (Spectra objects);
- ▌ Directive Engine Implementation · holobiomicslabUse when when you have intermediate JSON data that must be selectively transformed or enriched according to declarative conversion rules—for example, when extracting experimental metadata from tabular spreadsheets, you need to map certain fields to computed or filtered values, apply conditional.
- ▌ Distribution Channel Validation · holobiomicslabUse when when preparing a software release, testing contribution workflows, or auditing package availability: verify that matchms can be installed and imported successfully from all advertised distribution channels (PyPI and Bioconda) to confirm the package metadata, dependencies, and entry points.
- ▌ Fdr Correction Multiple Testing · holobiomicslabUse when you have computed empirical p-values from randomized sampling (e.
- ▌ Functional Group Classification · holobiomicslabUse when you have a set of query chemicals (chemical names or structures) and need to match them against a reference chemical library organized by type or category, with the goal of identifying structural similarity, functional group membership, or categorical assignment.
- ▌ Gradient Based Saliency Mapping · holobiomicslabUse when you have a trained graph neural network model for CCS prediction and need to identify which molecular structural features drive individual predictions or systematic biases.
- ▌ Igzip Header Structure Encoding · holobiomicslabUse when when implementing an igzip parser, decoder, or validator that must interpret the custom header format; when debugging igzip file corruption or encoding errors; or when extending pymzML's igzip support to handle new index schemes.
- ▌ Ion Mobility Reference Matching · holobiomicslabUse when you have raw arrival-time data from TWIM-MS and need to convert it to collision cross section (CCS) values for multi-omic analysis.
- ▌ JSON Schema Output Verification · holobiomicslabUse when after extracting tabular data into intermediate JSON form or after applying matrix conversion directives (e.
- ▌ Lipid Class Coverage Assessment · holobiomicslabUse when when you have acquired a CCS reference library (such as DTCCSN2 for U13C labeled lipids) and need to verify that it contains the expected lipid classes, CCS values are physically plausible for ion mobility data, and coverage matches the library's advertised documentation before using it.
- ▌ Lipid Class Stratified Analysis · holobiomicslabUse when you have IM-MS lipidomics data with measured CCS values, samples spiked with U13C-labeled lipid internal standards (e.
- ▌ Lipid Type Category Enumeration · holobiomicslabUse when when you have downloaded or cloned a lipidomics library repository (such as LipidMatch) and need to audit the breadth of lipid-type coverage to ensure the library meets minimum requirements for your analysis scope (e.g., ≥60 distinct lipid categories).
- ▌ Mass2motif Network Construction · holobiomicslabUse when after MS2LDA has inferred a motifset and you need to visualize and export the relationships between discovered Mass2Motifs for post-processing exploration, comparative annotation, or integration with external tools. Use this skill when you have motifset.json or motifset_optimized.
- ▌ Molecular Structure Attribution · holobiomicslabUse when you have a trained GNN model predicting CCS values from molecular graphs and need to understand which structural features (node and edge attributes) are most influential for specific predictions or across a test set.
- ▌ Multi Key Sorting And Filtering · holobiomicslabUse when when converting tabular data to JSON via the matrix directive, you need to exclude records that fail domain-specific validation (e.g., only retain records where a 'test' condition evaluates to true) and/or reorder the output list by one or more fields in ascending or descending order.
- ▌ Package Repository Verification · holobiomicslabUse when releasing a new version of a Python package to public repositories, when verifying that distribution pipelines are functioning after code changes, or when troubleshooting installation failures reported by users across different platforms (Linux, macOS) or architectures (x86_64, aarch64).
- ▌ Peak Removal Robustness Testing · holobiomicslabUse when when validating a metabolomics pathway analysis method (particularly decomposition-based approaches like PLAGE) against data quality degradation, or when comparing robustness across methods (PLAGE vs. ORA vs. GSEA).
- ▌ Post Hoc Model Interpretability · holobiomicslabUse when after training a GNN model on molecular structures with continuous targets (e.g., CCS values), when you need to understand which node-level (atom) or edge-level (bond) features contribute most to individual or aggregate predictions.
- ▌ Ppm Tolerance Window Adjustment · holobiomicslabUse when a mass spectrum calibration procedure initialized with a narrow ppm window (e.g., ±1.0 or ±5.0 ppm) finds fewer than 5 reference m/z matches.
- ▌ Prediction Sensitivity Analysis · holobiomicslabUse when you have a trained predictor (like BitterPredict) that accepts structured descriptors and want to understand feature importance without retraining.
- ▌ Python Automated Test Execution · holobiomicslabUse when when contributing code changes to a Python project (fork, feature branch, or pull request) that uses a setup.py-based test suite, before pushing changes to the remote repository or merging into the main branch.
- ▌ Rank Order Correlation Analysis · holobiomicslabUse when when you have run a pathway ranking method (such as PALS/PLAGE) on clean metabolomics data and wish to assess how sensitive the resulting pathway activity rankings are to realistic data quality issues—specifically Gaussian noise and random peak dropout—which are prevalent in untargeted.
- ▌ REST API Contract Documentation · holobiomicslabUse when a webservice component (like MAGMa's joblauncher) lacks formal API documentation but the source code is accessible, and downstream consumers (web applications, external services) need to understand available HTTP endpoints, parameter schemas, and response formats without manual.
- ▌ Service Health Check Definition · holobiomicslabUse when when deploying a multi-service microarchitecture (such as MAGMa's four distinct subprojects: magmaweb, joblauncher, job, and pubchem) via Docker Compose and you need to ensure that service startup order respects true readiness rather than container existence—particularly when services have.
- ▌ Spectral Data Format Validation · holobiomicslabUse when after converting existing mass spectrometry formats (mzML, vendor formats) into mzPeak using command-line tools or when receiving mzPeak files from external sources.
- ▌ Spectrum Subsetting And Merging · holobiomicslabUse when when you have a large MsBackend object and need to (1) select a contiguous or non-contiguous range of spectra for focused analysis, or (2) combine spectra from multiple independently-loaded backends (e.
- ▌ Structure Organism Pair Binning · holobiomicslabUse when when you have loaded a structure-organism pairs table from a natural products database (e.g., LOTUS) and need to answer questions about the distribution of chemical diversity—specifically, how many unique 2D structures appear in exactly 1 organism versus many organisms.
- ▌ Binary Spectral Data Extraction · holobiomicslabUse when you have parsed imzML XML metadata and loaded the corresponding .ibd binary intensity file, and need to extract specific ion images at one or more target m/z values.
- ▌ Bioinformatic Object Conversion · holobiomicslabUse when when you have processed Cardinal MSI data (normalized peak intensities, optional SSC segmentation results) and need to transition to Seurat-based workflows for differential expression, pathway analysis, or integration with spatial transcriptomics data.
- ▌ Blood Sample Handling Protocols · holobiomicslabUse when you are planning a blood sampling campaign and need to verify that lipids or polar metabolites of interest will remain stable through your anticipated sample-processing workflow (specific temperature regimes, delays before/after centrifugation).
- ▌ Centrality Metric Normalization · holobiomicslabUse when after computing raw betweenness centrality scores for metabolites in a bipartite pathway-metabolite igraph network, when you need to identify metabolites with high topological influence for visualization in ranked plots or when comparing centrality across multiple metabolite subsets or.
- ▌ Chromatographic Peak Processing · holobiomicslabUse when after peak detection when you have a table of detected peaks with m/z values and retention times from LC/HRMS data, and you observe systematic m/z drift across a batch or population-scale study (n > 500 samples).
- ▌ Complementary Score Integration · holobiomicslabUse when when you have scored a set of potential gene cluster family (GCF)–molecular feature (MF) links using two or more orthogonal methods (e.
- ▌ Compound Identifier Aggregation · holobiomicslabUse when you have metabolomics results from multiple studies with compound identifiers in heterogeneous formats (names, InChI strings, SMILES, ChEBI/KEGG/HMDB codes) and need to merge datasets for vote-counting or meta-analysis. Apply this skill before computing consensus measures (e.
- ▌ Configuration Schema Validation · holobiomicslabUse when when you have generated or edited a LipoCLEAN `options.txt` file using `--print MSD4` or `--print MSD5` and need to verify it is well-formed before running the analysis. Use this skill as a pre-flight check before invoking `--options options.
- ▌ Consensus Scoring Meta Analysis · holobiomicslabUse when when you have harmonized metabolite data from multiple studies with identifier, fold-change direction, and trend classification columns available, but lack standard deviations or variance estimates needed for quantitative meta-analysis.
- ▌ Cytoscape Network Format Export · holobiomicslabUse when after constructing or filtering a metabolomic network in MetaMapR (e.g., after applying the unique-edges hierarchy filter to resolve multiple edge types between node pairs), and you need to visualize, share, or further analyze the network using cytoscape.js or compatible tools.
- ▌ Data Format Compliance Checking · holobiomicslabUse when mSMetaEnhancer fetches metadata from external services (CIR, CTS, PubChem, IDSM, BridgeDb) and must write enriched annotations (SMILES, InChI, CAS numbers, formulas, inchikeys, IUPAC names) into .msp files.
- ▌ Differential Expression Ranking · holobiomicslabUse when you have completed a two-group or multi-group differential expression analysis with computed logFC values (e.
- ▌ Docker Image Layer Optimization · holobiomicslabUse when you are building a Docker image for a .NET Framework 4.8 application (like AirdPro) that must run on both Windows native containers and Linux+Wine environments, and you need to reduce build time and image size while maintaining separate targets for GUI, CLI, and development modes.
- ▌ Dotnet Assembly Path Resolution · holobiomicslabUse when your R package wraps a .NET assembly (e.g., RawFileReader) and you need to (1) report to users or logs which assembly file is actually being used, (2) verify that the assembly exists on disk before attempting to invoke it via system calls, or (3) enable diagnostic output showing the exact.
- ▌ Drift Correction Across Batches · holobiomicslabUse when your preprocessed metabolomics matrix (log2-scaled, with rows as features and columns as samples) exhibits batch-dependent signal drift or technical variation across multiple analytical runs or instrument sessions.
- ▌ Eawag Database Rule Application · holobiomicslabUse when when you have small-molecule structures (SMILES or structure format) and need to predict their environmental biotransformation pathways under microbial degradation conditions, particularly when the degradation rules must be drawn from curated, experimentally-validated biodegradation.
- ▌ Enrichment Score Interpretation · holobiomicslabUse when after differential expression analysis has produced gene lists with p-values and log fold-change values, and after pathway enrichment tools (clusterProfiler or biotranslator) have computed enrichment scores and adjusted p-values.
- ▌ Feature Table Format Conversion · holobiomicslabUse when you have raw feature tables exported from NPP tools (XCMS, MZmine 2, MS-DIAL, OpenMS, etc.) in their native formats and need to compare their peak detection and alignment performance against a mzRAPP benchmark dataset.
- ▌ Fold Change And P Value Ranking · holobiomicslabUse when after performing differential abundance testing (via limma or edgeR) on batch-corrected lipid abundance data, use this skill to rank and filter lipids when you need to prioritize results by both effect size and statistical confidence, especially in complex experimental designs with.
- ▌ Fold Change Log2 Transformation · holobiomicslabUse when when preparing fold-change measurements from multiple metabolomics studies for quantitative meta-analysis via weighted averaging.
- ▌ Force Field Minimization Mmff94 · holobiomicslabUse when after RDKit generates multiple 3D conformers from ionized molecular structures using distance-geometry embedding, before filtering with ASE-ANI or submitting to quantum calculations.
- ▌ Functional Orthology Annotation · holobiomicslabUse when after assigning hierarchical KEGG identifiers to a metabolomics count data frame (via assign_hierarchy with identifier='KEGG'), use this skill when you need to link metabolites to their functional roles (KO numbers) and associated genes for downstream functional analysis, pathway.