# Peak Network To Metabolite Assignment

> Use when after peak clustering has produced peak network groups (ideally from the same compound) and you need to assign chemical identities.

- Skill: `holobiomicslab/peak-network-to-metabolite-assignment` (Agent Skill)
- Install (CLI): `npx skillmds@latest add holobiomicslab/peak-network-to-metabolite-assignment`
- Raw SKILL.md: https://api.skillmd.com/api/skills/holobiomicslab/peak-network-to-metabolite-assignment/raw
- Safety review: PASS (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- License: CC-BY-4.0
- Author: HolobiomicsLab (https://skillmd.com/u/holobiomicslab)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/holobiomicslab/peak-network-to-metabolite-assignment

---


# peak-network-to-metabolite-assignment

## Summary

Match identified peak networks from INADEQUATE NMR spectra to a simulated metabolite database using similarity metrics to assign metabolite identities. This skill bridges clustering output to metabolite identification by correlating experimental peak network signatures against reference spectral data.

## When to use

Apply this skill after peak clustering has produced peak network groups (ideally from the same compound) and you need to assign chemical identities. Use when you have query INADEQUATE spectra with identified peak networks and access to a simulated INADEQUATE database containing reference metabolite signatures with known spectral characteristics.

## When NOT to use

- Query spectra have not yet been clustered into peak networks—run Clustering module first
- Reference database is incomplete or lacks INADEQUATE signatures for metabolites of interest
- Raw, unpicked peaks are provided instead of pre-processed peak networks

## Inputs

- Peak network clusters from pyINETA Clustering module output
- Simulated INADEQUATE metabolite database with reference spectral signatures
- Configuration file specifying matching parameters and thresholds

## Outputs

- Matched metabolite assignment table (peak network ID → metabolite name + match score)
- Match score matrix (query networks × database metabolites)
- Filtered high-confidence metabolite identifications

## How to apply

Load the peak network clusters from the upstream Clustering module output along with a simulated INADEQUATE metabolite database containing reference spectral signatures. Calculate similarity metrics (e.g., cosine similarity or spectral correlation) between each query peak network and all database metabolite signatures. Apply a similarity threshold to filter and retain only high-confidence metabolite assignments. Generate a matched metabolite output table that links peak network identifiers to assigned metabolite names and their corresponding match scores. The threshold selection is critical: set it high enough to exclude false positives but low enough to capture true metabolites present in your sample.

## Related tools

- **PyINETA** (Provides the Matching module that implements peak network–to–database similarity calculation and metabolite assignment) — https://github.com/edisonomics/PyINETA
- **Python** (Runtime language for executing the PyINETA Matching module and similarity metric computation)

## Examples

```
python <path_to_pyineta_repo>/run_pyineta.py -c config.ini -s match -o output_dir
```

## Evaluation signals

- Output table contains one row per query peak network with non-null metabolite assignments and match scores in expected numeric range (0–1 for normalized similarity)
- Match scores for assigned metabolites exceed the configured similarity threshold; unassigned networks have no entries or scores below threshold
- Assigned metabolite identities are present in the simulated database and have valid INADEQUATE spectral signatures
- Peak network identifiers in output match identifiers from Clustering module output (schema consistency)
- Match score distribution shows clear separation between high-confidence (assigned) and low-confidence (unassigned) peak networks, indicating threshold is appropriate

## Limitations

- Matching quality depends critically on the completeness and accuracy of the simulated INADEQUATE reference database—missing or poorly characterized metabolites will not be identified
- Similarity metrics assume peak network features are comparable to database signatures; preprocessing differences (shifting, normalization) between query and database can degrade matching
- Threshold selection requires manual tuning or validation; no principled default is provided in the source documentation
- The package README notes 'No changelog found', suggesting limited version history documentation for reproducibility tracking

## Evidence

- [readme] pyINETA matches identified peak networks to a simulated INADEQUATE database of metabolites to identify metabolites present in the query INADEQUATE spectra: "matches to a simulated INADEQUATE database of metabolites to identify metabolites present in the query INADEQUATE spectra"
- [other] Calculate similarity metrics (e.g., cosine similarity or spectral correlation) between each query peak network and database metabolite signatures: "Calculate similarity metrics (e.g., cosine similarity or spectral correlation) between each query peak network and database metabolite signatures."
- [other] Filter matches using a similarity threshold to retain high-confidence metabolite assignments: "Filter matches using a similarity threshold to retain high-confidence metabolite assignments."
- [other] Generate a matched metabolite output table linking peak network identifiers, assigned metabolite names, and match scores: "Generate a matched metabolite output table linking peak network identifiers, assigned metabolite names, and match scores using the pyINETA Matching module."
- [other] Load peak network clusters from the upstream Clustering module output: "Load peak network clusters from the upstream Clustering module output."

