lipid-feature-annotation-and-sorting
License: restricted — no clear open-source license detected for the underlying tool; verify licensing before commercial use or redistribution.
Summary
Annotate quantified lipid features with chemical identity metadata (lipid class, adduct form, neutral mass, m/z) and sort them by m/z for HDF5 export, ensuring machine-readable and standardized feature organization in MSI data containers.
When to use
After quantifying ion images in LipidQMap and before exporting to HDF5 format, when you need to organize per-feature metadata (lipid ID, class, adduct, m/z, internal standard flag) into aligned datasets that can be linked to intensity data via dimension scales and sorted for reproducible retrieval.
When NOT to use
- Input features are already sorted and linked via dimension scales in an existing HDF5 container — skip re-annotation and re-sorting to avoid data duplication.
- Lipid identity metadata is not available or cannot be reliably matched to m/z values — annotation will be incomplete or erroneous; resolve identifications first.
- Export format does not require Cardinal::HDF5 compliance (e.g., flat CSV or unstructured image stack) — this annotation workflow is specific to structured HDF5 schemas.
Inputs
- quantified ion image intensities (feature-by-pixel matrix, n_features × n_pixels, float32)
- lipid feature catalog with chemical annotations (lipid ID, class, adduct, neutral mass, m/z)
- internal standard designation flags per feature (boolean or TRUE/FALSE)
Outputs
- featureData HDF5 group with datasets: feature_index (int64), feature_id (UTF-8), lipid_class (UTF-8), adduct (UTF-8), neutral_id (UTF-8), mz (float32, sorted ascending), is_standard (bool)
- dimension scale linking intensity axis 0 to featureData/feature_index
- group attribute 'columns' listing all dataset names in featureData
How to apply
For each quantified lipid feature in the dataset, construct a per-feature metadata row containing: feature_index (int64), feature_id (UTF-8 lipid identifier), lipid_class (UTF-8 class name), adduct (UTF-8 adduct formula), neutral_id (UTF-8 reference mass identifier), mz (float32 measured m/z), and is_standard (boolean). Sort all feature datasets in ascending order by m/z value to establish a consistent feature axis order. Write these datasets into the featureData HDF5 group, aligning axis 0 of the intensity array (n_features × n_pixels) with featureData rows. Attach dimension scales linking intensity axis 0 to the featureData/feature_index dataset, and add group-level metadata listing the column names. This enables Cardinal-compatible feature lookup and cross-linking with external lipid databases.
Related tools
- LipidQMap (Generates quantified ion images and lipid feature list (ID, class, adduct, m/z) for annotation and HDF5 export) — https://github.com/swinnenteam/LipidQMap
- Cardinal (Defines the HDF5 schema (Cardinal::HDF5 convention) that specifies featureData group structure and dimension scale linking) — https://cardinalmsi.org
Evaluation signals
- All feature metadata datasets (feature_index, feature_id, lipid_class, adduct, neutral_id, mz, is_standard) exist in the featureData group with correct dtypes (int64, UTF-8 strings, float32, bool).
- Feature rows are sorted in ascending order by mz; verify mz[i] ≤ mz[i+1] for all i.
- Dataset shapes align: all featureData datasets have length n_features, matching intensity matrix axis 0.
- Dimension scale is correctly attached: intensity.dims[0] references featureData/feature_index with scale labels.
- Group attribute 'columns' lists all dataset names present in featureData (e.g., ['feature_index', 'feature_id', 'lipid_class', 'adduct', 'neutral_id', 'mz', 'is_standard']).
Limitations
- Feature annotation relies on accurate lipid identification from the input catalog; mismatched m/z or erroneous lipid assignments propagate into the HDF5 output.
- Sorting by m/z alone does not account for isobaric features (same m/z, different identities); post-hoc filtering or manual curation may be needed for disambiguation.
- Sparse or incomplete metadata (e.g., missing neutral_id or unknown adduct) must be handled consistently (e.g., empty strings or 'NA' placeholders) to maintain schema validity.
- No changelog or versioning in the current LipidQMap release; feature annotation logic may change between versions without backwards-compatibility guarantees.
Evidence
- [methods] Populate the featureData group with per-feature metadata datasets (feature_index as int64, feature_id and lipid_class and adduct and neutral_id as UTF-8 strings, mz as float32 for sorting, is_standard as bool) aligned with intensity axis 0, sorted ascending by mz; add group attribute columns.: "Populate the featureData group with per-feature metadata datasets (feature_index as int64, feature_id and lipid_class and adduct and neutral_id as UTF-8 strings, mz as float32 for sorting,"
- [methods] LipidQMap writes MSI exports as HDF5 containers that follow the Cardinal::HDF5 conventions, enabling standardized storage and interchange of quantified imaging data.: "LipidQMap writes MSI exports as HDF5 containers that follow the Cardinal::HDF5 conventions, enabling standardized storage and interchange of quantified imaging data."
- [readme] Shows ion images for an easily editable list of lipids (list is read from an excel file).: "Shows ion images for an easily editable list of lipids (list is read from an excel file)."
- [readme] Each row in the Excel database represents a different species, and the file should contain the following columns: ID, Class, Neutral Formula, Adducts, M-2 Isotope, Na+ Isotope, Is standard, Standard amount (pmol / mm2), IS.: "Each row in the Excel database represents a different species, and the file should contain the following columns (the column titles need to match exactly): ID, Class, Neutral Formula, Adducts, M-2"
1---2name: lipid-feature-annotation-and-sorting3description: Use when after quantifying ion images in LipidQMap and before exporting to HDF5 format, when you need to organize per-feature metadata (lipid ID, class, adduct, m/z, internal standard flag) into aligned datasets that can be linked to intensity data via dimension scales and sorted for reproducible.4license: CC-BY-4.05---67# lipid-feature-annotation-and-sorting89> **License: restricted** — no clear open-source license detected for the underlying tool; verify licensing before commercial use or redistribution. <!-- asb-license-banner -->10## Summary1112Annotate quantified lipid features with chemical identity metadata (lipid class, adduct form, neutral mass, m/z) and sort them by m/z for HDF5 export, ensuring machine-readable and standardized feature organization in MSI data containers.1314## When to use1516After quantifying ion images in LipidQMap and before exporting to HDF5 format, when you need to organize per-feature metadata (lipid ID, class, adduct, m/z, internal standard flag) into aligned datasets that can be linked to intensity data via dimension scales and sorted for reproducible retrieval.1718## When NOT to use1920- Input features are already sorted and linked via dimension scales in an existing HDF5 container — skip re-annotation and re-sorting to avoid data duplication.21- Lipid identity metadata is not available or cannot be reliably matched to m/z values — annotation will be incomplete or erroneous; resolve identifications first.22- Export format does not require Cardinal::HDF5 compliance (e.g., flat CSV or unstructured image stack) — this annotation workflow is specific to structured HDF5 schemas.2324## Inputs2526- quantified ion image intensities (feature-by-pixel matrix, n_features × n_pixels, float32)27- lipid feature catalog with chemical annotations (lipid ID, class, adduct, neutral mass, m/z)28- internal standard designation flags per feature (boolean or TRUE/FALSE)2930## Outputs3132- featureData HDF5 group with datasets: feature_index (int64), feature_id (UTF-8), lipid_class (UTF-8), adduct (UTF-8), neutral_id (UTF-8), mz (float32, sorted ascending), is_standard (bool)33- dimension scale linking intensity axis 0 to featureData/feature_index34- group attribute 'columns' listing all dataset names in featureData3536## How to apply3738For each quantified lipid feature in the dataset, construct a per-feature metadata row containing: feature_index (int64), feature_id (UTF-8 lipid identifier), lipid_class (UTF-8 class name), adduct (UTF-8 adduct formula), neutral_id (UTF-8 reference mass identifier), mz (float32 measured m/z), and is_standard (boolean). Sort all feature datasets in ascending order by m/z value to establish a consistent feature axis order. Write these datasets into the featureData HDF5 group, aligning axis 0 of the intensity array (n_features × n_pixels) with featureData rows. Attach dimension scales linking intensity axis 0 to the featureData/feature_index dataset, and add group-level metadata listing the column names. This enables Cardinal-compatible feature lookup and cross-linking with external lipid databases.3940## Related tools4142- **LipidQMap** (Generates quantified ion images and lipid feature list (ID, class, adduct, m/z) for annotation and HDF5 export) — https://github.com/swinnenteam/LipidQMap43- **Cardinal** (Defines the HDF5 schema (Cardinal::HDF5 convention) that specifies featureData group structure and dimension scale linking) — https://cardinalmsi.org4445## Evaluation signals4647- All feature metadata datasets (feature_index, feature_id, lipid_class, adduct, neutral_id, mz, is_standard) exist in the featureData group with correct dtypes (int64, UTF-8 strings, float32, bool).48- Feature rows are sorted in ascending order by mz; verify mz[i] ≤ mz[i+1] for all i.49- Dataset shapes align: all featureData datasets have length n_features, matching intensity matrix axis 0.50- Dimension scale is correctly attached: intensity.dims[0] references featureData/feature_index with scale labels.51- Group attribute 'columns' lists all dataset names present in featureData (e.g., ['feature_index', 'feature_id', 'lipid_class', 'adduct', 'neutral_id', 'mz', 'is_standard']).5253## Limitations5455- Feature annotation relies on accurate lipid identification from the input catalog; mismatched m/z or erroneous lipid assignments propagate into the HDF5 output.56- Sorting by m/z alone does not account for isobaric features (same m/z, different identities); post-hoc filtering or manual curation may be needed for disambiguation.57- Sparse or incomplete metadata (e.g., missing neutral_id or unknown adduct) must be handled consistently (e.g., empty strings or 'NA' placeholders) to maintain schema validity.58- No changelog or versioning in the current LipidQMap release; feature annotation logic may change between versions without backwards-compatibility guarantees.5960## Evidence6162- [methods] Populate the featureData group with per-feature metadata datasets (feature_index as int64, feature_id and lipid_class and adduct and neutral_id as UTF-8 strings, mz as float32 for sorting, is_standard as bool) aligned with intensity axis 0, sorted ascending by mz; add group attribute columns.: "Populate the featureData group with per-feature metadata datasets (feature_index as int64, feature_id and lipid_class and adduct and neutral_id as UTF-8 strings, mz as float32 for sorting,"63- [methods] LipidQMap writes MSI exports as HDF5 containers that follow the Cardinal::HDF5 conventions, enabling standardized storage and interchange of quantified imaging data.: "LipidQMap writes MSI exports as HDF5 containers that follow the Cardinal::HDF5 conventions, enabling standardized storage and interchange of quantified imaging data."64- [readme] Shows ion images for an easily editable list of lipids (list is read from an excel file).: "Shows ion images for an easily editable list of lipids (list is read from an excel file)."65- [readme] Each row in the Excel database represents a different species, and the file should contain the following columns: ID, Class, Neutral Formula, Adducts, M-2 Isotope, Na+ Isotope, Is standard, Standard amount (pmol / mm2), IS.: "Each row in the Excel database represents a different species, and the file should contain the following columns (the column titles need to match exactly): ID, Class, Neutral Formula, Adducts, M-2"