Computational Metabolomics At Scale From Open Data To Insight
Computational metabolomics at scale: from open data to insight
Terms
- endogenous: produced by an organism (amino acids etc)
- exogenous: not naturally produced by the organism (environmental contiminants, etc)
- FAIR: Findablility, Accessibility, Interoperability, Reusability
Introduction
Metabolomics measures endogenous and exogenous molecules smaller than 1500 Daltons in order th better understand biological systems. Collection is mature but integration and interprability are still in need of further development.
Experts assert that there are 4 pillars for maintaining the metabolomics data lifecycle: 1) public data availability 2) open data standards 3) data and knowledge integration 4) education Public data can be valuable fo the format becomes the issue. The FAIR principles have a guiding framework for aligning data lifecycles with open science goals.
Availability, usability, and utility of public data and benchmark datasets
There are many public datasets available online, from places like:
- MetaboLights
- Metabolomics Workbench
- GNPS / MassIVE
- MetaboBank
- KMAP The amount of MS1 data continues to increase. However the metadata is not increasing at the same rate/quality.
Data reuse depends on:
- visibility
- availability
- annotations of metadata Reasons for lack of connectivity of data include:
- software producers
- data standards Public data works well as a benchmark to assess other new data and reproducibility of work.
Open data standards and knowledge sources advance metabolomics analyses
Open standards improve interoperability. Thus open standards can contribute to reproducibility in the field, and allow for new findings with cross-analysis.
In metabolomics there can be different ID types, and therefore need to have cross-linking betweeen ID types, their example is BridgeDb, MetLinkR.
Standardization also leads to better interop for ML and statistical purposes.
Across institutions, tools, companies, standards vary widely and thus following of and implementing an open standard is clearly needed.
^ refer to table two for a good list of their open standards in metabolomics
Toward integrating data and knoweldge at scale
Thee is a growing emphasis for adding prior bio knowledge to data collection, such as compound databases, better meta-data, etc.