Uncovering informative content in metabolomics data : from pre-processing of 1H NMR spectra to biomarkers discovery in multifactorial designs

(2020)

Files

Thesis_MMartin.pdf
  • Open Access
  • Adobe PDF
  • 6.28 MB

Details

Authors
Supervisors
Govaerts, Bernadette
Abstract
Metabolomics is a quite young and multidisciplinary science at the interface between life sciences and statistics/bioinformatics; and NMR data, one of its pillar. The objectives of this thesis are threefold: (1) provide a clearer assessment on the existing tools, (2) suggest innovative methodologies and processing workflows for topics of current interest in metabolomics, and (3) offer online open source resources to researchers interested in this field. First, a new R package, called PepsNMR, to obtain interpretable 1H NMR spectral profiles from FID raw data has been published in Bioconductor and is integrated to a web platform dedicated to metabolomics data treatment (Workflow4Metabolomics). It involves common and innovative non- or semi-parametric methods to deal with artefacts and sources of unwanted variability. After a series of improvements and validation steps, the workflow is successfully compared to a conventional pre-processing based on spectral repeatability measurements. Then, the classic and sparse (O)PLS algorithms are presented, explained in a unified fashion and further compared to each other for feature selection in an innovative evaluation scheme with the resampling of a dataset where the exact location of the true biomarkers is known in advance, and under different and realistic conditions. Finally, the use of ASCA in metabolomics is emphasized. ASCA and global methods (AComDim, PARAFASCA, AMOPLS) are presented under a common framework, with enhancements and extensions to general linear models, and are compared to each other extensively based on an experimental dataset. Then, ASCA is successfully extended to linear mixed models to incorporate random effects. New suggestions for each step are made to appropriately compare and test the fixed and random sources of variability. Except for the pre-processing, the suggested developments can be more generally applied to other datasets with multicollinearity and low (high) number of samples (variables).
Affiliations

Citations

Martin, M. (2020). Uncovering informative content in metabolomics data : from pre-processing of 1H NMR spectra to biomarkers discovery in multifactorial designs. https://hdl.handle.net/2078.5/212518