Format

Send to

Choose Destination
Methods Mol Biol. 2014;1198:333-53. doi: 10.1007/978-1-4939-1258-2_22.

Statistical analysis and modeling of mass spectrometry-based metabolomics data.

Author information

1
Department of Statistics, Purdue University, 250 North University Street, West Lafayette, IN, 47907, USA, xbw@purdue.edu.

Abstract

Multivariate statistical techniques are used extensively in metabolomics studies, ranging from biomarker selection to model building and validation. Two model independent variable selection techniques, principal component analysis and two sample t-tests are discussed in this chapter, as well as classification and regression models and model related variable selection techniques, including partial least squares, logistic regression, support vector machine, and random forest. Model evaluation and validation methods, such as leave-one-out cross-validation, Monte Carlo cross-validation, and receiver operating characteristic analysis, are introduced with an emphasis to avoid over-fitting the data. The advantages and the limitations of the statistical techniques are also discussed in this chapter.

PMID:
25270940
PMCID:
PMC4319703
DOI:
10.1007/978-1-4939-1258-2_22
[Indexed for MEDLINE]
Free PMC Article

Supplemental Content

Full text links

Icon for Springer Icon for PubMed Central
Loading ...
Support Center