Amb motiu del tancament d'estiu, la validació de documents es reprendrà a partir del 28 d'agost de 2026. Disculpeu les molèsties.
Con motivo del cierre de verano, la validación de documentos se reanudará a partir del 28 de agosto de 2026. Disculpad las molestias
Due to the summer closure, document validation will resume starting August 28, 2026. We apologize for any inconvenience.

Document type

Article

Version

Published version

Publication date

Publication license

cc by (c) Viallon, Vivian et al, 2021
Please use this identifier to cite or link to this item: https://hdl.handle.net/2445/180519

A New Pipeline for the Normalization and Pooling of Metabolomics Data

Journal Title

Director/Tutor

Journal ISSN

Volume Title

Abstract

Pooling metabolomics data across studies is often desirable to increase the statistical power of the analysis. However, this can raise methodological challenges as several preanalytical and analytical factors could introduce differences in measured concentrations and variability between datasets. Specifically, different studies may use variable sample types (e.g., serum versus plasma) collected, treated, and stored according to different protocols, and assayed in different laboratories using different instruments. To address these issues, a new pipeline was developed to normalize and pool metabolomics data through a set of sequential steps: (i) exclusions of the least informative observations and metabolites and removal of outliers; imputation of missing data; (ii) identification of the main sources of variability through principal component partial R-square (PC-PR2) analysis; (iii) application of linear mixed models to remove unwanted variability, including samples' originating study and batch, and preserve biological variations while accounting for potential differences in the residual variances across studies. This pipeline was applied to targeted metabolomics data acquired using Biocrates AbsoluteIDQ kits in eight case-control studies nested within the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort. Comprehensive examination of metabolomics measurements indicated that the pipeline improved the comparability of data across the studies. Our pipeline can be adapted to normalize other molecular data, including biomarkers as well as proteomics data, and could be used for pooling molecular datasets, for example in international consortia, to limit biases introduced by inter-study variability. This versatility of the pipeline makes our work of potential interest to molecular epidemiologists.

Subject (English)

Citation

Citation

VIALLON, Vivian, et al. A New Pipeline for the Normalization and Pooling of Metabolomics Data. Metabolites. 2021. Vol. 11, num. 9, pags. 631. ISSN 2218-1989. [consulted: 20 of August of 2026]. Available at: https://hdl.handle.net/2445/180519

Export metadata

JSON - METS

Share record