Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection

Since the turn of the century, technological advances have made it possible to obtain the molecular profile of any tissue in a cost-effective manner. Among these advances are sophisticated high-throughput assays that measure the relative abundances of microorganisms, RNA molecules, and metabolites....

Full description

Bibliographic Details
Authors: Quinn, Thomas P., Erb, Ionas
Format: article
Status:Published version
Publication Date:2020
Country:España
Institution:Universitat Pompeu Fabra
Repository:Repositorio Digital de la UPF
OAI Identifier:oai:repositori.upf.edu:10230/44491
Online Access:http://hdl.handle.net/10230/44491
http://dx.doi.org/10.1128/mSystems.00230-19
Access Level:Open access
Keyword:Balances
Classification
Coda
Compositional data
Log contrast
Log ratio
Machine learning
Microbiome
Prediction
id ES_9e673c9b71c09f0f3b43b82aef9cbf64
oai_identifier_str oai:repositori.upf.edu:10230/44491
network_acronym_str ES
network_name_str España
spelling Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection Quinn, Thomas P. Erb, Ionas Balances Classification Coda Compositional data Log contrast Log ratio Machine learning Microbiome Prediction Since the turn of the century, technological advances have made it possible to obtain the molecular profile of any tissue in a cost-effective manner. Among these advances are sophisticated high-throughput assays that measure the relative abundances of microorganisms, RNA molecules, and metabolites. While these data are most often collected to gain new insights into biological systems, they can also be used as biomarkers to create clinically useful diagnostic classifiers. How best to classify high-dimensional -omics data remains an area of active research. However, few explicitly model the relative nature of these data and instead rely on cumbersome normalizations. This report (i) emphasizes the relative nature of health biomarkers, (ii) discusses the literature surrounding the classification of relative data, and (iii) benchmarks how different transformations perform for regularized logistic regression across multiple biomarker types. We show how an interpretable set of log contrasts, called balances, can prepare data for classification. We propose a simple procedure, called discriminative balance analysis, to select groups of 2 and 3 bacteria that can together discriminate between experimental conditions. Discriminative balance analysis is a fast, accurate, and interpretable alternative to data normalization.IMPORTANCE High-throughput sequencing provides an easy and cost-effective way to measure the relative abundance of bacteria in any environmental or biological sample. When these samples come from humans, the microbiome signatures can act as biomarkers for disease prediction. However, because bacterial abundance is measured as a composition, the data have unique properties that make conventional analyses inappropriate. To overcome this, analysts often use cumbersome normalizations. This article proposes an alternative method that identifies pairs and trios of bacteria whose stoichiometric presence can differentiate between diseased and nondiseased samples. By using interpretable log contrasts called balances, we developed an entirely normalization-free classification procedure that reduces the feature space and improves the interpretability, without sacrificing classifier performance. American Society for Microbiology http://hdl.handle.net/10230/44491 http://dx.doi.org/10.1128/mSystems.00230-19
title Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection
spellingShingle Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection
Quinn, Thomas P.
Balances
Classification
Coda
Compositional data
Log contrast
Log ratio
Machine learning
Microbiome
Prediction
title_short Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection
title_full Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection
title_fullStr Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection
title_full_unstemmed Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection
title_sort Interpretable log contrasts for the classification of health biomarkers: a new approach to balance selection
author Quinn, Thomas P.
author_facet Quinn, Thomas P.
Erb, Ionas
author_role author
author2 Erb, Ionas
author2_role author
topic Balances
Classification
Coda
Compositional data
Log contrast
Log ratio
Machine learning
Microbiome
Prediction
topic_facet Balances
Classification
Coda
Compositional data
Log contrast
Log ratio
Machine learning
Microbiome
Prediction
description Since the turn of the century, technological advances have made it possible to obtain the molecular profile of any tissue in a cost-effective manner. Among these advances are sophisticated high-throughput assays that measure the relative abundances of microorganisms, RNA molecules, and metabolites. While these data are most often collected to gain new insights into biological systems, they can also be used as biomarkers to create clinically useful diagnostic classifiers. How best to classify high-dimensional -omics data remains an area of active research. However, few explicitly model the relative nature of these data and instead rely on cumbersome normalizations. This report (i) emphasizes the relative nature of health biomarkers, (ii) discusses the literature surrounding the classification of relative data, and (iii) benchmarks how different transformations perform for regularized logistic regression across multiple biomarker types. We show how an interpretable set of log contrasts, called balances, can prepare data for classification. We propose a simple procedure, called discriminative balance analysis, to select groups of 2 and 3 bacteria that can together discriminate between experimental conditions. Discriminative balance analysis is a fast, accurate, and interpretable alternative to data normalization.IMPORTANCE High-throughput sequencing provides an easy and cost-effective way to measure the relative abundance of bacteria in any environmental or biological sample. When these samples come from humans, the microbiome signatures can act as biomarkers for disease prediction. However, because bacterial abundance is measured as a composition, the data have unique properties that make conventional analyses inappropriate. To overcome this, analysts often use cumbersome normalizations. This article proposes an alternative method that identifies pairs and trios of bacteria whose stoichiometric presence can differentiate between diseased and nondiseased samples. By using interpretable log contrasts called balances, we developed an entirely normalization-free classification procedure that reduces the feature space and improves the interpretability, without sacrificing classifier performance.
publishDate 2020
format article
status_str publishedVersion
url http://hdl.handle.net/10230/44491
http://dx.doi.org/10.1128/mSystems.00230-19
eu_rights_str_mv openAccess
publisher American Society for Microbiology
institution Universitat Pompeu Fabra
collection Repositorio Digital de la UPF
reponame_str Repositorio Digital de la UPF
instname_str Universitat Pompeu Fabra
_version_ 1878438532230414336
publishDateSort 2020
author_browse Erb, Ionas
Quinn, Thomas P.
publisherStr American Society for Microbiology
score 6,9303427