» Articles » PMID: 37600970

Omics Data Integration in Computational Biology Viewed Through the Prism of Machine Learning Paradigms

Overview
Journal Front Bioinform
Specialty Biology
Date 2023 Aug 21
PMID 37600970
Authors
Affiliations
Soon will be listed here.
Abstract

Important quantities of biological data can today be acquired to characterize cell types and states, from various sources and using a wide diversity of methods, providing scientists with more and more information to answer challenging biological questions. Unfortunately, working with this amount of data comes at the price of ever-increasing data complexity. This is caused by the multiplication of data types and batch effects, which hinders the joint usage of all available data within common analyses. Data integration describes a set of tasks geared towards embedding several datasets of different origins or modalities into a joint representation that can then be used to carry out downstream analyses. In the last decade, dozens of methods have been proposed to tackle the different facets of the data integration problem, relying on various paradigms. This review introduces the most common data types encountered in computational biology and provides systematic definitions of the data integration problems. We then present how machine learning innovations were leveraged to build effective data integration algorithms, that are widely used today by computational biologists. We discuss the current state of data integration and important pitfalls to consider when working with data integration tools. We eventually detail a set of challenges the field will have to overcome in the coming years.

Citing Articles

A single-cell multimodal view on gene regulatory network inference from transcriptomics and chromatin accessibility data.

Loers J, Vermeirssen V Brief Bioinform. 2024; 25(5).

PMID: 39207727 PMC: 11359808. DOI: 10.1093/bib/bbae382.


scCross: a deep generative model for unifying single-cell multi-omics with seamless integration, cross-modal generation, and in silico exploration.

Yang X, Mann K, Wu H, Ding J Genome Biol. 2024; 25(1):198.

PMID: 39075536 PMC: 11285326. DOI: 10.1186/s13059-024-03338-z.

References
1.
Duren Z, Chen X, Zamanighomi M, Zeng W, Satpathy A, Chang H . Integrative analysis of single-cell genomics data by coupled nonnegative matrix factorizations. Proc Natl Acad Sci U S A. 2018; 115(30):7723-7728. PMC: 6065048. DOI: 10.1073/pnas.1805681115. View

2.
Chen H, Albergante L, Hsu J, Lareau C, Lo Bosco G, Guan J . Single-cell trajectories reconstruction, exploration and mapping of omics data with STREAM. Nat Commun. 2019; 10(1):1903. PMC: 6478907. DOI: 10.1038/s41467-019-09670-4. View

3.
Stoeckius M, Hafemeister C, Stephenson W, Houck-Loomis B, Chattopadhyay P, Swerdlow H . Simultaneous epitope and transcriptome measurement in single cells. Nat Methods. 2017; 14(9):865-868. PMC: 5669064. DOI: 10.1038/nmeth.4380. View

4.
Buenrostro J, Wu B, Litzenburger U, Ruff D, Gonzales M, Snyder M . Single-cell chromatin accessibility reveals principles of regulatory variation. Nature. 2015; 523(7561):486-90. PMC: 4685948. DOI: 10.1038/nature14590. View

5.
Li B, Zhang W, Guo C, Xu H, Li L, Fang M . Benchmarking spatial and single-cell transcriptomics integration methods for transcript distribution prediction and cell type deconvolution. Nat Methods. 2022; 19(6):662-670. DOI: 10.1038/s41592-022-01480-9. View