» Articles » PMID: 40038746

EnSCAN: ENsemble Scoring for Prioritizing CAusative VariaNts Across Multiplatform GWASs for Late-onset Alzheimer's Disease

Overview
Journal BioData Min
Publisher Biomed Central
Specialty Biology
Date 2025 Mar 4
PMID 40038746
Authors
Affiliations
Soon will be listed here.
Abstract

Late-onset Alzheimer's disease (LOAD) is a progressive and complex neurodegenerative disorder of the aging population. LOAD is characterized by cognitive decline, such as deterioration of memory, loss of intellectual abilities, and other cognitive domains resulting from due to traumatic brain injuries. Alzheimer's Disease (AD) presents a complex genetic etiology that is still unclear, which limits its early or differential diagnosis. The Genome-Wide Association Studies (GWAS) enable the exploration of individual variants' statistical interactions at candidate loci, but univariate analysis overlooks interactions between variants. Machine learning (ML) algorithms can capture hidden, novel, and significant patterns while considering nonlinear interactions between variants to understand the genetic predisposition for complex genetic disorders. When working on different platforms, majority voting cannot be applied because the attributes differ. Hence, a new post-ML ensemble approach was developed to select significant SNVs via multiple genotyping platforms. We proposed the EnSCAN framework using a new algorithm to ensemble selected variants even from different platforms to prioritize candidate causative loci, which consequently helps improve ML results by combining the prior information captured from each dataset. The proposed ensemble algorithm utilizes the chromosomal locations of SNVs by mapping to cytogenetic bands, along with the proximities between pairs and multimodel Random Forest (RF) validations to prioritize SNVs and candidate causative genes for LOAD. The scoring method is scalable and can be applied to any multiplatform genotyping study. We present how the proposed EnSCAN scoring algorithm prioritizes candidate causative variants related to LOAD among three GWAS datasets.

References
1.
Saez-Orellana F, Octave J, Pierrot N . Alzheimer's Disease, a Lipid Story: Involvement of Peroxisome Proliferator-Activated Receptor α. Cells. 2020; 9(5). PMC: 7290654. DOI: 10.3390/cells9051215. View

2.
Greene C, Krishnan A, Wong A, Ricciotti E, Zelaya R, Himmelstein D . Understanding multicellular function and disease with human tissue-specific networks. Nat Genet. 2015; 47(6):569-76. PMC: 4828725. DOI: 10.1038/ng.3259. View

3.
Bagyinszky E, Youn Y, An S, Kim S . The genetics of Alzheimer's disease. Clin Interv Aging. 2014; 9:535-51. PMC: 3979693. DOI: 10.2147/CIA.S51571. View

4.
Byeon H . Is the Random Forest Algorithm Suitable for Predicting Parkinson's Disease with Mild Cognitive Impairment out of Parkinson's Disease with Normal Cognition?. Int J Environ Res Public Health. 2020; 17(7). PMC: 7178031. DOI: 10.3390/ijerph17072594. View

5.
Goldstein B, Polley E, Briggs F . Random forests for genetic association studies. Stat Appl Genet Mol Biol. 2012; 10(1):32. PMC: 3154091. DOI: 10.2202/1544-6115.1691. View