» Articles » PMID: 31979006

Flexible Data Trimming Improves Performance of Global Machine Learning Methods in Omics-Based Personalized Oncology

Overview
Journal Int J Mol Sci
Publisher MDPI
Date 2020 Jan 26
PMID 31979006
Citations 13
Authors
Affiliations
Soon will be listed here.
Abstract

(1) Background: Machine learning (ML) methods are rarely used for an omics-based prescription of cancer drugs, due to shortage of case histories with clinical outcome supplemented by high-throughput molecular data. This causes overtraining and high vulnerability of most ML methods. Recently, we proposed a hybrid global-local approach to ML termed floating window projective separator (FloWPS) that avoids extrapolation in the feature space. Its core property is data trimming, i.e., sample-specific removal of irrelevant features. (2) Methods: Here, we applied FloWPS to seven popular ML methods, including linear SVM, nearest neighbors (kNN), random forest (RF), Tikhonov (ridge) regression (RR), binomial naïve Bayes (BNB), adaptive boosting (ADA) and multi-layer perceptron (MLP). (3) Results: We performed computational experiments for 21 high throughput gene expression datasets (41-235 samples per dataset) totally representing 1778 cancer patients with known responses on chemotherapy treatments. FloWPS essentially improved the classifier quality for all global ML methods (SVM, RF, BNB, ADA, MLP), where the area under the receiver-operator curve (ROC AUC) for the treatment response classifiers increased from 0.61-0.88 range to 0.70-0.94. We tested FloWPS-empowered methods for overtraining by interrogating the importance of different features for different ML methods in the same model datasets. (4) Conclusions: We showed that FloWPS increases the correlation of feature importance between the different ML methods, which indicates its robustness to overtraining. For all the datasets tested, the best performance of FloWPS data trimming was observed for the BNB method, which can be valuable for further building of ML classifiers in personalized oncology.

Citing Articles

Bioinformatics in Russia: history and present-day landscape.

Nawaz M, Pamirsky I, Golokhvast K Brief Bioinform. 2024; 25(6).

PMID: 39402695 PMC: 11473191. DOI: 10.1093/bib/bbae513.


Machine learning algorithms' application to predict childhood vaccination among children aged 12-23 months in Ethiopia: Evidence 2016 Ethiopian Demographic and Health Survey dataset.

Demsash A, Chereka A, Walle A, Kassie S, Bekele F, Bekana T PLoS One. 2023; 18(10):e0288867.

PMID: 37851705 PMC: 10584162. DOI: 10.1371/journal.pone.0288867.


Uniformly shaped harmonization combines human transcriptomic data from different platforms while retaining their biological properties and differential gene expression patterns.

Borisov N, Tkachev V, Simonov A, Sorokin M, Kim E, Kuzmin D Front Mol Biosci. 2023; 10:1237129.

PMID: 37745690 PMC: 10511763. DOI: 10.3389/fmolb.2023.1237129.


Machine learning for predicting accuracy of lung and liver tumor motion tracking using radiomic features.

Li G, Zhang X, Song X, Duan L, Wang G, Xiao Q Quant Imaging Med Surg. 2023; 13(3):1605-1618.

PMID: 36915317 PMC: 10006135. DOI: 10.21037/qims-22-621.


Transcriptomic Harmonization as the Way for Suppressing Cross-Platform Bias and Batch Effect.

Borisov N, Buzdin A Biomedicines. 2022; 10(9).

PMID: 36140419 PMC: 9496268. DOI: 10.3390/biomedicines10092318.


References
1.
Tkachev V, Sorokin M, Mescheryakov A, Simonov A, Garazha A, Buzdin A . FLOating-Window Projective Separator (FloWPS): A Data Trimming Tool for Support Vector Machines (SVM) to Improve Robustness of the Classifier. Front Genet. 2019; 9:717. PMC: 6341065. DOI: 10.3389/fgene.2018.00717. View

2.
Tabl A, Alkhateeb A, ElMaraghy W, Rueda L, Ngom A . A Machine Learning Approach for Identifying Gene Biomarkers Guiding the Treatment of Breast Cancer. Front Genet. 2019; 10:256. PMC: 6446069. DOI: 10.3389/fgene.2019.00256. View

3.
Ioannidis J, Hozo I, Djulbegovic B . Optimal type I and type II error pairs when the available sample size is fixed. J Clin Epidemiol. 2013; 66(8):903-910.e2. DOI: 10.1016/j.jclinepi.2013.03.002. View

4.
Korde L, Lusa L, McShane L, Lebowitz P, Lukes L, Camphausen K . Gene expression pathway analysis to predict response to neoadjuvant docetaxel and capecitabine for breast cancer. Breast Cancer Res Treat. 2009; 119(3):685-99. PMC: 5892182. DOI: 10.1007/s10549-009-0651-3. View

5.
Miller W, Larionov A, Anderson T, Evans D, Dixon J . Sequential changes in gene expression profiles in breast cancers during treatment with the aromatase inhibitor, letrozole. Pharmacogenomics J. 2010; 12(1):10-21. DOI: 10.1038/tpj.2010.67. View