» Articles » PMID: 39232851

Processing Imbalanced Medical Data at the Data Level with Assisted-reproduction Data As an Example

Overview
Journal BioData Min
Publisher Biomed Central
Specialty Biology
Date 2024 Sep 5
PMID 39232851
Authors
Affiliations
Soon will be listed here.
Abstract

Objective: Data imbalance is a pervasive issue in medical data mining, often leading to biased and unreliable predictive models. This study aims to address the urgent need for effective strategies to mitigate the impact of data imbalance on classification models. We focus on quantifying the effects of different imbalance degrees and sample sizes on model performance, identifying optimal cut-off values, and evaluating the efficacy of various methods to enhance model accuracy in highly imbalanced and small sample size scenarios.

Methods: We collected medical records of patients receiving assisted reproductive treatment in a reproductive medicine center. Random forest was used to screen the key variables for the prediction target. Various datasets with different imbalance degrees and sample sizes were constructed to compare the classification performance of logistic regression models. Metrics such as AUC, G-mean, F1-Score, Accuracy, Recall, and Precision were used for evaluation. Four imbalance treatment methods (SMOTE, ADASYN, OSS, and CNN) were applied to datasets with low positive rates and small sample sizes to assess their effectiveness.

Results: The logistic model's performance was low when the positive rate was below 10% but stabilized beyond this threshold. Similarly, sample sizes below 1200 yielded poor results, with improvement seen above this threshold. For robustness, the optimal cut-offs for positive rate and sample size were identified as 15% and 1500, respectively. SMOTE and ADASYN oversampling significantly improved classification performance in datasets with low positive rates and small sample sizes.

Conclusions: The study identifies a positive rate of 15% and a sample size of 1500 as optimal cut-offs for stable logistic model performance. For datasets with low positive rates and small sample sizes, SMOTE and ADASYN are recommended to improve balance and model accuracy.

References
1.
Dablain D, Krawczyk B, Chawla N . DeepSMOTE: Fusing Deep Learning and SMOTE for Imbalanced Data. IEEE Trans Neural Netw Learn Syst. 2022; 34(9):6390-6404. DOI: 10.1109/TNNLS.2021.3136503. View

2.
Beam A, Kohane I . Big Data and Machine Learning in Health Care. JAMA. 2018; 319(13):1317-1318. DOI: 10.1001/jama.2017.18391. View

3.
Ahsan M, Siddique Z . Machine learning-based heart disease diagnosis: A systematic literature review. Artif Intell Med. 2022; 128:102289. DOI: 10.1016/j.artmed.2022.102289. View

4.
Ren Y, Wu D, Tong Y, Lopez-Defede A, Gareau S . Issue of Data Imbalance on Low Birthweight Baby Outcomes Prediction and Associated Risk Factors Identification: Establishment of Benchmarking Key Machine Learning Models With Data Rebalancing Strategies. J Med Internet Res. 2023; 25:e44081. PMC: 10267797. DOI: 10.2196/44081. View

5.
Kim K, Sohn S . Hybrid neural network with cost-sensitive support vector machine for class-imbalanced multimodal data. Neural Netw. 2020; 130:176-184. DOI: 10.1016/j.neunet.2020.06.026. View