» Articles » PMID: 26685993

Reverse Engineering and Evaluation of Prediction Models for Progression to Type 2 Diabetes: An Application of Machine Learning Using Electronic Health Records

Overview
Specialty Endocrinology
Date 2015 Dec 22
PMID 26685993
Citations 43
Authors
Affiliations
Soon will be listed here.
Abstract

Background: Application of novel machine learning approaches to electronic health record (EHR) data could provide valuable insights into disease processes. We utilized this approach to build predictive models for progression to prediabetes and type 2 diabetes (T2D).

Methods: Using a novel analytical platform (Reverse Engineering and Forward Simulation [REFS]), we built prediction model ensembles for progression to prediabetes or T2D from an aggregated EHR data sample. REFS relies on a Bayesian scoring algorithm to explore a wide model space, and outputs a distribution of risk estimates from an ensemble of prediction models. We retrospectively followed 24 331 adults for transitions to prediabetes or T2D, 2007-2012. Accuracy of prediction models was assessed using an area under the curve (AUC) statistic, and validated in an independent data set.

Results: Our primary ensemble of models accurately predicted progression to T2D (AUC = 0.76), and was validated out of sample (AUC = 0.78). Models of progression to T2D consisted primarily of established risk factors (blood glucose, blood pressure, triglycerides, hypertension, lipid disorders, socioeconomic factors), whereas models of progression to prediabetes included novel factors (high-density lipoprotein, alanine aminotransferase, C-reactive protein, body temperature; AUC = 0.70).

Conclusions: We constructed accurate prediction models from EHR data using a hypothesis-free machine learning approach. Identification of established risk factors for T2D serves as proof of concept for this analytical approach, while novel factors selected by REFS represent emerging areas of T2D research. This methodology has potentially valuable downstream applications to personalized medicine and clinical research.

Citing Articles

Development and validation of a chronic kidney disease progression model using patient-level simulations.

Ramos M, Gerlier L, Uster A, Muttram L, Steubl D, Frankel A Ren Fail. 2024; 46(2):2406402.

PMID: 39431558 PMC: 11494709. DOI: 10.1080/0886022X.2024.2406402.


Enhancing severe hypoglycemia prediction in type 2 diabetes mellitus through multi-view co-training machine learning model for imbalanced dataset.

Agraz M, Deng Y, Karniadakis G, Mantzoros C Sci Rep. 2024; 14(1):22741.

PMID: 39349500 PMC: 11444036. DOI: 10.1038/s41598-024-69844-z.


Machine learning-based evaluation of prognostic factors for mortality and relapse in patients with acute lymphoblastic leukemia: a comparative simulation study.

Mehrbakhsh Z, Hassanzadeh R, Behnampour N, Tapak L, Zarrin Z, Khazaei S BMC Med Inform Decis Mak. 2024; 24(1):261.

PMID: 39285373 PMC: 11404043. DOI: 10.1186/s12911-024-02645-6.


PyCaret for Predicting Type 2 Diabetes: A Phenotype- and Gender-Based Approach with the "Nurses' Health Study" and the "Health Professionals' Follow-Up Study" Datasets.

Gul S, Ayturan K, Hardalac F J Pers Med. 2024; 14(8).

PMID: 39201996 PMC: 11355927. DOI: 10.3390/jpm14080804.


Predictive modeling of multi-class diabetes mellitus using machine learning and filtering iraqi diabetes data dynamics.

Sahid M, Babar M, Uddin M PLoS One. 2024; 19(5):e0300785.

PMID: 38753669 PMC: 11098411. DOI: 10.1371/journal.pone.0300785.


References
1.
Marchesini G, Brizi M, Bianchi G, Tomassetti S, Bugianesi E, Lenzi M . Nonalcoholic fatty liver disease: a feature of the metabolic syndrome. Diabetes. 2001; 50(8):1844-50. DOI: 10.2337/diabetes.50.8.1844. View

2.
Vozarova B, Stefan N, Lindsay R, Saremi A, Pratley R, Bogardus C . High alanine aminotransferase is associated with decreased hepatic insulin sensitivity and predicts the development of type 2 diabetes. Diabetes. 2002; 51(6):1889-95. DOI: 10.2337/diabetes.51.6.1889. View

3.
Wilson P, Meigs J, Sullivan L, Fox C, Nathan D, DAgostino Sr R . Prediction of incident diabetes mellitus in middle-aged adults: the Framingham Offspring Study. Arch Intern Med. 2007; 167(10):1068-74. DOI: 10.1001/archinte.167.10.1068. View

4.
Dehghan A, van Hoek M, Sijbrands E, Stijnen T, Hofman A, Witteman J . Risk of type 2 diabetes attributable to C-reactive protein and other risk factors. Diabetes Care. 2007; 30(10):2695-9. DOI: 10.2337/dc07-0348. View

5.
Sanchez-Alavez M, Tabarean I, Osborn O, Mitsukawa K, Schaefer J, Dubins J . Insulin causes hyperthermia by direct inhibition of warm-sensitive neurons. Diabetes. 2009; 59(1):43-50. PMC: 2797943. DOI: 10.2337/db09-1128. View