» Articles » PMID: 25155798

Calibration of Risk Prediction Models: Impact on Decision-analytic Performance

Overview
Publisher Sage Publications
Date 2014 Aug 27
PMID 25155798
Citations 141
Authors
Affiliations
Soon will be listed here.
Abstract

Decision-analytic measures to assess clinical utility of prediction models and diagnostic tests incorporate the relative clinical consequences of true and false positives without the need for external information such as monetary costs. Net Benefit is a commonly used metric that weights the relative consequences in terms of the risk threshold at which a patient would opt for treatment. Theoretical results demonstrate that clinical utility is affected by a model';s calibration, the extent to which estimated risks correspond to observed event rates. We analyzed the effects of different types of miscalibration on Net Benefit and investigated whether and under what circumstances miscalibration can make a model clinically harmful. Clinical harm is defined as a lower Net Benefit compared with classifying all patients as positive or negative by default. We used simulated data to investigate the effect of overestimation, underestimation, overfitting (estimated risks too extreme), and underfitting (estimated risks too close to baseline risk) on Net Benefit for different choices of the risk threshold. In accordance with theory, we observed that miscalibration always reduced Net Benefit. Harm was sometimes observed when models underestimated risk at a threshold below the event rate (as in underestimation and overfitting) or overestimated risk at a threshold above event rate (as in overestimation and overfitting). Underfitting never resulted in a harmful model. The impact of miscalibration decreased with increasing discrimination. Net Benefit was less sensitive to miscalibration for risk thresholds close to the event rate than for other thresholds. We illustrate these findings with examples from the literature and with a case study on testicular cancer diagnosis. Our findings strengthen the importance of obtaining calibrated risk models.

Citing Articles

Cause-specific mortality after spousal bereavement in a Danish register-based cohort.

Sloth M, Hruza J, Mortensen L, Bhatt S, Katsiferis A Sci Rep. 2025; 15(1):6240.

PMID: 39979402 PMC: 11842573. DOI: 10.1038/s41598-025-90657-1.


Against reflexive recalibration: towards a causal framework for addressing miscalibration.

Swaminathan A, Srivastava U, Tu L, Lopez I, Shah N, Vickers A Diagn Progn Res. 2025; 9(1):4.

PMID: 39930530 PMC: 11812191. DOI: 10.1186/s41512-024-00184-2.


An Explainable Artificial Intelligence Text Classifier for Suicidality Prediction in Youth Crisis Text Line Users: Development and Validation Study.

Thomas J, Lucht A, Segler J, Wundrack R, Miche M, Lieb R JMIR Public Health Surveill. 2025; 11:e63809.

PMID: 39879608 PMC: 11822322. DOI: 10.2196/63809.


The mCHEST Score for Incident Atrial Fibrillation: MESA (Multi-Ethnic Study of Atherosclerosis).

Li Y, Li Q, Wang L, Zhang T, Gao H, Pastori D JACC Adv. 2025; 4(2):101521.

PMID: 39877666 PMC: 11773033. DOI: 10.1016/j.jacadv.2024.101521.


The Harms of Class Imbalance Corrections for Machine Learning Based Prediction Models: A Simulation Study.

Carriero A, Luijken K, de Hond A, Moons K, Van Calster B, van Smeden M Stat Med. 2025; 44(3-4):e10320.

PMID: 39865585 PMC: 11771573. DOI: 10.1002/sim.10320.