Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging

Overview

Journal Proc ACM Conf Health Inference Learn (2020)

Publisher Association for Computing Machinery

Date 2020 Nov 16

PMID 33196064

Citations 87

Authors

Luke Oakden-Rayner

Jared Dunnmon

Gustavo Carneiro

Christopher Re

Affiliations

Soon will be listed here.

Abstract

Machine learning models for medical image analysis often suffer from poor performance on important subsets of a population that are not identified during training or testing. For example, overall performance of a cancer detection model may be high, but the model may still consistently miss a rare but aggressive cancer subtype. We refer to this problem as , and observe that it results from incompletely describing the meaningful variation in a dataset. While hidden stratification can substantially reduce the clinical efficacy of machine learning models, its effects remain difficult to measure. In this work, we assess the utility of several possible techniques for measuring hidden stratification effects, and characterize these effects both via synthetic experiments on the CIFAR-100 benchmark dataset and on multiple real-world medical imaging datasets. Using these measurement techniques, we find evidence that hidden stratification can occur in unidentified imaging subsets with low prevalence, low label quality, subtle distinguishing features, or spurious correlates, and that it can result in relative performance differences of over 20% on clinically important subsets. Finally, we discuss the clinical implications of our findings, and suggest that evaluation of hidden stratification should be a critical component of any machine learning deployment in medical imaging.

Citing Articles

Evaluating the pathological and clinical implications of errors made by an artificial intelligence colon biopsy screening tool.

Evans H, Sivakumar N, Bhanderi S, Graham S, Snead D, Patel A BMJ Open Gastroenterol. 2025; 12(1.

PMID: 39762071 PMC: 11749196. DOI: 10.1136/bmjgast-2024-001649.

Calibrating Multi-modal Representations: A Pursuit of Group Robustness without Annotations.

You C, Min Y, Dai W, Sekhon J, Staib L, Duncan J Proc IEEE Comput Soc Conf Comput Vis Pattern Recognit. 2024; 2024:26140-26150.

PMID: 39640960 PMC: 11620289. DOI: 10.1109/cvpr52733.2024.02470.

The risk of shortcutting in deep learning algorithms for medical imaging research.

Hill B, Koback F, Schilling P Sci Rep. 2024; 14(1):29224.

PMID: 39587148 PMC: 11589829. DOI: 10.1038/s41598-024-79838-6.

A data-driven framework for identifying patient subgroups on which an AI/machine learning model may underperform.

Subbaswamy A, Sahiner B, Petrick N, Pai V, Adams R, Diamond M NPJ Digit Med. 2024; 7(1):334.

PMID: 39572755 PMC: 11582698. DOI: 10.1038/s41746-024-01275-6.

Thick Data Analytics (TDA): An Iterative and Inductive Framework for Algorithmic Improvement.

Nguyen M, Eulalio T, Marafino B, Rose C, Chen J, Baiocchi M Am Stat. 2024; 78(4):456-464.

PMID: 39524529 PMC: 11545316. DOI: 10.1080/00031305.2024.2327535.

References

Dunnmon J, Yi D, Langlotz C, Re C, Rubin D, Lungren M . Assessment of Convolutional Neural Networks for Automated Classification of Chest Radiographs. Radiology. 2018; 290(2):537-544. PMC: 6358056. DOI: 10.1148/radiol.2018181422. View

Mazurowski M, Habas P, Zurada J, Lo J, Baker J, Tourassi G . Training neural network classifiers for medical decision making: the effects of imbalanced datasets on classification performance. Neural Netw. 2008; 21(2-3):427-36. PMC: 2346433. DOI: 10.1016/j.neunet.2007.12.031. View

Badgeley M, Zech J, Oakden-Rayner L, Glicksberg B, Liu M, Gale W . Deep learning predicts hip fracture using confounding patient and healthcare variables. NPJ Digit Med. 2019; 2:31. PMC: 6550136. DOI: 10.1038/s41746-019-0105-1. View

Bien N, Rajpurkar P, Ball R, Irvin J, Park A, Jones E . Deep-learning-assisted diagnosis for knee magnetic resonance imaging: Development and retrospective validation of MRNet. PLoS Med. 2018; 15(11):e1002699. PMC: 6258509. DOI: 10.1371/journal.pmed.1002699. View

Chilamkurthy S, Ghosh R, Tanamala S, Biviji M, Campeau N, Venugopal V . Deep learning algorithms for detection of critical findings in head CT scans: a retrospective study. Lancet. 2018; 392(10162):2388-2396. DOI: 10.1016/S0140-6736(18)31645-3. View

Wang P, Berzin T, Glissen Brown J, Bharadwaj S, Becq A, Xiao X . Real-time automatic detection system increases colonoscopic polyp and adenoma detection rates: a prospective randomised controlled study. Gut. 2019; 68(10):1813-1819. PMC: 6839720. DOI: 10.1136/gutjnl-2018-317500. View

Fries J, Varma P, Chen V, Xiao K, Tejeda H, Saha P . Weakly supervised classification of aortic valve malformations using unlabeled cardiac MRI sequences. Nat Commun. 2019; 10(1):3111. PMC: 6629670. DOI: 10.1038/s41467-019-11012-3. View

Cardon L, Palmer L . Population stratification and spurious allelic association. Lancet. 2003; 361(9357):598-604. DOI: 10.1016/S0140-6736(03)12520-2. View

Chen V, Wu S, Weng Z, Ratner A, Re C . : A Programming Model for Residual Learning in Critical Data Slices. Adv Neural Inf Process Syst. 2019; 32:9392-9402. PMC: 6927210. View

10.

Rajpurkar P, Irvin J, Ball R, Zhu K, Yang B, Mehta H . Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practicing radiologists. PLoS Med. 2018; 15(11):e1002686. PMC: 6245676. DOI: 10.1371/journal.pmed.1002686. View

11.

Buda M, Maki A, Mazurowski M . A systematic study of the class imbalance problem in convolutional neural networks. Neural Netw. 2018; 106:249-259. DOI: 10.1016/j.neunet.2018.07.011. View

12.

Oakden-Rayner L . Exploring Large-scale Public Medical Image Datasets. Acad Radiol. 2019; 27(1):106-112. DOI: 10.1016/j.acra.2019.10.006. View

13.

Mulherin S, Miller W . Spectrum bias or spectrum effect? Subgroup variation in diagnostic test evaluation. Ann Intern Med. 2002; 137(7):598-602. DOI: 10.7326/0003-4819-137-7-200210010-00011. View

14.

Gulshan V, Peng L, Coram M, Stumpe M, Wu D, Narayanaswamy A . Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs. JAMA. 2016; 316(22):2402-2410. DOI: 10.1001/jama.2016.17216. View

15.

Esteva A, Kuprel B, Novoa R, Ko J, Swetter S, Blau H . Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017; 542(7639):115-118. PMC: 8382232. DOI: 10.1038/nature21056. View

16.

Dunnmon J, Ratner A, Saab K, Khandwala N, Markert M, Sagreiya H . Cross-Modal Data Programming Enables Rapid Medical Machine Learning. Patterns (N Y). 2020; 1(2). PMC: 7413132. DOI: 10.1016/j.patter.2020.100019. View

17.

Agniel D, Kohane I, Weber G . Biases in electronic health record data due to processes within the healthcare system: retrospective observational study. BMJ. 2018; 361:k1479. PMC: 5925441. DOI: 10.1136/bmj.k1479. View

18.

Campanella G, Hanna M, Geneslaw L, Miraflor A, Silva V, Busam K . Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat Med. 2019; 25(8):1301-1309. PMC: 7418463. DOI: 10.1038/s41591-019-0508-1. View

19.

Mahajan V, Venugopal V, Murugavel M, Mahajan H . The Algorithmic Audit: Working with Vendors to Validate Radiology-AI Algorithms-How We Do It. Acad Radiol. 2019; 27(1):132-135. DOI: 10.1016/j.acra.2019.09.009. View

20.

Ratner A, Ehrenberg H, Hussain Z, Dunnmon J, Re C . Learning to Compose Domain-Specific Transformations for Data Augmentation. Adv Neural Inf Process Syst. 2018; 30:3239-3249. PMC: 5786274. View