» Articles » PMID: 38042599

Natural Language Processing with Machine Learning Methods to Analyze Unstructured Patient-reported Outcomes Derived from Electronic Health Records: A Systematic Review

Overview
Date 2023 Dec 2
PMID 38042599
Authors
Affiliations
Soon will be listed here.
Abstract

Objective: Natural language processing (NLP) combined with machine learning (ML) techniques are increasingly used to process unstructured/free-text patient-reported outcome (PRO) data available in electronic health records (EHRs). This systematic review summarizes the literature reporting NLP/ML systems/toolkits for analyzing PROs in clinical narratives of EHRs and discusses the future directions for the application of this modality in clinical care.

Methods: We searched PubMed, Scopus, and Web of Science for studies written in English between 1/1/2000 and 12/31/2020. Seventy-nine studies meeting the eligibility criteria were included. We abstracted and summarized information related to the study purpose, patient population, type/source/amount of unstructured PRO data, linguistic features, and NLP systems/toolkits for processing unstructured PROs in EHRs.

Results: Most of the studies used NLP/ML techniques to extract PROs from clinical narratives (n = 74) and mapped the extracted PROs into specific PRO domains for phenotyping or clustering purposes (n = 26). Some studies used NLP/ML to process PROs for predicting disease progression or onset of adverse events (n = 22) or developing/validating NLP/ML pipelines for analyzing unstructured PROs (n = 19). Studies used different linguistic features, including lexical, syntactic, semantic, and contextual features, to process unstructured PROs. Among the 25 NLP systems/toolkits we identified, 15 used rule-based NLP, 6 used hybrid NLP, and 4 used non-neural ML algorithms embedded in NLP.

Conclusions: This study supports the potential utility of different NLP/ML techniques in processing unstructured PROs available in EHRs for clinical care. Though using annotation rules for NLP/ML to analyze unstructured PROs is dominant, deploying novel neural ML-based methods is warranted.

Citing Articles

Leveraging Natural Language Processing and Machine Learning Methods for Adverse Drug Event Detection in Electronic Health/Medical Records: A Scoping Review.

Golder S, Xu D, OConnor K, Wang Y, Batra M, Hernandez G Drug Saf. 2025; 48(4):321-337.

PMID: 39786481 PMC: 11903561. DOI: 10.1007/s40264-024-01505-6.


The Frontiers of Smart Healthcare Systems.

Lin N, Paul R, Guerra S, Liu Y, Doulgeris J, Shi M Healthcare (Basel). 2024; 12(23).

PMID: 39684952 PMC: 11641075. DOI: 10.3390/healthcare12232330.


Identifying stigmatizing and positive/preferred language in obstetric clinical notes using natural language processing.

Scroggins J, Hulchafo I, Harkins S, Scharp D, Moen H, Davoudi A J Am Med Inform Assoc. 2024; 32(2):308-317.

PMID: 39569431 PMC: 11756426. DOI: 10.1093/jamia/ocae290.


Nursing Records Regarding Decision-Making in Cancer Supportive Care: A Retrospective Study in Japan.

Kawasaki Y, Nii M, Nishioka E Healthc Inform Res. 2024; 30(4):364-374.

PMID: 39551923 PMC: 11570663. DOI: 10.4258/hir.2024.30.4.364.


The recent history and near future of digital health in the field of behavioral medicine: an update on progress from 2019 to 2024.

Arigo D, Jake-Schoffman D, Pagoto S J Behav Med. 2024; 48(1):120-136.

PMID: 39467924 PMC: 11893649. DOI: 10.1007/s10865-024-00526-x.


References
1.
Topaz M, Adams V, Wilson P, Woo K, Ryvicker M . Free-Text Documentation of Dementia Symptoms in Home Healthcare: A Natural Language Processing Study. Gerontol Geriatr Med. 2020; 6:2333721420959861. PMC: 7520927. DOI: 10.1177/2333721420959861. View

2.
Jensen K, Soguero-Ruiz C, Mikalsen K, Lindsetmo R, Kouskoumvekaki I, Girolami M . Analysis of free text in electronic health records for identification of cancer patient trajectories. Sci Rep. 2017; 7:46226. PMC: 5384191. DOI: 10.1038/srep46226. View

3.
Yang Z, Dehmer M, Yli-Harja O, Emmert-Streib F . Combining deep learning with token selection for patient phenotyping from electronic health records. Sci Rep. 2020; 10(1):1432. PMC: 6989657. DOI: 10.1038/s41598-020-58178-1. View

4.
Sorup F, Eriksson R, Westergaard D, Hallas J, Brunak S, Andersen S . Sex differences in text-mined possible adverse drug events associated with drugs for psychosis. J Psychopharmacol. 2020; 34(5):532-539. DOI: 10.1177/0269881120903466. View

5.
Fodeh S, Finch D, Bouayad L, Luther S, Ling H, Kerns R . Classifying clinical notes with pain assessment using machine learning. Med Biol Eng Comput. 2017; 56(7):1285-1292. PMC: 6014866. DOI: 10.1007/s11517-017-1772-1. View