Towards Automated Clinical Coding

Overview

Journal Int J Med Inform

Specialty Medical Informatics

Date 2018 Nov 10

PMID 30409346

Citations 7

Authors

Finneas Catling

Georgios P Spithourakis

Sebastian Riedel

Affiliations

Soon will be listed here.

Abstract

Background: Patients' encounters with healthcare services must undergo clinical coding. These codes are typically derived from free-text notes. Manual clinical coding is expensive, time-consuming and prone to error. Automated clinical coding systems have great potential to save resources, and realtime availability of codes would improve oversight of patient care and accelerate research. Automated coding is made challenging by the idiosyncrasies of clinical text, the large number of disease codes and their unbalanced distribution.

Methods: We explore methods for representing clinical text and the labels in hierarchical clinical coding ontologies. Text is represented as term frequency-inverse document frequency counts and then as word embeddings, which we use as input to recurrent neural networks. Labels are represented atomically, and then by learning representations of each node in a coding ontology and composing a representation for each label from its respective node path. We consider different strategies for initialisation of the node representations. We evaluate our methods using the publicly-available Medical Information Mart for Intensive Care III dataset: we extract the history of presenting illness section from each discharge summary in the dataset, then predicting the International Classification of Diseases, ninth revision, Clinical Modification codes associated with these.

Results: Composing the label representations from the clinical-coding-ontology nodes increased weighted F1 for prediction of the 17,561 disease labels to 0.264-0.281 from 0.232-0.249 for atomic representations. Recurrent neural network text representation improved weighted F1 for prediction of the 19 disease-category labels to 0.682-0.701 from 0.662-0.682 using term frequency-inverse document frequency. However, term frequency-inverse document frequency outperformed recurrent neural networks for prediction of the 17,561 disease labels.

Conclusions: This study demonstrates that hierarchically-structured medical knowledge can be incorporated into statistical models, and produces improved performance during automated clinical coding. This performance improvement results primarily from improved representation of rarer diseases. We also show that recurrent neural networks improve representation of medical text in some settings. Learning good representations of the very rare diseases in clinical coding ontologies from data alone remains challenging, and alternative means of representing these diseases will form a major focus of future work on automated clinical coding.

Citing Articles

Classification of user queries according to a hierarchical medical procedure encoding system using an ensemble classifier.

Deng Y, Denecke K Front Artif Intell. 2022; 5:1000283.

PMID: 36406473 PMC: 9672500. DOI: 10.3389/frai.2022.1000283.

Conversion of Automated 12-Lead Electrocardiogram Interpretations to OMOP CDM Vocabulary.

Choi S, Joo H, Kim Y, Kim J, Seok J Appl Clin Inform. 2022; 13(4):880-890.

PMID: 36130711 PMC: 9492322. DOI: 10.1055/s-0042-1756427.

Consultation analysis: use of free text versus coded text.

Millares Martin P Health Technol (Berl). 2021; 11(2):349-357.

PMID: 33520588 PMC: 7829039. DOI: 10.1007/s12553-020-00517-3.

Natural language processing algorithms for mapping clinical text fragments onto ontology concepts: a systematic review and recommendations for future studies.

Kersloot M, van Putten F, Abu-Hanna A, Cornet R, Arts D J Biomed Semantics. 2020; 11(1):14.

PMID: 33198814 PMC: 7670625. DOI: 10.1186/s13326-020-00231-z.

Automated ICD coding via unsupervised knowledge integration (UNITE).

Sonabend W A, Cai W, Ahuja Y, Ananthakrishnan A, Xia Z, Yu S Int J Med Inform. 2020; 139:104135.

PMID: 32361145 PMC: 9410729. DOI: 10.1016/j.ijmedinf.2020.104135.