Automated identification of incidental hepatic steatosis on Emergency Department imaging using large language models

Tyrus Vong; Nicholas Rizer; Vedant Jain; Valerie L Thompson; Mark Dredze; Eili Y Klein; Jeremiah S Hinson; Tanjala Purnell; Stephen Kwak; Tinsay Woreta; Alexandra T Strauss

doi:10.1097/HC9.0000000000000638

Automated identification of incidental hepatic steatosis on Emergency Department imaging using large language models

Hepatol Commun. 2025 Feb 19;9(3):e0638. doi: 10.1097/HC9.0000000000000638. eCollection 2025 Mar 1.

Authors

Affiliations

¹ Division of Gastroenterology & Hepatology, Johns Hopkins University School of Medicine, Baltimore, Maryland, USA.
² Department of Emergency Medicine, Johns Hopkins University School of Medicine, Baltimore, Maryland, USA.
³ Department of Gastroenterology & Hepatology Carle Illinois College of Medicine, University of Illinois at Urbana-Champaign, Urbana, Illinois, USA.
⁴ Department of Computer Science, Johns Hopkins University, Baltimore, Maryland, USA.
⁵ Department of Epidemiology, Johns Hopkins University Bloomberg School of Public Health, Baltimore, Maryland, USA.
⁶ Department of Radiology, Johns Hopkins University School of Medicine, Baltimore, Maryland, USA.

Abstract

Background: Hepatic steatosis is a precursor to more severe liver disease, increasing morbidity and mortality risks. In the Emergency Department, routine abdominal imaging often reveals incidental hepatic steatosis that goes undiagnosed due to the acute nature of encounters. Imaging reports in the electronic health record contain valuable information not easily accessible as discrete data elements. We hypothesized that large language models could reliably detect hepatic steatosis from reports without extensive natural language processing training.

Methods: We identified 200 adults who had CT abdominal imaging in the Emergency Department between August 1, 2016, and December 31, 2023. Using text from imaging reports and structured prompts, 3 Azure OpenAI models (ChatGPT 3.5, 4, 4o) identified patients with hepatic steatosis. We evaluated model performance regarding accuracy, inter-rater reliability, sensitivity, and specificity compared to physician reviews.

Results: The accuracy for the models was 96.2% for v3.5, 98.3% for v4, and 98.8% for v4o. Inter-rater reliability ranged from 0.99 to 1.00 across 10 iterations. Mean model confidence scores were 2.9 (SD 0.8) for v3.5, 3.9 (SD 0.3) for v4, and 4.0 (SD 0.07) for v4o. Incorrect evaluations were 76 (3.8%) for v3.5, 34 (1.7%) for v4, and 25 (1.3%) for v4o. All models showed sensitivity and specificity above 0.9.

Conclusions: Large language models can assist in identifying incidental conditions from imaging reports that otherwise may be missed opportunities for early disease intervention. Large language models are a democratization of natural language processing by allowing for a user-friendly, expansive analyses of electronic medical records without requiring the development of complex natural language processing models.

MeSH terms

Adult
Aged
Electronic Health Records
Emergency Service, Hospital*
Fatty Liver* / diagnostic imaging
Female
Humans
Incidental Findings*
Large Language Models
Male
Middle Aged
Natural Language Processing*
Reproducibility of Results
Sensitivity and Specificity
Tomography, X-Ray Computed