Automated extraction of functional biomarkers of verbal and ambulatory ability from multi-institutional clinical notes using large language models

Levi Kaster; Ethan Hillis; Inez Y Oh; Bhooma R Aravamuthan; Virginia C Lanzotti; Casey R Vickstrom; Brain Gene Registry Consortium; Christina A Gurnett; Philip R O Payne; Aditi Gupta

doi:10.1186/s11689-025-09612-w

Automated extraction of functional biomarkers of verbal and ambulatory ability from multi-institutional clinical notes using large language models

J Neurodev Disord. 2025 Apr 30;17(1):24. doi: 10.1186/s11689-025-09612-w.

Authors

Levi Kaster¹, Ethan Hillis¹, Inez Y Oh¹, Bhooma R Aravamuthan², Virginia C Lanzotti³, Casey R Vickstrom²; Brain Gene Registry Consortium; Christina A Gurnett², Philip R O Payne¹, Aditi Gupta⁴

Collaborators

Brain Gene Registry Consortium:
M Wasserstein, M Chopra, M Sahin, M Wangler, B Schultz, K Izumi, S Bergner, A Gropman, C Smith-Hicks, L Abbeduto, H Hazlett, D Doherty, K German, L DaWalt, J Neul, J Constantino, D Baldridge, S Srivastava, S Molholm, S Walkley, E Storch, R Samaco, J Cohen, S Shankar, J Piven, S Mahida, A Sveden, K Dies, E R Riggs, J M Savatt, B Minor

Affiliations

¹ Institute for Informatics, Data Science and Biostatistics, Washington University School of Medicine in St. Louis, St. Louis, MO, USA.
² Department of Neurology, Washington University School of Medicine in St. Louis, St. Louis, MO, USA.
³ Department of Psychiatry, Washington University School of Medicine in St. Louis, St. Louis, MO, USA.
⁴ Institute for Informatics, Data Science and Biostatistics, Washington University School of Medicine in St. Louis, St. Louis, MO, USA. agupta24@wustl.edu.

Abstract

Background: Functional biomarkers in neurodevelopmental disorders, such as verbal and ambulatory abilities, are essential for clinical care and research activities. Treatment planning, intervention monitoring, and identifying comorbid conditions in individuals with intellectual and developmental disabilities (IDDs) rely on standardized assessments of these abilities. However, traditional assessments impose a burden on patients and providers, often leading to longitudinal inconsistencies and inequities due to evolving guidelines and associated time-cost. Therefore, this study aimed to develop an automated approach to classify verbal and ambulatory abilities from EHR data of IDD and cerebral palsy (CP) patients. Application of large language models (LLMs) to clinical notes, which are rich in longitudinal data, may provide a low-burden pipeline for extracting functional biomarkers efficiently and accurately.

Methods: Data from the multi-institutional National Brain Gene Registry (BGR) and a CP clinic cohort were utilized, comprising 3,245 notes from 125 individuals and 5,462 clinical notes from 260 individuals, respectively. Employing three LLMs-GPT-3.5 Turbo, GPT-4 Turbo, and GPT-4 Omni-we provided the models with a clinical note and utilized a detailed conversational format to prompt the models to answer: "Does the individual use any words?" and "Can the individual walk without aid?" These responses were evaluated against ground-truth abilities, which were established using neurobehavioral assessments collected for each dataset.

Results: LLM pipelines demonstrated high accuracy (weighted-F1 scores > .90) in predicting ambulatory ability for both cohorts, likely due to the consistent use of Gross Motor Functional Classification System (GMFCS) as a consistent ground-truth standard. However, verbal ability predictions were more accurate in the BGR cohort, likely due to higher adherence between the prompt and ground-truth assessment questions. While LLMs can be computationally expensive, analysis of our protocol affirmed the cost effectiveness when applied to select notes from the EHR.

Conclusions: LLMs are effective at extracting functional biomarkers from EHR data and broadly generalizable across variable note-taking practices and institutions. Individual verbal and ambulatory ability were accurately extracted, supporting the method's ability to streamline workflows by offering automated, efficient data extraction for patient care and research. Future studies are needed to extend this methodology to additional populations and to demonstrate more granular functional data classification.

Keywords: Electronic health records; Functional biomarkers; Large language models; Neurodevelopmental disorders.

Publication types

Multicenter Study

MeSH terms

Adolescent
Adult
Biomarkers
Cerebral Palsy* / diagnosis
Cerebral Palsy* / physiopathology
Child
Child, Preschool
Developmental Disabilities* / diagnosis
Developmental Disabilities* / physiopathology
Electronic Health Records*
Female
Humans
Intellectual Disability* / diagnosis
Intellectual Disability* / physiopathology
Large Language Models*
Male
Natural Language Processing
Registries
Young Adult

Substances

Biomarkers

Abstract

Publication types

MeSH terms

Substances

Grants and funding