Robert Kalfus, MD1; Christine Chen, MD1; Su Yin Ong, MSc2; Amy R. Mulick, PhD3; Simone Heeg, PhD4; Kiliana
Suzart-Woischik, MD MPH3; Anand Shroff1
Background
Key clinical outcomes are often absent from structured EHR fields and rarely documented explicitly. Relevant information is instead embedded in narrative notes. The modified Rankin Scale (mRS), widely used in acute stroke trials, exemplifies this limitation. Large language models (LLMs) offer a promising approach for automated inference of outcome-related data from routine clinical documentation.
Objectives
To evaluate whether LLMs can accurately infer a critical clinical outcome measure, the mRS score, from routinely collected narrative clinical documentation, as a proof of concept for scalable outcome ascertainment.
Methods
We performed a retrospective study using routinely collected data from a single integrated delivery network in the Midwest region of the United States. Eligible patients were adults (≥18 years) with ischemic stroke receiving emergency or inpatient care from 2018–2024. Functional status at discharge and during follow-up was assessed using the mRS score. To identify documentation of functional status at discharge, we examined physical and occupational therapy notes from inpatient admissions. For functional status during follow-up, we used outpatient visit notes occurring 7–425 days post-stroke. Notes were screened for explicitly stated mRS scores using a keyword-based heuristic and then randomly sampled into training (n=150), validation (n=100), and test (n=100) sets separately for discharge and follow-up groups. All sampled notes were independently annotated by two clinicians, with only current and definitive mRS scores retained as the reference standard. Explicit scores and accompanying descriptions were redacted prior to LLM processing. Prompt engineering was performed using the training and validation sets, and model accuracy was assessed on the held-out test sets. Based on preliminary review and prior mRS research, we targeted evaluation of dichotomized mRS outcomes (0–2 vs. 3–5) rather than exact score derivation.
Results
The LLM achieved strong performance in inferring dichotomized mRS. For discharge mRS score assessments, final test performance was precision 0.84, recall 0.83, and F1-score 0.84. For follow-up mRS score assessments, performance was 0.93, 0.85, and 0.88, respectively. All metrics exceeded the predefined threshold of 0.80, demonstrating robust accuracy across both discharge and follow-up.
Conclusions
LLMs can accurately infer outcome measures from routinely collected clinical documentation, demonstrating feasibility with the mRS score as a proof of concept. This approach offers a scalable strategy for capturing outcomes rarely documented in structured data, with broad potential to enhance outcome ascertainment and real-world evidence generation across clinical domains.