Development and validation of prediction models for stroke and myocardial infarction in type 2 diabetes based on health insurance claims: does machine learning outperform traditional regression approaches?

Anna-Janina Stephan; Michael Hanselmann; Medina Bajramovic; Simon Schosser; Michael Laxy

doi:10.1186/s12933-025-02640-9

Development and validation of prediction models for stroke and myocardial infarction in type 2 diabetes based on health insurance claims: does machine learning outperform traditional regression approaches?

Cardiovasc Diabetol. 2025 Feb 18;24(1):80. doi: 10.1186/s12933-025-02640-9.

Authors

Anna-Janina Stephan^{1

2}, Michael Hanselmann^{3

4}, Medina Bajramovic^{3

5}, Simon Schosser^{3

5}, Michael Laxy^{3

4}

Affiliations

¹ Professorship for Public Health and Prevention, TUM School of Medicine and Health, Technical University of Munich, Munich, Germany. anna-janina.stephan@tum.de.
² German Center for Diabetes Research (DZD), Munich, Germany. anna-janina.stephan@tum.de.
³ Professorship for Public Health and Prevention, TUM School of Medicine and Health, Technical University of Munich, Munich, Germany.
⁴ German Center for Diabetes Research (DZD), Munich, Germany.
⁵ Department of Statistics, Ludwig-Maximilians-Universität München, Munich, Germany.

Abstract

Background: Digitalization and big health system data open new avenues for targeted prevention and treatment strategies. We aimed to develop and validate prediction models for stroke and myocardial infarction (MI) in patients with type 2 diabetes based on routinely collected high-dimensional health insurance claims and compared predictive performance of traditional regression with state-of-the-art machine learning including deep learning methods.

Methods: We used German health insurance claims from 2014 to 2019 with 287 potentially relevant literature-derived variables to predict 3-year risk of MI and stroke. Following a train-test split approach, we compared the performance of logistic methods with and without forward selection, LASSO-regularization, random forests (RF), gradient boosting (GB), multi-layer-perceptrons (MLP) and feature-tokenizer transformers (FTT). We assessed discrimination (Areas Under the Precision-Recall and Receiver-Operator Curves, AUPRC and AUROC) and calibration.

Results: Among n = 371,006 patients with type 2 diabetes (mean age: 67.2 years), 3.5% (n = 13,030) had MIs and 3.4% (n = 12,701) strokes. AUPRCs were 0.035 (MI) and 0.034 (stroke) for a null model, between 0.082 (MLP) and 0.092 (GB) for MI, and between 0.061 (MLP) and 0.073 (GB) for stoke. AUROCs were 0.5 for null models, between 0.70 (RF, MLP, FTT) and 0.71 (all other models) for MI, and between 0.66 (MLP) and 0.69 (GB) for stroke. All models were well calibrated.

Conclusions: Discrimination performance of claims-based models reached a ceiling at around 0.09 AUPRC and 0.7 AUROC. While for AUROC this performance was comparable to existing epidemiological models incorporating clinical information, comparison of other, potentially more relevant metrics, such as AUPRC, sensitivity and Positive Predictive Value was hampered by lack of reporting in the literature. The fact that machine learning including deep learning methods did not outperform more traditional approaches may suggest that feature richness and complexity were exploited before the choice of algorithm could become critical to maximize performance. Future research might focus on the impact of different feature derivation approaches on performance ceilings. In the absence of other more powerful screening alternatives, applying transparent regression-based models in routine claims, though certainly imperfect, remains a promising scalable low-cost approach for population-based cardiovascular risk prediction and stratification.

Keywords: Claims database analysis; Deep learning; Health insurance; Logistic regression; Machine learning; Myocardial infarction; Prediction model; Predictive algorithms; Risk scores; Stroke; Type 2 diabetes.

Publication types

Validation Study
Comparative Study

MeSH terms

Administrative Claims, Healthcare*
Aged
Databases, Factual
Decision Support Techniques*
Deep Learning
Diabetes Mellitus, Type 2* / diagnosis
Diabetes Mellitus, Type 2* / epidemiology
Female
Germany / epidemiology
Humans
Insurance Claim Review*
Machine Learning
Male
Middle Aged
Myocardial Infarction* / diagnosis
Myocardial Infarction* / epidemiology
Predictive Value of Tests
Prognosis
Reproducibility of Results
Risk Assessment
Risk Factors
Stroke* / diagnosis
Stroke* / epidemiology
Time Factors