Artificial intelligence models for predicting future cardiovascular disease risk show promising performance, but major validation and reporting gaps currently limit their clinical use. A scoping review accepted for publication in BMC Medical Informatics and Decision Making assessed 30 studies of prognostic AI models in adults without established cardiovascular disease. All included studies were published after 2017 and covered a geographically diverse set of datasets, although the U.S. and UK accounted for most of the evidence. The models generally used established machine learning methods and routine clinical data, while external validation, calibration and measures of clinical utility remained uncommon or absent.
Models Rely Mainly on Routine Clinical Data
The 30 studies used a range of prospective, retrospective, observational and cross-sectional designs, with prospective cohorts the most common. Frequently used sources included the Framingham Heart Study, UK Biobank and Korea National Health and Nutrition Examination Survey. Dataset sizes varied widely, and several studies drew on the same publicly available resources, reducing the independence of some cross-study comparisons. Long-term prediction over 10 years was the most common time horizon, although shorter and multiple prediction periods were also used.
Must Read: Secure Data Pipelines in Regulated Settings
Most approaches relied on structured clinical information. Twenty-four studies used unimodal data, mainly routine or real-world clinical variables, while six combined different data types. The multimodal studies incorporated additional sources including proteomics and imaging alongside structured clinical information. Their performance was consistently strong, but the small and heterogeneous evidence base limits firm conclusions about the value of multimodal approaches.
Random Forest models and artificial neural networks were among the most frequently used methods, with support vector machines, machine-learning logistic regression and boosting algorithms also common. Established cardiovascular risk factors remained prominent. Age appeared in every study, while sex, diabetes status, smoking and systolic blood pressure were used in most. Only half of the studies included all seven factors used in the Framingham Risk Score. Feature selection methods also varied, and 15 studies did not report how features were selected.
Comparisons With Traditional Scores Remain Uncertain
Twelve studies directly compared AI or machine-learning models with traditional clinical risk scores, including the Framingham Risk Score, ASCVD and QRISK. Across these comparisons, AI models produced discrimination that was comparable with or modestly better than traditional approaches. The available evidence, however, does not establish that AI models are broadly superior for cardiovascular risk prediction.
Differences between studies make the comparisons difficult to interpret. Outcome definitions and prediction horizons varied, with some models estimating long-term incident cardiovascular disease and others addressing shorter-term or composite outcomes. The models also did not always use the same predictor sets as the traditional scores against which they were tested. When an AI model had access to more or different variables, any performance advantage could reflect the additional predictors rather than the modelling method itself.
Validation strategies also differed between competing approaches. Some AI models were assessed only using internal or cross-validation, while traditional scores were applied to external populations. This imbalance can favour the internally evaluated model and weakens claims of comparative advantage. Higher discrimination values therefore cannot, on their own, demonstrate superior clinical prediction. Future comparisons require the same datasets, predictor sets, outcome definitions and validation strategies. Without that consistency, modest gains in discrimination cannot demonstrate that AI-based risk prediction is more clinically reliable than established scoring systems.
Validation and Reporting Limit Clinical Readiness
External validation was uncommon. Only seven of the 30 studies tested models on external data, while 22 relied on internal validation and one did not perform validation. Cross-validation was the most common internal strategy. None of the externally validated models was tested across geographically or ethnically distinct populations. This is particularly relevant because 19 studies used U.S. or UK populations, leaving performance across other populations insufficiently tested.
Reporting of model performance also focused heavily on discrimination. Area under the curve was the most frequently reported metric, followed by the C-index and accuracy. Calibration reporting remained a major gap. Sensitivity was reported in only four studies, specificity in seven and no study reported a full decision curve analysis. Reliance on discrimination and accuracy without calibration or clinical utility measures limits assessment of whether predicted risks are dependable in practice.
Reporting completeness was assessed against TRIPOD+AI, while PROBAST+AI was used to assess methodological quality, risk of bias and applicability. Twenty-one studies were rated as having high overall quality concern for model development, with participants and data sources a major contributor. Analysis was the dominant source of high evaluation risk, particularly because calibration was missing. Eight studies did not report how missing data were handled and none provided a rationale for sample size. These gaps constrain reproducibility, transportability and assessment of real-world clinical applicability.
AI-based cardiovascular risk prediction has developed rapidly, with models using large datasets, established machine-learning techniques and, in a smaller number of cases, multimodal information. Some comparisons show discrimination comparable with or modestly better than traditional risk scores, but the evidence does not establish broad superiority. Limited external validation, incomplete reporting and weak assessment of calibration and clinical utility remain major barriers to deployment. More consistent evaluation across diverse populations, clearer reporting and routine use of calibration alongside discrimination and clinically relevant performance measures are needed before these models can be judged ready for clinical decision-making.
Source: BMC Medical Informatics & Decision Making
Image Credit: iStock
References:
Pai S, Subramanian R, Krishnan L et al. (2026) A scoping review on artificial intelligence-based tools for cardiovascular disease risk prediction. BMC Med Inform Decis Mak. https://doi.org/10.1186/s12911-026-03699-4