Clinical predictive artificial intelligence models can support clinical decisions and risk communication, but they can also reproduce bias when fairness is not carefully assessed. Bias can enter model development through non-representative data, different disease patterns, existing health disparities, biased medical devices and other mechanisms. A scoping review published in The Lancet Digital Health assessed how fairness is measured in clinical predictive AI models. Searches covered English-language literature from 2014 to 2024 across PubMed, ACM Digital Library, IEEE Xplore, arXiv and medRxiv. After screening 820 records, 42 publications were included and 63 fairness metrics were identified. The results point to a varied and unsettled field, with many metrics lacking clear clinical validation and many depending on decision thresholds.

 

Fairness Requires Clearer Definitions
Clinical prediction models estimate the probability of current or future health outcomes using baseline predictors. They can be diagnostic or prognostic, depending on whether they assess a current condition or a future outcome over a defined period. Their performance is often judged at population level, but this can hide important differences between subgroups. Such hidden variation can create or maintain unfairness if model behaviour differs across people grouped by sensitive attributes.

 

Fairness in clinical predictive AI depends on how sensitive attributes are defined and used. These can include age, race or ethnicity, sex or gender and socioeconomic status. The meaning and relevance of these attributes can vary by setting. In healthcare, some attributes may also reflect biological differences that can legitimately affect outcomes, which makes fairness assessment more complex.

 

A fairness metric quantifies whether a model output disadvantages individuals or groups defined by sensitive attributes. The 63 identified metrics were grouped by whether they depend on model performance, whether they use estimated probabilities or predicted classes and which underlying performance measure they compare. Most metrics assess group fairness. Individual fairness metrics were rare, with only two identified.

 

Must Read: Fairness in Unsupervised Healthcare AI

 

Metric Selection Remains Fragmented
The identified metrics came from different fields. Most originated from AI, while fewer came from biomedical or applied ethics research. This contributes to a fragmented landscape, with overlapping definitions, inconsistent terminology and varying levels of detail. Some metrics were clearly defined and linked to ethical or legal principles, while others had limited justification or insufficient information for reliable use.

Performance-independent metrics assess model behaviour without using outcome labels. These metrics can be useful when historical labels may reflect health inequities, such as unequal access to care. They often focus on whether positivity rates are similar across groups. However, such parity can be difficult to interpret when outcome prevalence differs across groups. A model that is accurate overall may still fail to meet parity-based fairness criteria if underlying risks differ.

 

Performance-dependent metrics were more common. These compare model performance across individuals or groups. They can assess whether discrimination, calibration, overall performance or clinical usefulness remains consistent across subgroups. However, they may preserve patterns already present in the data. If historical data contain unequal care pathways or biased labels, performance-dependent metrics may reinforce rather than challenge those patterns. This makes the choice of metric closely tied to whether existing data patterns are considered acceptable for clinical use.

 

Clinical Utility Receives Limited Attention
Many fairness metrics rely on thresholds that turn estimated probabilities into predicted classes. These threshold-dependent metrics can be easier to calculate, but they depend on the chosen cutoff. If that cutoff lacks a clinical rationale, the fairness result may have limited meaning. Results from one threshold may also not apply to another threshold or clinical setting.

 

Probability-based metrics avoid this issue by using estimated probabilities directly. These metrics can assess areas such as discrimination, calibration or overall performance without first converting risk estimates into classes. Metrics comparing discrimination, such as AUROC parity, can help identify subgroup differences, but they need to be paired with calibration-related assessment. Calibration metrics can also be informative, especially when supported by subgroup calibration plots.

 

Clinical utility was much less represented. Only one metric was explicitly identified as focusing on clinical utility. This matters because clinical prediction models are intended to support decision making. Metrics based only on statistical parity may not show whether model use improves or worsens clinical decisions for different groups. The limited attention to clinical utility leaves an important gap in fairness evaluation, especially for models intended to guide interventions, referrals or risk-based decisions.


Fairness assessment in clinical predictive AI remains uneven, with many metrics lacking clear definitions, clinical validation or practical guidance. Performance-dependent and threshold-dependent metrics dominate, while uncertainty, intersectionality, individual fairness and clinical utility receive less attention. Fairness evaluation needs more than a single metric. Subgroup performance, calibration, decision-focused assessment and transparent reporting all have a role. More clinically meaningful approaches are needed to show how models behave across groups and whether their use supports fairer decision making.

 

Source: The Lancet Digital Health

Image Credit: iStock  


References:

Matos J, Van Calster B, Celi L et al. (2026) Critical appraisal of fairness metrics for artificial intelligence-based clinical prediction models: a scoping review. The Lancet Digital Health: Online first.




Latest Articles

clinical AI fairness, predictive AI bias, healthcare AI metrics, AI model equity, clinical prediction models, algorithmic fairness healthcare Fairness metrics in clinical AI remain fragmented and poorly validated, risking bias in predictive models across patient subgroups and care decisions.