Effective clinical use of artificial intelligence depends on clear evidence of model performance, especially as medical AI now handles images, audio, structured health information and unstructured text. A 2026 special report published in Radiology: Artificial Intelligence sets out a machine-interpretable taxonomy of performance metrics within the Radiology Ontology of AI Datasets, Models and Projects, known as ROADMAP. The framework formalises the language and descriptions of metrics used to measure AI model performance across clinical and multimodal applications. It gives developers, regulators, purchasers and users a shared vocabulary for comparing AI systems, checking reporting completeness and detecting issues such as bias, calibration problems and uncertainty across different clinical settings and over time.
Shared Language for Performance Metrics
The ROADMAP taxonomy contains 207 metrics relevant to clinical and multimodal AI evaluation. It covers widely used measures such as sensitivity, Dice similarity coefficient and area under the receiver operating characteristic curve, alongside metrics associated with large language models, fairness, readability and model drift. The breadth reflects the expanding range of medical AI data types and use cases, from image segmentation and binary classification to diagnosis, free-text generation and time-to-event prediction.
The framework treats a metric as any form in which performance information appears. The model performance metric class has three main subclasses: graphical metrics, matrix metrics and scalar metrics. Graphical metrics show performance by plotting one variable against another. Receiver operating characteristic analysis, for example, plots true-positive rate against false-positive rate across varying decision thresholds, while overall performance can be summarised through the area under the receiver operating characteristic curve. Other graphical metrics include the Kaplan-Meier estimator, precision-recall curve and calibration curve.
Matrix metrics include the confusion matrix, with entities for binary and multiclass classification tasks. In a binary confusion matrix, cells contain counts of true positives, false positives, true negatives and false negatives. Scalar metrics provide a single numerical value and form the largest group, with 193 entities including Brier score and positive predictive value.
Linking Metrics to Clinical AI Tasks
The performance criterion class defines the characteristics by which AI performance can be judged. Its 18 subclasses include calibration performance, classification performance, fairness, image processing performance, image segmentation performance, information content, model drift, prediction performance, readability performance, regression performance, signal analysis performance, text analysis performance, time-to-event analysis performance, user agreement performance, user response performance and utility. A dedicated relationship links each metric to one or more criteria, clarifying what a given measure evaluates.
This linking structure matters because one metric can serve more than one evaluation purpose. Dice similarity coefficient, for example, relates to both classification and image segmentation performance. Graphical and matrix metrics can also connect to scalar metrics, allowing a performance curve or matrix to remain associated with its derived numerical summary. That structure helps underlying data and summary statistics stay interpretable together, rather than appearing as isolated values.
Each metric includes an English-language label and a human-readable definition. Some entries also contain synonyms, abbreviations, terms in other languages, numerical bounds and mathematical formulae. The true positive rate entry includes English and Spanish labels, alternate labels such as recall and sensitivity, lower and upper bounds, a formula and an association with classification performance. Some common names appear as alternate labels rather than preferred labels, with sensitivity serving as an alternate name for true positive rate and specificity as an alternate name for true negative rate.
Must Read: AI Use in Medication Reconciliation Remains Limited
Structured Representation for Responsible Evaluation
The ontology organises metrics through hierarchical relationships and categorical groupings, creating a navigable knowledge base for AI evaluation. Its representation remains readable by humans and machines through formats such as the Resource Description Framework and query languages such as SPARQL. Automated processing and reasoning can support complex queries, define the context in which a metric applies and clarify the relevance of particular metrics to specific machine learning tasks or datasets.
ROADMAP connects with SNOMED CT and RadLex to support harmonisation across studies, data repositories and regulatory submissions. SNOMED CT contributes standard terminology for anatomic structures, diseases and medical interventions. RadLex contributes vocabulary for medical imaging findings, relevant anatomy, clinical conditions and related concepts. The metrics component supports consistent reporting, model comparison on common grounds and identification of bias or reporting gaps.
The taxonomy complements Metrics Reloaded and MIDRC-MetricTree, which guide users through decision trees for selecting metrics in particular problem settings. ROADMAP standardises the representation of those metrics and could support partial automation of such tools. Future work centres on alignment with the Information Artifact Ontology and the Basic Formal Ontology, extraction of ROADMAP metadata from manuscripts with a large language model–based tool and a community submission platform for proposed metrics, revisions and use-case examples.
ROADMAP’s metrics taxonomy gives medical AI evaluation a structured language for performance measurement, metric selection and interpretation. By organising graphical, matrix and scalar metrics within a machine-interpretable ontology, it supports transparent reporting, reproducible comparison and more consistent communication across clinical and multimodal AI applications. Its links to performance criteria, biomedical terminologies and future community contributions create a foundation for rigorous evaluation as AI systems continue to expand across data types, tasks and clinical settings. It also preserves context by linking metric categories, criteria and data modalities within the same framework.
Source: Radiology: Artificial Intelligence
Image Credit: iStock
References:
Gonzales RA, Takahashi MS, Retson T et al. (2026) Metrics for Artificial Intelligence in Medicine: A Reference Resource. Radiology: Artificial Intelligence; 8:3.