Synthetic medical imaging is advancing rapidly, but methods used to judge whether generated images are reliable remain inconsistent. A systematic review accepted for publication in the European Journal of Radiology assessed evaluation approaches across 47 studies and found substantial variation in how fidelity, realism, diversity and clinical validity are measured. No single metric captured all four dimensions, while metric selection was often driven by data availability rather than clinical purpose. The findings support a task-specific, layered approach that combines complementary evaluation methods according to the intended use, with additional expert assessment when synthetic images affect clinical decisions. 

 

Common Metrics Capture Different Dimensions 

Evaluation practices were spread across four main categories: reference-based metrics, no-reference metrics, expert evaluation and task-based assessment. Expert and reference-based approaches were the most common, each appearing in just over half of the included studies, while no-reference and task-based approaches were used at similar rates. Most studies combined more than one category, but only a minority used more than two, leaving key dimensions of synthetic image quality insufficiently covered. Most included studies originated from Europe, North America and Asia, with only one from Africa. Fidelity concerns correspondence with a paired reference, while realism concerns perceptual plausibility and diversity reflects coverage of variation in real data. 

 

Must Read: Structured Reporting Improves Radiography Workflow 

 

Reference-based metrics compare a synthetic image with a paired ground-truth image and are useful for measuring reconstruction fidelity. Common measures included peak signal-to-noise ratio, structural similarity index and mean absolute error. Their computational simplicity and widespread use make them practical engineering benchmarks, but they cannot establish anatomical or clinical validity on their own. They are also sensitive to preprocessing choices such as normalisation, windowing and spatial alignment, which can limit comparisons across studies and datasets. 

 

No-reference approaches are used when paired images are unavailable and commonly assess the distributional similarity or quality of real and synthetic images. Fréchet Inception Distance was widely used, although its application to medical imaging is limited because its feature extractor was trained on natural rather than medical images. Distribution-based metrics may also fail to capture rare pathologies or subtle clinically relevant features. 

 

Clinical Utility Requires More Than Image Quality 

Expert evaluation remains important because it can directly assess realism and clinical validity. Reader studies can use blinded presentation, balanced case selection and measures of inter-reader agreement, but they are resource-intensive, difficult to scale and vulnerable to variability and cognitive bias. A synthetic image may appear realistic without preserving the information needed for a clinical task, while an image that appears less natural may still be useful for a constrained application. Expert conclusions are therefore specific to the task and context in which the images are assessed. 

 

Task-based evaluation focuses on downstream performance rather than image appearance alone. Examples include lesion counting, segmentation performance, relaxometry measurements and dosimetric accuracy. This approach is particularly relevant for data augmentation and modality translation, where the aim is to support performance on real clinical data. However, results depend on task design, dataset composition and the downstream model. Strong performance on one task does not establish overall image quality, and weak performance does not necessarily indicate poor synthesis. Optimising synthetic images for a narrow task can also come at the expense of other clinically relevant information. 

 

Technical image quality and clinical utility therefore remain related but distinct. Conventional metrics can assign high scores despite anatomical distortions or pathological changes, while lower-scoring images may preserve the information required for a particular purpose. Realism should not be treated as equivalent to clinical validity, and evaluation needs to reflect the intended diagnostic, therapeutic or computational use. 

 

Layered Evaluation Matches the Intended Use 

A layered framework is proposed to match evaluation methods with data availability and intended application. When paired ground-truth data exist, reference-based fidelity measures can provide an initial assessment. Distribution-based measures and diversity assessment are more appropriate for unpaired generation, while data augmentation should be tested through downstream performance on independent real-world datasets. Whenever synthetic images influence clinical decisions, expert reader evaluation should also be included. A single approach is not considered suitable across different imaging modalities, anatomical regions and clinical purposes. 

 

The framework also calls for transparent reporting of preprocessing, registration and metric selection. Proposed minimum reporting includes the intended clinical application, synthesis task, pairing status, preprocessing specifications, registration procedures, the metrics used and the reasons for choosing them. Domain-specific measures, including voxel-wise intensity or radiodensity agreement and morphometric accuracy, may help identify anatomically or clinically relevant discrepancies that conventional image-quality metrics miss. Evaluation strategies should be explicitly justified according to the available data and use case. 

 

Evaluation also extends beyond image quality. Synthetic data may retain identifying information from training datasets, creating a potential risk of re-identification. Existing privacy frameworks such as the General Data Protection Regulation were not designed specifically for synthetic data, and classification and validation remain unclear. Synthetic images may also reproduce or amplify demographic, sex, scanner-related or site-specific biases, potentially affecting performance in underrepresented patient groups. 

 

Synthetic medical imaging cannot be validated adequately through a single similarity or realism measure. The available evidence shows a heterogeneous evaluation landscape in which common metrics capture only selected aspects of image quality and may not reflect clinical validity or downstream utility. A multidimensional, task-aware approach can combine fidelity, distributional, expert and task-based assessment according to the intended application and available data. Transparent reporting of evaluation choices, alongside attention to privacy, bias and clinically relevant discrepancies, is needed to support interpretable comparisons and safer clinical translation. 

 

Source: European Journal of Radiology 

Image Credit: iStock


References:

Wilde D, Schärli B, Ackermann K et al. (2026) Evaluation metrics for synthetic medical imaging. European Journal of Radiology: In Press. 




Latest Articles

synthetic medical imaging, medical image synthesis, synthetic image evaluation, clinical validity, medical imaging AI, radiology AI, image quality metrics Synthetic medical imaging evaluation varies widely. A systematic review highlights task-specific metrics for fidelity, realism, diversity and clinical validity.