Foundation models are widening the scope of medical imaging while raising new concerns about privacy. Trained on very large datasets, they can support screening, diagnosis, treatment planning and report generation across different imaging tasks. Their broad utility sits alongside a key risk: learned features may retain demographic and identity-related signals that could support patient re-identification, particularly when combined with external data. The challenge is to preserve clinical value while reducing the risk to confidentiality.

 

Demographic Inference and Re-Identification

Foundation models can extract features that enable high-accuracy prediction of demographic attributes and, separately, support re-identification from retinal and radiology images. That distinction is critical. Characteristics such as age, biological sex, ethnicity or diagnosis do not on their own identify a single person. Identification would usually require those features to be combined with additional identifiers, allowing one individual to be isolated from a wider population. The privacy concern therefore lies less in demographic inference itself than in the possibility that inferred features could be matched with external data.

 

Retinal imaging illustrates why the debate has intensified. High-resolution vascular patterns may in principle act as biometric signals. Even so, that would generally require access to raw images and a way to relink those images to external biometric or reference databases. In many clinical and research settings, such a sequence is unlikely without deliberate intent and substantial auxiliary data. The risk is therefore real, but it is shaped by both technical possibility and practical plausibility.

 

Must Read: Secure and Rapid Encryption Method for Medical Images

 

The issue is also not entirely new. Earlier deep learning models in retinal imaging showed that latent demographic and biometric information could be extracted from medical images. What distinguishes foundation models is their scale, transferability and broader generalisation capacity. Those properties may amplify existing privacy concerns and extend them across a wider clinical and social context.

 

Evidence Gaps and Technical Safeguards

Important questions remain unresolved. Strong predictive performance for demographic traits does not by itself explain how a model could identify an individual person. That gap calls for closer mechanistic investigation, including methods such as feature ablation and attention mapping. Privacy breaches appear most plausible when model outputs are combined with external datasets, but that scenario still requires empirical simulation. Another limitation is that current work often relies on cohorts with limited diversity, which may not capture how privacy risks behave in broader populations.

 

A practical response starts with technical safeguards. Differential privacy mechanisms tailored to foundation models can reduce privacy leakage during large-scale training. Federated pretraining frameworks can support heterogeneous data sources without centralising patient data. Feature disentanglement can help separate clinically useful signals from identifying details. Synthetic data generation with strong privacy guarantees can reduce exposure of real patient records while preserving research utility.

 

Protection also needs to extend across the full model lifecycle. During training, data curation can include scrubbing personally identifiable information, applying privacy-preserving transformations and minimising data to essential clinical features. During inference, input sanitisation can reduce identifiable content, output filtering can remove rare or sensitive outputs, differential privacy can limit leakage and audit logging can support monitoring. Privacy protection depends on multiple controls working together rather than on a single intervention.

 

Governance, Regulation and Clinical Use

Technical measures alone are not enough. Participatory governance, privacy audits and transparent risk reporting are part of the proposed direction. Clinical AI requires coordination across disciplines, including clinicians, data scientists, ethicists, privacy engineers and legal experts. Early collaboration can help identify bias and privacy risks during development, align standards across institutions and assess fairness after deployment. Shared governance frameworks and continuous feedback can keep privacy and fairness central to implementation.

 

Regulation forms another part of the response. The GDPR, the EU AI Act and the Cyber Resilience Act are all relevant to foundation model deployment in healthcare. The GDPR provides principles such as lawful, fair and transparent processing, purpose limitation and rights relating to access, correction, erasure and portability. The EU AI Act introduces a risk-based framework for trustworthy, human-centric AI and places obligations on providers of general-purpose and high-risk systems. The Cyber Resilience Act adds cybersecurity requirements across the lifecycle of digital products, including vulnerability handling and incident reporting.

 

A clinical ophthalmology example shows how these protections may operate in practice. In a multi-institutional consortium, raw data remained on local hospital servers while training used personally identifiable information scrubbing, differential privacy transformation and data minimisation. Federated learning supported cross-site development without centralising data. Differential privacy-SGD limited the influence of individual records. A feature disentanglement module separated disease-relevant signals from sensitive or confounding attributes. During inference, input sanitisation, output filtering, differential privacy and audit logging all contributed to reducing privacy leakage.

 

Foundation models in medical imaging combine significant clinical promise with a clear privacy challenge. Demographic or diagnostic inference does not in itself identify an individual, but the risk increases when learned features can be linked to external identified data. Addressing that risk requires stronger evidence, layered technical safeguards and governance that spans development, deployment and oversight. The proposed direction brings together privacy-preserving model design, multidisciplinary collaboration and regulatory alignment. That approach supports continued progress in medical AI while protecting confidentiality and sustaining trust in clinical use.

 

Source: npj Digital Medicine

Image Credit: iStock


References:

Santos R, DeBuc DC & Somfai GM (2026) Cautious optimism on foundation models in medical imaging balancing privacy and innovation. npj Digit Med; 9, 215.




Latest Articles

foundation models healthcare, medical imaging AI, patient privacy risks, AI re-identification, healthcare data security, differential privacy AI, federated learning healthcare, clinical AI governance Foundation models in medical imaging boost diagnosis but raise privacy risks, including re-identification from data and need for safeguards.