A medical imaging foundation model has been developed to interpret seven modalities across several clinical specialities. A paper published in The Lancet Digital Health details MerMED-FM, which was trained on about 3.3 million unlabelled images using self-supervised learning and a memory module. Evaluation across radiology, histopathology, ophthalmology and dermatology showed strong performance with limited labelled data, although results varied by task and dataset.
One Model Across Diverse Imaging Modalities
Most medical imaging AI models are designed for one modality or disease, while clinical pathways commonly rely on several types of imaging. MerMED-FM was designed as a vision-only model that can be adapted to chest x-rays, CT, ultrasound, histopathology, colour fundus photography, optical coherence tomography and dermatoscopy. Its pretraining corpus covered 12 specialities and 53 publicly available datasets. No synthetic data, textual annotations or metadata guidance were used during training.
Must Read: Imaging Biomarker Catalogue Aims to Improve Reuse
Self-supervised learning allowed the model to learn from unlabelled images, reducing dependence on large datasets annotated by clinical experts. A memory component stored compact representations from earlier samples and refreshed them during training. This approach was intended to preserve learning across distinct image types and limit the loss of previously acquired information when new modalities were introduced. Modality-aware and speciality-aware sampling also sought to prevent larger domains from dominating training.
After pretraining, the model was fine-tuned for disease classification and tested on 26 public and 5 private datasets. These covered conditions including pneumonia, tuberculosis, lung carcinoma, breast cancer, liver disease, retinal disorders, glaucoma, skin lesions and several tissue types. Evaluations used different proportions of labelled data, with the main comparisons based on 10%. Saliency maps were reviewed by clinicians in radiology, pathology and ophthalmology to assess whether influential image regions were clinically plausible.
Strong Results with Limited Labelled Data
With 10% of the available labels, MerMED-FM achieved the highest overall mean performance among the models tested. Results were strongest in optical coherence tomography, histopathology and CT, while performance was also robust for chest x-rays, ultrasound, fundus photography and dermatoscopy. The model generally retained much of its full-data performance with limited labels, and gains from increasing labelled data were greater between 10% and 30–50% than between 50% and full labelling.
Performance nevertheless differed across individual tasks. MerMED-FM exceeded specialist or multispecialty comparators in several CT, chest x-ray, ultrasound, histopathology and ophthalmology datasets. It performed particularly strongly on retinal imaging and remained competitive in real-world private datasets. However, it did not lead every comparison. A multispecialty model performed better for one pneumonia dataset and one breast ultrasound dataset, while specialist models retained advantages in some dermatology and pathology tasks. On a private chest x-ray dataset containing varied pathologies, all models performed less well than on standard public datasets and no model was clearly superior.
Out-of-distribution testing produced results within a few percentage points of the best comparator across the seven modalities, but MerMED-FM was not consistently the strongest model in every setting. The findings therefore support local validation and calibration before clinical use. Fairness testing was possible only for datasets with demographic information. Differences by age and sex were moderate in fundus photography and small in optical coherence tomography and dermatoscopy, leaving broader demographic performance unresolved.
Potential Deployment Requires Further Validation
A single imaging backbone could reduce the need to deploy and maintain separate models for different departments. It could be adapted for cross-modality triage, multidisciplinary cancer assessment, systemic disease screening or joint clinical pathways. The same architecture could also support consistent governance where AI is used across several specialities and labelled local data are limited. The model is therefore positioned as a generalist platform rather than a universal replacement for specialist systems across complex hospital environments.
A hybrid deployment approach remains necessary. MerMED-FM could provide screening or triage across multiple modalities, with specialist models used where peak performance is required for a narrowly defined task. Additional fine-tuning with larger labelled datasets could also improve performance for dedicated applications. The model currently processes two-dimensional images and does not support volumetric imaging. It has not been evaluated for combined reasoning across different images from the same patient, prospective clinical use or health economic outcomes.
Further development would need to address segmentation, prognosis and report generation, alongside testing on low-quality images and broader fairness assessment. Regulatory requirements also remain unresolved. Patient data used in development were de-identified, and diverse populations were included in training and evaluation, but local testing remains important because imaging protocols, equipment, disease presentation and patient characteristics can vary between clinical settings.
MerMED-FM shows that one vision-only foundation model can learn across seven medical imaging modalities while remaining competitive with specialised systems. Its strongest advantage appears when labelled data are scarce, particularly in radiology and ophthalmology applications. The results also show that generalist performance is not uniform, with specialist models retaining an advantage in some tasks and performance declining on more complex real-world data. Clinical adoption would require local validation, calibration, prospective assessment and regulatory review. The model provides a basis for further development of cross-speciality imaging support rather than a ready replacement for existing diagnostic tools.
Source: The Lancet Digital Health
Image Credit: iStock
References:
Zhou Y, Quek C, Zhou J et al. (2026) MerMED-FM: Multimodal, Multi-Disease Medical Imaging Foundation Model. The Lancet Digital Health, 8: 101007.