Vision-language models showed uneven performance when classifying suspicious axillary lymph nodes on ultrasound in patients with breast cancer, while experienced radiologists remained the most accurate readers. The retrospective comparison, accepted by the American Journal of Roentgenology, used images from a tertiary academic hospital in China and included 718 biopsied nodes. Three models assessed each node as metastatic or non-metastatic from ultrasound images alone, and their results were compared with those of radiologists at different experience levels. The best-performing model achieved accuracy similar to that of inexperienced radiologists, but it detected fewer metastatic nodes and produced fewer false-positive findings. The results support supervised decision support rather than independent interpretation. 

 

Must Read: Image Noise Reveals Weaknesses in Breast Ultrasound AI 

 

Images, Readers and Models Compared 

The analysis included patients with primary invasive breast cancer who underwent ultrasound-guided needle sampling of a suspicious axillary lymph node during initial clinical assessment. One target node was selected for each patient, with greyscale imaging available in all cases and Doppler imaging available in most. Biopsy findings were used to classify each node as metastatic or non-metastatic. The sample was almost entirely female, and nearly half the sampled nodes were metastatic.  

 

Two experienced and two less-experienced breast imaging radiologists reviewed the saved images. The less-experienced readers had encountered breast and axillary ultrasound during residency, while the experienced readers had more than a decade of post-training experience. Each reader first assessed the images independently, after which the readers within each experience group resolved disagreements and produced a consensus classification.  

 

GPT-5.2, Gemini-3-Pro and Claude-Opus-4.5 received the same image-only prompt and were required to provide one of two diagnostic labels. The main comparison assessed accuracy, sensitivity and specificity for the two radiologist groups and the three models. Additional testing explored whether performance changed when structured descriptions of ultrasound features or clinicopathological information were added. The models were also tested repeatedly over time to assess the consistency of their outputs. These steps allowed comparison of overall performance, error patterns and repeatability rather than relying on a single accuracy measure. 

 

Experience Produced the Strongest Overall Results 

Experienced radiologists achieved the highest overall accuracy and the highest specificity. Their assessments correctly classified 83% of nodes and were particularly effective at identifying non-metastatic findings. Less-experienced radiologists achieved lower overall accuracy but showed the highest sensitivity, identifying most metastatic nodes while also producing more false-positive assessments. GPT-5.2 performed best among the three models and returned a valid answer for every case. Its overall accuracy was close to that of the less-experienced readers, with no significant difference between them.  

 

However, the pattern of errors differed. The model was less sensitive, meaning that it missed more metastatic nodes, but it was more specific and therefore generated fewer false-positive classifications. Compared with experienced radiologists, it remained less accurate and less specific. Gemini-3-Pro showed lower accuracy than GPT-5.2, while Claude-Opus-4.5 performed least well overall. Experienced radiologists were more accurate than every model. Less-experienced radiologists were also more accurate than Claude-Opus-4.5.  

 

Adding structured text, imaging descriptions or clinicopathological details did not consistently improve performance. More complex input therefore did not produce a reliable advantage. Results varied according to the model and the information supplied, rather than improving steadily as more data were added. Performance depended not only on the clinical task but also on the model and input format used. 

 

Supervised Use Requires Caution and Repeat Testing 

The different error profiles shape the potential role of these models. Less-experienced radiologists favoured sensitivity, which reduced missed metastases but increased false-positive assessments and could lead to more biopsies. The strongest model showed the opposite tendency, reducing false positives but missing more metastatic nodes. This balance limits the case for autonomous interpretation or independent biopsy recommendations.  

 

A supervised role may be more appropriate. The model could, for example, prompt an additional review when a less-experienced radiologist classifies a node as positive. Even this narrower use would require careful testing because any improvement in specificity could be offset by lower sensitivity. Missed nodal metastases could affect staging and treatment decisions involving surgery, regional nodal irradiation or systemic therapy. Repeat testing also showed that model behaviour was not fully stable over time. GPT-5.2 maintained substantial agreement across repeated runs, but its later performance shifted towards markedly lower sensitivity and higher specificity. The reason could not be identified because the models were accessed through evolving service-side systems whose changes are not externally visible.  

 

The findings therefore support version-aware evaluation, repeat validation before implementation and continued monitoring after deployment. They also argue against assuming that a model tested once will continue to behave in the same way in later use. Clinical workflow studies remain necessary because the comparison used selected static images rather than live examinations with radiologist interaction. 

 

Vision-language models did not match experienced radiologists in assessing suspicious axillary lymph nodes on ultrasound. The strongest model reached similar overall accuracy to less-experienced readers but achieved this through a different balance of errors, with fewer false positives and more missed metastases. That trade-off prevents support for autonomous use. A supervised decision-support role may warrant further evaluation, particularly as a prompt for reviewing positive assessments by less-experienced radiologists. Any clinical use would require validation in routine workflows, repeat testing over time and monitoring for changes in performance. 

 

Source: American Journal of Roentgenology 

Image Credit: iStock 


References:

He P, Chen C, Dong HD et al. (2026) Vision Language Models for Ultrasound Assessment of Suspicious Axillary Lymph Nodes in Breast Cancer. American Journal of Roentgenology: New Articles.




Latest Articles

breast ultrasound AI, axillary lymph nodes, vision-language models, breast cancer imaging, metastatic lymph nodes, radiology AI, ultrasound diagnosis Vision-language models showed uneven accuracy in detecting metastatic axillary lymph nodes on breast ultrasound compared with experienced radiologists.