Artificial intelligence models can classify breast ultrasound images accurately when image quality is good, but their performance may fall when noise alters lesion texture, boundaries and intensity. A study in the Journal of Imaging Informatics in Medicine tested this problem using a controlled framework covering three forms of synthetic image degradation. Two deep learning models were compared on breast ultrasound images, with results measured through overall accuracy, malignant recall and false-negative behaviour. The findings show that the effect of noise depends on the model, the type of degradation and its severity. They also show that higher accuracy after noise-aware training does not always mean better detection of malignant lesions. 

 

Must Read: Axillary US Helps Assess Nodal Burden in Breast Cancer 

 

Framework Tests Different Forms of Image Noise 

The evaluation used the public breast ultrasound images dataset, which contains 780 scans from 600 female patients. Images were labelled as normal, benign or malignant. The dataset also provides lesion masks. These were used as additional classification inputs rather than for segmentation. Original scans and mask images were treated as separate samples before the data were divided into training, validation and test sets. 

 

A small custom convolutional neural network was compared with Inception V3, a pretrained model that can capture visual features at different scales. Both models were first tested on clean images. They were then exposed to Gaussian, Poisson and speckle noise at low, moderate and high levels. Gaussian noise represented additive image fluctuations, Poisson noise represented intensity-dependent variation and speckle reflected the granular interference commonly seen in ultrasound. 

 

The models were assessed in three main settings. They were trained and tested on clean images, trained on clean images and tested on noisy images, then trained and tested with matching noise conditions. Mixed-noise testing was also performed. Overall accuracy was reported together with class-specific precision, recall and error patterns. Particular attention was given to malignant recall, which measures the proportion of malignant cases correctly identified. A fall in this measure indicates a greater risk of missed cancers during classification. 

 

Clean-Image Accuracy Does Not Ensure Robustness 

On clean images, Inception V3 performed much better than the custom network. It reached about 91% accuracy and identified 83% of malignant cases. The custom model reached 72% accuracy and identified 49% of malignant cases. It therefore classified nearly half of malignant images as benign. Inception V3 reduced this error, although some malignant cases were still missed. 

 

Performance declined when models trained on clean data were tested with noisy images. Gaussian noise caused the largest accuracy loss for the custom network, while Poisson noise had the strongest effect on Inception V3. At the highest noise level, the custom model lost about 26 percentage points of accuracy under Gaussian noise. Inception V3 lost about 32 points under Poisson noise. Speckle caused smaller changes in overall accuracy, but the ability to identify malignant cases still fell under severe degradation. 

 

The two models also failed in different ways. As noise increased, Inception V3 was more likely to classify malignant images as benign, raising the risk of false negatives. The custom network sometimes showed higher malignant recall under severe Gaussian or Poisson noise. However, this did not reflect better overall discrimination. The model had simply become more likely to classify benign images as malignant, increasing false positives and reducing precision. This changed the balance of errors rather than improving robustness. 

 

Noise-Aware Training Creates New Trade-Offs 

Training each model on the same type and level of noise used during testing improved overall accuracy in several conditions. The strongest gain appeared for the custom network under severe Gaussian noise, where accuracy rose from about 46% to 72%. Inception V3 also improved under the same condition, although the increase was smaller. Across all experiments, noise-matched training raised accuracy by as much as 56.5% at the most severe noise level. 

 

These gains did not always improve the detection of malignant lesions. Under severe Gaussian noise, malignant recall fell for both models after noise-matched training even though accuracy increased. Similar trade-offs appeared with Poisson and speckle noise, depending on the noise level. Higher overall accuracy could therefore hide a greater number of missed malignant cases. 

 

The controlled design also limits how far the findings can be applied to clinical practice. The noise was added synthetically and cannot reproduce every effect of scanner settings, operator technique, patient movement, probe pressure or image reconstruction. The work used only one public dataset. Mask images were included as separate inputs, and related scans and masks were not grouped during data splitting, creating a possible risk of information leakage. The independent value of mask information could not be separated from the results. Testing on external images from different institutions, devices and patient groups remains necessary. 

 

Breast ultrasound AI models do not respond to image noise in the same way, and strong results on clean images do not guarantee reliable performance when quality falls. Inception V3 usually achieved higher accuracy, but severe noise reduced its ability to detect malignant cases. Training with matching noise restored accuracy in several settings while sometimes increasing the risk of missed malignancies. Robustness evaluation therefore needs several noise conditions, class-specific measures and analysis of the full error pattern. The findings provide a controlled benchmark, while real-world testing and external validation remain necessary before clinical use. 

 

Source: Journal of Imaging Informatics in Medicine   

Image Credit: iStock 


References:

Arriz-Jorquiera M, Bakht A, Uysal I et al. (2026) A Noise-Aware Robustness Evaluation Framework for Breast Cancer Classification in Ultrasound Imaging. J Digit Imaging Inform med. https://doi.org/10.1007/s10278-026-02114-8 




Latest Articles

breast ultrasound AI, image noise, artificial intelligence, breast cancer detection, deep learning, Inception V3, medical imaging Image noise reduces breast ultrasound AI accuracy and malignant lesion detection, highlighting the need for robust validation before clinical use.