A report-guided artificial intelligence framework uses detailed descriptions of radiological abnormalities to locate and outline them on chest X-rays. Published in npj Digital Medicine, the approach aims to turn information normally recorded in narrative findings into spatial data for retrospective analysis. Known as Clinical-Findings-Guided Segmentation, or CF2Seg, the framework was evaluated on more than 53,000 examinations covering a range of thoracic abnormalities. It achieved the strongest overall results among the compared methods and remained comparatively stable when labelled data were limited, target conditions were absent from training or image and text inputs were altered. 

 

Turning Findings-Style Text into Spatial Guidance 

Radiologists generally describe abnormalities in the findings section of written reports rather than marking their exact boundaries on images. Hospitals therefore accumulate large collections of chest X-rays and reports, while detailed expert-drawn masks remain comparatively scarce. This limits retrospective uses that require information about where an abnormality is located and how extensive it is. 

 

Must Read: LLM Filtering Supports Whole-Body CT Reference Charts

 

CF2Seg is designed to use detailed clinical findings rather than relying only on a disease label or short descriptive phrase. The text can contain information about an abnormality’s location, appearance, extent and anatomical relationships. It is first adjusted in relation to the image so that relevant descriptions can guide visual analysis. Text information then enters the image-processing pathway before the final segmentation stage and is used again during decoding to refine lesion boundaries across different image scales. 

 

An important qualification concerns the text used for training and evaluation. Most public chest X-ray segmentation datasets do not include original radiology reports paired with expert masks. The accompanying findings-style descriptions were therefore constructed through a retrieval-guided generation process. Radiologist-authored report sentences, image information and the target abnormality were used to generate case-specific descriptions, which then underwent automated assessment and human verification. The resulting text was intended to approximate report-based conditioning rather than serve as a complete substitute for original paired reports. 

 

Performance Under Varied Conditions 

The evaluation compared three forms of text input: category labels, short descriptive phrases and findings-style descriptions. For localised abnormalities, richer text did not consistently improve the comparison models. CF2Seg, however, improved progressively as the input moved from labels to phrases and then to findings. Differences between methods were smaller for broader lung conditions, where simpler wording was often sufficient, but CF2Seg still produced the strongest overall performance. 

 

Across the combined benchmark, CF2Seg achieved the best average segmentation results. Its advantages were most apparent for small, subtle or boundary-sensitive abnormalities, including nodules, calcifications and pneumothorax. Image-only approaches remained competitive for some broader or anatomically more obvious targets, while report guidance provided clearer gains when more precise lesion localisation was required. 

 

The framework was also tested under conditions intended to reflect variation in clinical archives. It was evaluated on abnormalities whose masks were excluded from training, on substantially reduced labelled training data and on images or text that had been deliberately altered. Performance declined across models when annotations were limited, but CF2Seg generally remained more stable. It also showed smaller changes than the comparison approaches when images were affected by noise or blur and when wording in the accompanying text was modified. These evaluations tested whether segmentation could remain useful when annotations, image quality or reporting language differed from the original training conditions. 

 

Potential Use and Important Limitations 

The predicted masks were also assessed for their ability to preserve broad patterns in lesion extent. A focused retrospective assessment examined 100 chest X-rays containing either pleural effusion or atelectasis. The abnormal area identified by the model was compared with that marked by experts, using total image area as a common reference. 

 

Model-derived measurements showed strong correlations with expert-derived estimates for both abnormalities. This suggests that the masks preserved proportional trends in lesion extent beyond conventional measures of pixel overlap. However, the calculation remained a simplified two-dimensional imaging measure rather than a direct estimate of pathological volume or a clinically established severity score. Larger affected areas also showed some proportional deviation. The assessment was intended to explore retrospective quantitative utility rather than provide full clinical validation. 

 

Several limitations remain. Because the findings-style text was generated rather than consistently extracted from original paired reports, the wording may not fully represent routine reporting across different institutions. Incorrect, vague or conflicting descriptions can also provide misleading spatial guidance. Testing was restricted to two-dimensional chest radiographs and adaptation to CT or MRI would require changes to accommodate three-dimensional information. The framework also requires more computation than simpler language-guided approaches. No reader study or prospective workflow evaluation has established whether the generated masks improve clinical interpretation or decision-making. The proposed role therefore remains a human-reviewed tool for annotation and quantitative analysis rather than a standalone diagnostic system. 

 

CF2Seg combines chest X-rays with detailed findings-style radiology text to guide the localisation and outlining of thoracic abnormalities. It performed strongly across a large, mixed benchmark and showed particular advantages for subtle or boundary-sensitive findings. Its results also remained comparatively stable when annotations were scarce or inputs were altered, while its predicted masks preserved broad trends in lesion extent. However, the use of generated findings-style supervision rather than consistently paired original reports is an important limitation. Further validation with original clinical reports, diverse datasets and expert review remains necessary before routine clinical use. 

 

Source: npj Digital Medicine 

Image Credit: iStock 


References:

Reference: Xi S, Hu S, Wang S et al. (2026) Grounding Radiology Report Findings into Medical Image Segmentation. npj Digit Med. https://doi.org/10.1038/s41746-026-03051-0 




Latest Articles

CF2Seg, chest X-ray segmentation, radiology AI, report-guided AI, medical imaging, thoracic abnormalities, npj Digital Medicine CF2Seg AI uses radiology report findings to improve chest X-ray segmentation, enabling precise lesion localisation and robust retrospective analysis.