A deep learning model using PET/CT images alone estimated the likelihood that indeterminate pulmonary nodules were malignant with performance similar to established clinical assessment. The single-centre retrospective evaluation, published in European Radiology, compared the AITO-PETCT-MP model with the Herder risk model and seven experienced clinicians. The dataset included 533 nodules from 436 patients treated at a Dutch academic centre, with malignancy confirmed by tissue diagnosis and benign status supported by at least 2 years of cancer registry follow-up. Performance was then assessed in a separate test set of 161 nodules, about half of which were malignant.
Imaging-Only Risk Estimation
AITO-PETCT-MP combines information from computed tomography and positron emission tomography images centred on each nodule. It produces a malignancy probability without using clinical variables. The Herder model, by contrast, combines imaging findings with factors such as age, smoking history and previous cancer. This difference allowed a direct comparison between an imaging-only approach and a guideline-based method that depends on both imaging and patient information.
The model was developed using scans collected over several years. The cases came from a biopsy registry, electronic health record searches and automated nodule detection. These sources were combined to create a more balanced set of malignant and benign nodules, as the biopsy registry contained a much higher proportion of cancers. Several trained versions were combined into the final model, while a separate test set remained unused during development.
Must Read: CT Density Pattern Helps Assess Small Lung Nodules
On the test set, the deep learning model achieved an area under the curve of 0.78, compared with 0.73 for the Herder model. Statistical testing showed that its performance was non-inferior to the guideline-recommended method. The average clinician reached 0.80, while results for individual readers ranged from 0.74 to 0.87. The imaging-only model therefore performed within the range of experienced clinical readers, although the clinicians also had access to limited clinical details and earlier CT scans where available.
Follow-Up Decisions Differed
The main differences emerged when malignancy probabilities were translated into British Thoracic Society management categories. These categories direct patients towards CT surveillance, biopsy or possible direct treatment when biopsy is not feasible. Although overall diagnostic performance was similar, the three approaches did not distribute patients across these options in the same way.
The Herder model placed no malignant nodules in the surveillance group. AITO-PETCT-MP assigned 15 malignant nodules to surveillance, while the average clinician assigned 18. This meant that the Herder model captured more cancers within the biopsy and treatment categories. However, it also directed substantially more benign nodules towards possible treatment.
Among 81 benign nodules in the test set, the Herder model assigned 26 to potential direct treatment. The deep learning model and the average clinician each assigned only 3 benign nodules to that category. The model and clinicians therefore showed greater specificity when recommending treatment, but they also accepted a higher risk that malignant nodules would remain under imaging follow-up.
Across the individual clinicians, 41 different malignant nodules were placed under surveillance by at least one reader. Most were primary lung cancers and the remainder were metastases. These findings show that similar overall accuracy can lead to different management patterns, particularly when balancing the risk of treating benign disease against the risk of delaying further investigation of malignancy.
Broader Validation Remains Necessary
A simulated second-reader assessment explored whether combining evaluations could improve performance. Pairing clinicians with either another clinician or the deep learning model produced a small increase in diagnostic accuracy and reduced variation between readers. The improvement was not statistically significant. The possible value of AITO-PETCT-MP as a second reader therefore remains uncertain and requires evaluation in settings where clinicians can view its predictions during assessment.
The evaluation also had several limitations. All data came from one academic centre and was collected retrospectively. The dataset combined cases from three sources with very different malignancy rates, so it may not represent a consecutive clinical population. The model may also have learned differences linked to how and when the scans were collected.
The Herder criteria were applied to a broader range of cases than in their original development, including some patients with previous lung cancer and larger nodules. Management decisions were also simplified into guideline categories, while clinical practice may take account of other health conditions and benign findings that influence follow-up. Some missing clinical information, including smoking status and cancer history, was set to negative because the data were pseudonymised.
Further evaluation should use consecutive multicentre datasets from different institutions and regions. The effects of small lesion measurement and respiratory motion were not specifically assessed. Direct clinical testing is also needed to determine whether access to model predictions improves decision-making and reduces variation between clinicians.
The imaging-only AITO-PETCT-MP model estimated pulmonary nodule malignancy with performance comparable to the Herder model and experienced clinicians. Its management recommendations, however, followed a different pattern. The Herder model sent more benign nodules towards possible treatment, while the deep learning model and clinicians placed more malignant nodules under CT surveillance. These differences matter because similar diagnostic performance does not necessarily produce similar follow-up decisions. Wider multicentre validation and direct assessment of model-assisted reading are required before the model’s clinical role can be established.
Source: European Radiology
Image Credit: iStock
References:
Leijten L, Aarntzen EHJG, Verhoeven RLJ et al. (2026) Deep learning-based malignancy probability estimation of pulmonary nodules in PET/CT imaging. Eur Radiol. https://doi.org/10.1007/s00330-026-12744-9