Automated labelling of radiology records can reduce the time required to prepare datasets, but it can also add new costs through model use, energy consumption and operational complexity. A recent analysis published in Insights into Imaging compared manual, rule-based, large language model and hybrid workflows for labelling structured CT pulmonary angiography records related to pulmonary embolism. The comparison focused on accuracy, processing time, direct cost and estimated carbon emissions. The clearest result was not simply that language models can accelerate annotation. Workflow design mattered more. A rules-first approach, in which only fields that failed rule-based extraction were sent to a language model, delivered stronger accuracy while using fewer computational resources.

 

Must Read: LLM Reasoning Format Shapes Radiology Accuracy

 

Rules Reduce the Need for Model Inference

The comparison used 2923 structured CT pulmonary angiography records created with a pulmonary embolism template. Each record contained 14 fields, covering clinical and technical information such as pulmonary embolism presence and burden, right-heart strain, image quality and key impression items. This structure made the dataset suitable for testing whether deterministic rules could handle routine fields and reserve language model use for more difficult cases.

 

Four labelling routes populated the same field structure. Manual labelling relied on eight radiologists using an electronic form. A rule-based extractor mapped predefined wording in the template to allowed categories and flagged entries that were missing or invalid. A language-model-only route sent all fields to 22 models, including open-weight and proprietary options. The hybrid route applied the rule-based extractor first, retained valid outputs and sent only invalid fields to a language model.

 

This design limited model use to the parts of the record where rules were insufficient. It also created a direct comparison between broad model deployment and targeted model deployment. The same language models could be assessed under both conditions, making the effect of workflow design clearer than a comparison based only on model selection.

 

Hybrid Labelling Improves the Operational Balance

Manual labelling achieved high accuracy but required the most labour. Across the full dataset, radiologist labelling took about 33 hours and cost more than €1200 in direct labour costs. The rule-based extractor was much faster and reached a lower, but still substantial, level of report-level accuracy. Its full-cohort runtime was below one second in repeated testing, reflecting the speed advantage of deterministic extraction when structured fields follow expected patterns.

 

Language-model-only workflows reduced direct cost and time compared with manual labelling, but they did not match the strongest accuracy results. Their mean report-level accuracy was lower than that achieved by radiologists. They also required every field from every record to pass through a language model, including fields that could already be handled by rules. This created avoidable model calls and additional resource use.

 

The hybrid route produced the strongest overall balance. Mean report-level accuracy reached 98.5%, exceeding both manual labelling and language-model-only extraction. It also reduced the cohort-level time, direct cost and estimated emissions compared with language-model-only processing. Across matched model configurations, moving from full language-model extraction to hybrid processing cut median processing time from several hours to about one hour. Median direct cost also dropped substantially, and estimated emissions for open-weight models fell from under 1 kg of CO2 to a much lower level.

 

These results point to a practical operational pattern. In structured radiology records, rules can manage routine extraction, while language models can address fields that fall outside the expected format. This limits computational work without removing language model support from the workflow.

 

Model Choice Still Affects Sustainability

Workflow design had a major effect, but model selection still influenced the trade-off between performance and resource use. Some models achieved similar accuracy with very different time, cost and emissions profiles. Mid-sized open-weight configurations offered a favourable balance, combining strong labelling performance with shorter runtimes, very low direct electricity cost and measurable emissions. Smaller models used fewer resources but had lower report-level accuracy. A proprietary model also performed efficiently for time and accuracy, but provider-level emissions data were unavailable.

 

The comparison between open-weight and proprietary models was therefore incomplete for environmental assessment. Local open-weight models allowed direct measurement of incremental energy during inference on the workstation used in the experiment. Proprietary models provided latency and token-based cost data, but not emissions data. This limitation matters for procurement and governance decisions, because operational sustainability cannot be assessed fully when energy and carbon data are unavailable.

 

The resource estimates also had a defined boundary. They covered marginal execution during the experiment and did not include model training, cooling, embodied hardware emissions, continuous infrastructure baselines, software development or ongoing maintenance of rule-based and hybrid pipelines. These omissions do not invalidate the comparison between workflows, but they do limit how far the numbers can be used as total lifecycle estimates.

 

The dataset itself also came from one regional network using a single structured pulmonary embolism template. That setting favoured rule-based extraction because fields were already organised around expected categories. Different results may occur with less structured free text, other modalities or other clinical domains. Even so, the main operational lesson remains relevant: unnecessary language model calls increase resource use, and rules can reduce that burden when record structure allows.

 

Rules-first hybrid labelling offered the best combination of accuracy and resource efficiency for structured CT pulmonary embolism records. Manual labelling remained accurate but required far more labour, while language-model-only processing reduced time and direct cost at the expense of lower accuracy. The hybrid route preserved the value of language models while limiting their use to fields that needed additional handling. Sustainable use of language models in radiology therefore depends on more than model performance. It also requires careful routing, clear measurement of resource use and workflow designs that avoid unnecessary computation.

 

Source: Insights into Imaging

Image Credit: iStock


References:

Fink MA, Bischoff A, Atsiatorme E et al. (2026) How green are large language models for radiology report labelling? Comparing human, rule-based and hybrid workflows. Insights Imaging; 17, 142.




Latest Articles

MRI Muscle Markers, Cardiometabolic Risk, Paraspinal Muscle MRI, Intermuscular Adipose Tissue, Lean Muscle Mass, Metabolic Health MRI muscle markers, paraspinal muscle MRI, cardiometabolic risk, intermuscular adipose tissue, IMAT, lean muscle mass, LMM, metabolic risk, hypertension, dysglycaemia