Brain MRI differential diagnosis depends on both imaging expertise and the quality of information given to AI tools. A 2026 investigation published in Radiology assessed how reader experience affects large language model support in challenging brain MRI cases. Neuroradiologists, radiology residents and neurology or neurosurgery residents created imaging descriptions and initial differential diagnoses. Three large language models then generated ranked diagnoses from those descriptions. The findings show that expert inputs improve model accuracy, while less experienced readers gain more from AI assistance.

 

Experience Improves Model Input

The dataset included 40 brain MRI scans from a single academic centre, with diagnoses confirmed either histopathologically or through independent agreement by at least two neuroradiologists using available clinical and follow-up information. Cases needed sufficient diagnostic complexity for use in neuroradiology subspecialty examinations. Exclusions covered non-brain MRI imaging, inadequate image quality and unconfirmed diagnoses. The scans came from TUM University Hospital in Munich and covered a broad neuroradiological spectrum, including neoplastic, vascular, degenerative, infectious, toxic or metabolic, congenital or developmental and other abnormalities.

 

Reader groups differed markedly in training and radiological experience. The neurology and neurosurgery residents had no formal radiological training but had clinical experience in their specialties. Radiology residents had early radiology experience dedicated to neuroradiology training. Neuroradiologists had the highest overall radiology and neuroradiology experience.

 

Each reader reviewed brain MRI scans with condensed medical history and demographic characteristics, then produced a textual description of the main imaging finding. Readers also listed up to three differential diagnoses and recorded confidence. These imaging descriptions then became prompts for GPT-4.1, Gemini 2.5 Pro and DeepSeek-R1, which generated ranked differential diagnoses. During the assisted condition, readers reviewed GPT-4.1 suggestions generated from their own descriptions and then revised their top-three diagnoses and confidence scores.

 

Must Read: Two-Minute Deep Learning Brain Quantitative Mapping

 

Model Accuracy Rises with Expert Descriptions

Large language model performance followed the experience gradient of the reader input. Across all three models, differential diagnoses achieved the highest top-three accuracy when based on imaging descriptions from neuroradiologists. Accuracy was lower with radiology resident input and lowest with neurology or neurosurgery resident input. Gemini 2.5 Pro showed the strongest overall performance across reader groups, with GPT-4.1 and DeepSeek-R1 producing similar results in the comparisons described.

 

The pattern links model output to the quality of the human-generated description. Neuroradiologist descriptions received the highest ratings for correctness and completeness. Radiology resident descriptions showed high correctness but more moderate completeness. Neurology and neurosurgery resident descriptions had the lowest ratings for both dimensions. Statistical modelling found positive associations between radiological experience and both correctness and completeness of imaging descriptions.

 

This result matters because the models received textual information rather than directly interpreting the full brain MRI examination. Their diagnostic performance depended on the clarity, precision and completeness of the findings supplied by readers. The more experienced the reader, the stronger the input available to the model. However, that same expertise also reduced the room for assisted improvement, creating a practical mismatch between the conditions that optimise model performance and the settings where readers gain most from support.

 

Reader Gains Are Greatest Among Less Experienced Clinicians

GPT-4.1 assistance improved top-three diagnostic accuracy across reader groups, but the magnitude of improvement declined as experience increased. Neurology and neurosurgery residents showed the largest mean absolute gain, rising from lower unassisted accuracy to substantially higher assisted accuracy. Radiology residents also gained meaningfully from assistance. Neuroradiologists started from the highest unassisted accuracy and showed only a small increase after GPT-4.1 support.

 

The cumulative link mixed model found a negative association between reader experience and diagnostic benefit from large language model assistance. In practical terms, each additional year of radiological experience corresponded to a smaller benefit from the assistive model. Reader confidence followed a similar pattern. GPT-4.1 assistance increased confidence across all groups, with the largest improvement in neurology and neurosurgery residents, a smaller increase in radiology residents and only minimal change among neuroradiologists, whose confidence was already high.

 

Reader interaction with model suggestions also shaped outcomes. Across groups, GPT-4.1 assistance caused a small number of correct answers to change to incorrect ones, but a greater number of incorrect answers changed to correct ones. Some readers rejected correct model suggestions, while fewer abandoned a correct diagnosis in favour of an incorrect suggestion. These transitions show that assistance depends not only on model output but also on how readers judge, accept or disregard suggested diagnoses.

 

Large language model support in brain MRI differential diagnosis shows an expertise paradox. More experienced readers provide stronger imaging descriptions, allowing models to perform better, yet those readers gain less because their baseline diagnostic accuracy is already high. Less experienced clinicians gain more from assistance, although their final performance remains below that of neuroradiologists. The findings place human-AI interaction at the centre of clinical value, showing that model accuracy alone does not determine usefulness. Successful deployment depends on input quality, reader expertise and the capacity to evaluate model-generated suggestions.

 

Source: Radiology

Image Credit: iStock


References:

Schramm S, Le Guellec B, Topka M et al. (2026) Performing Best When Needed Least: Reader Experience Shapes Accuracy Gains in Large Language Model–assisted Brain MRI Differential Diagnosis. Radiology; 319:2.




Latest Articles

Brain MRI, AI diagnosis, neuroradiology, radiology AI, GPT-4.1, Gemini 2.5 Pro, DeepSeek-R1, differential diagnosis, MRI interpretation Brain MRI AI diagnosis improves with expert imaging descriptions. Study shows neuroradiologists boost LLM accuracy most.