Clinical AI tools are increasingly judged not only by what they know, but by how reliably they reason through the order of events. An informal prompt-based evaluation published by HLTH raises concerns about how large language models (LLMs) handle chronology in clinical reasoning. The exercise centred on a cough vignette in which enalapril was prescribed after the symptom had already begun, making the medication an impossible trigger. The correct cause was laryngopharyngeal reflux, also known as silent gastroesophageal reflux disease (GERD). In September 2025, eight models blamed enalapril despite the timeline. In March 2026, after model updates, only Gemini was consistently correct, while Grok succeeded only under specific conditions.

 

A Simple Prompt with a Critical Sequence
The prompt asked each model to act as a pulmonologist and determine the trigger of the cough, requesting clarification until it could identify the cause with confidence. The patient had asthma and recent weight gain, then developed a cough two months earlier. She later visited her doctor for a routine check-up, forgot to mention the cough, received a coincidental hypertension diagnosis and left with an enalapril prescription.

 

Must Read: AI Use in Medication Reconciliation Remains Limited

 

The chronology was central. Enalapril could not be the trigger because the cough began before the first prescription. The expected approach involved asking follow-up questions about symptoms such as heartburn or regurgitation, then recognising that silent GERD could still be present even without classic symptoms. The September run did not follow that route. The models did not ask for clarification and instead selected the familiar medication side-effect explanation. Some still failed to correct after being prompted to construct a timeline and challenge their initial answer. The error did not require rare medical knowledge; it required a basic check that cause preceded effect. That makes the vignette relevant to routine clinical reasoning rather than only specialist diagnostic performance. It also shows how a short patient history can expose a basic ordering problem in AI-generated clinical logic.

 

Anchoring Overrides Clinical Chronology
The most important failure was an overreliance on a familiar association between dry cough and ACE inhibitor treatment. The models appeared to treat that pattern as more persuasive than the sequence of events. The result was a causal claim that contradicted the timeline in the prompt. The problem was therefore not only diagnostic accuracy, but also the way an LLM weighs pattern recognition against temporal evidence. In the September run, the stated failure modes included temporal reasoning failure, deprioritised contradiction, confirmation bias, anchoring bias, overweighted pattern recognition and prior bias linked to training data.

 

The safety concern is concrete. Some models recommended stopping enalapril, although the medication could not have caused a cough that started before treatment. A patient following that advice without informing a physician could stop an indicated therapy while the real cause remains unaddressed. The same reasoning pattern could create greater risk if applied to another drug with more immediate consequences, such as warfarin in a patient with a potentially fatal blood clot. The point is not that one vignette can define model performance, but that confident reasoning can look clinically plausible while missing a basic chronological constraint entirely. For healthcare organisations considering AI-supported triage, symptom checking or decision support, such errors require attention to verification steps rather than output fluency alone.

 

Multi-Agent Debate Gives Mixed Signals
The March rerun included 13 named LLMs after new releases and updates. Gemini gave the correct answer consistently by checking the timeline early and refusing to let the ACE inhibitor association override chronology. Grok showed a different pattern. It could reach the correct answer when its multi-agent debate process created a role focused on temporal verification before the system reached its confidence threshold.

 

That distinction matters because multi-agent debate does not automatically create safer reasoning. Across the 40 Grok runs, the system sometimes generated several internal agents that argued, critiqued and converged, but failed when none of them focused on chronology. When a temporal verifier appeared early enough, performance improved. Success varied strongly with debate duration: 0% when the debate lasted under 25 seconds, 2.5% when it lasted under 33 seconds and 97.5% when it lasted over 33 seconds. The result suggests that architecture, compute time and explicit verification roles may matter as much as model scale. A multi-agent process still needs the right dissenting function at the right moment. In clinical settings, that function could include timeline verification, medication review and a safety check before any treatment recommendation.

 

The prompt-based evaluation does not establish definitive comparative performance across clinical AI systems, but it gives a clear cautionary signal. Several LLMs produced a plausible answer while missing a basic chronological contradiction, and some recommended stopping a medication that could not have caused the symptom. For healthcare organisations considering AI-supported symptom assessment, triage or decision support, the priority is therefore not only model capability, but also the design of verification steps. Clinical AI tools need safeguards that check timelines, challenge premature confidence and prevent familiar patterns from overriding causality.

 

Source: HLTH

Image Credit: iStock  




Latest Articles

clinical AI reasoning failure, LLM temporal reasoning, large language model clinical safety, AI diagnostic errors, anchoring bias AI, clinical decision support AI, LLM chronology failure, healthcare AI safety LLMs fail basic clinical timeline reasoning in cough vignette test, revealing anchoring bias and chronological errors that pose real patient safety risks.