Large language models can accelerate medical information work, but rapid synthesis does not give their outputs the status of clinical evidence. A recent article in npj Digital Medicine places their role within a data–information–evidence–practice hierarchy and argues that appraisal, validation and contextual judgement remain essential before information guides care. These systems can retrieve, reorganise and summarise material at scale while also generating unsupported statements, obscuring uncertainty and expanding weakly verified content. Their value therefore depends less on speed than on how clearly workflows separate information processing from the methodological and clinical decisions that establish evidentiary quality.
Why Speed Does Not Equal Evidence
Evidence-based medicine developed to help clinicians manage an overwhelming volume of biomedical information. It introduced structured appraisal, evidence hierarchies and specialised synthesis processes that turn fragmented findings into reviews, guidelines and practical recommendations. Large language models now add a faster layer by supporting literature gathering, summarisation and tailored responses, but they also lower the barriers to producing text that may resemble evidence without meeting its standards.
The proposed hierarchy begins with raw observations, including measurements, laboratory results and study outputs. Once organised and interpreted, these become information. Information becomes evidence only when it meets methodological requirements and undergoes appraisal, validation and synthesis. Evidence then informs clinical practice alongside patient circumstances, professional judgement and local context.
Must Read: LLMs Fall Short on Clinical Reasoning Tasks
Large language models can operate across workflow stages including question formulation, screening, extraction and draft synthesis. Participation in these stages does not automatically confer evidentiary status. Performance can deteriorate when clinical questions are contested, inclusion and exclusion criteria are balanced or causal and cross-disciplinary elements require deeper reasoning. At risk-of-bias and evidence-grading stages, systems may reproduce the language of appraisal without performing the evaluative reasoning that gives conclusions their reliability.
Further risks arise when outputs lack underlying data, when genuine sources are summarised without transparent evidence chains or when uncertainty is not clearly signalled. The resulting text may appear complete even when its foundations remain fragile.
Where LLMs Can Support Evidence Workflows
The most reliable contribution lies in the transition from data to information. Large language models can retrieve, reorganise and present existing material, while human oversight remains necessary when information must be judged as evidence. Responsible workflow design therefore requires clear structural boundaries.
LLM components can support literature queries, title and abstract screening, structured extraction, result tabulation and draft comparisons. These tasks can reduce time spent on information-heavy stages when outputs remain traceable to their sources. Retrieval-augmented systems can ground responses in curated material and reduce fabricated citations, although retrieving evidence is not the same as generating an evidence-based conclusion.
The transition to evidence still depends on methodological quality assessment, applicability, consistency, indirectness, precision and publication bias. These functions require judgement about whether study populations match the patient or setting, whether findings agree across studies and how confidence should be calibrated. The full pipeline, including human governance, therefore becomes the appropriate unit for audit rather than the model output alone.
Large language models may also identify where evidence is absent or incomplete. They can map gaps across published material, organise information from basic science and clinical reports and surface patient preferences scattered across sources. Such maps may be more transparent than answers that imply certainty across uneven evidence. Connections to updated literature feeds, trial registries and regulatory databases could reduce delays in horizon scanning and evidence maintenance, while leaving protocol development, scope definition and critical appraisal under human control.
Human Governance Becomes More Important
Greater automation changes the distribution of clinical work rather than removing the need for expertise. As information becomes easier to access, clinicians spend less effort on retrieval and initial synthesis and more on appraisal, interpretation and accountable decisions. Large language models can present structured comparisons and competing options, but clinicians must decide which evidence is relevant, how risks relate to patient priorities and how uncertainty should be managed.
This responsibility is particularly important where formal evidence is limited. Rare diseases, complex multimorbidity, emerging treatments and resource-constrained settings may lack clear guidance. Decisions in these areas draw on incomplete studies, mechanistic reasoning, local experience and patient values. Large language models can organise that material, but they cannot independently determine what should count as evidence for an individual case.
As agentic systems take on question formulation, search strategy design and report generation, the human role shifts from co-executor to governor. This increases the need for evidential literacy. Clinicians and institutions must distinguish information from evidence at each checkpoint and recognise when an output only resembles a validated conclusion.
Research culture faces a related challenge. Efficiency gains can support deeper methodological work, but publication incentives may reward high volumes of repetitive AI-assisted output. Evaluation therefore needs to preserve data integrity, reproducibility, methodological soundness and demonstrable impact. The central safeguard is disciplined judgement about which outputs have earned the status required for clinical action.
Large language models can support evidence-based medicine when they are limited to functions they perform reliably and when human appraisal governs the transition from information to evidence. Their strongest contribution is rapid retrieval, organisation and synthesis, together with clearer mapping of evidence gaps. They do not remove the need for methodological judgement, contextual validation or accountability to patient values. As systems become more autonomous, clinicians and institutions need stronger evidential literacy, clear workflow boundaries and auditable governance. The central task is to preserve the distinction between information produced quickly and evidence validated sufficiently for practice.
Source: npj Digital Medicine
Image Credit: iStock
References:
Wu Y, Ong AY, Hu K et al. (2026) Fast information and slow evidence in the large language models era. npj Digit Med. https://doi.org/10.1038/s41746-026-02909-7