HealthManagement, Volume 26 - Issue 3, 2026
Healthcare AI can improve visible productivity while creating hidden work for clinicians. Time saved on documentation may be offset by the effort needed to detect, assess and correct unreliable output. This verification burden also shifts clinical work from producing notes, messages or decisions to reviewing and correcting machine-generated content, often without matching changes in roles, training, governance or workload models.
Key Points
- AI can save time while creating hidden verification work for clinicians.
- After-hours EHR work did not fall despite documentation time savings.
- Unreliable AI output demands detection, evaluation and correction effort.
- Clinical roles shift from producing content to reviewing machine-generated output.
- Governance often misses the tools clinicians already use and trust.
The Problem Healthcare Leaders Are Not Measuring
The largest multi-site study of AI-powered clinical documentation to date tracked 8,581 clinicians across five U.S. academic health systems for over two years. AI scribe adoption was associated with 13.4 fewer minutes of total electronic health record time per day. But the same study found no significant reduction in after-hours EHR work (Rotenstein et al. 2026). If AI scribes were eliminating the documentation problem, after-hours work should have declined. It did not.
That gap points to something most healthcare organisations are not yet accounting for. Generative AI is built on probabilistic architectures that produce useful output alongside factually unreliable output: hallucinations, omissions, outdated claims and assertions that cannot be verified without additional effort (Ji et al. 2023). In accountable clinical work, that unreliability does not disappear. It creates verification labour: the detection, evaluation and correction work required before AI-assisted output can be signed off, acted on or passed forward. That labour is real. It consumes time and cognitive resources. And in most organisations, it is invisible to operational dashboards because current metrics track what AI produces, not what humans must do to make that output operational (Kamara 2026).
Underlying research article argues that the central operating risk in healthcare AI is not slow adoption. It is the accumulation of unmeasured verification cost hidden beneath reassuring productivity metrics and the unacknowledged professional role shift that follows (ibid).
Adoption Is Not the Problem. The Wrong Metrics Are
The prevailing narrative frames healthcare AI as an adoption challenge. The evidence tells a different story. A pragmatic randomised trial of 238 outpatient physicians found adoption rates of only 29–34% for AI scribes, even in a structured, supported trial setting (Lukac et al. 2025). The largest multi-site study reported that only 32% of adopters used the tool for more than half of their visits (Rotenstein et al. 2026). The American Medical Association’s 2026 survey found 81% of physicians using AI in their practices, yet 92% wanting more education and 88% citing safety and efficacy validation as preconditions for adoption (American Medical Association 2026).
These are not the numbers of a profession resisting innovation. They are the numbers of a profession exercising rational caution toward tools that do not yet meet the reliability standard that clinical accountability requires.
Many healthcare AI tools that organisations are deploying, testing and studying, are typically vendor-specific wrappers of narrower AI models with constrained capabilities. These contrast with the frontier models, which offer advanced reasoning, broader functionality and constant improvement, and which clinicians can privately access. Cross-sector evidence shows that 78% of AI users bring their own tools to work (Microsoft 2024), and a 2026 hospital survey found that unsanctioned AI tools are broadly present across health systems, often adopted for workflow speed and superior functionality (Wolters Kluwer 2026). Governance and adoption measurements cover the tools the organisation provides. The AI tools professionals trust are often outside that perimeter and therefore invisible to operational analytics.
The metric problem compounds this. A separate study confirmed that clinicians perceive greater documentation time reductions than objective measurement supports (Adler-Milstein et al. 2026). When organisations report AI productivity gains, they are typically measuring gross output acceleration. They are not measuring the verification and correction effort those outputs require before they become clinically usable. That distinction is the difference between apparent efficiency and realised productivity.
Why Verification Burden Is an Operating Cost
The mechanism is straightforward. AI output in clinical settings carries factual unreliability across three dimensions: hallucination, where content is fabricated; omission, where clinically relevant information is missing; and unverifiability, where claims cannot be confirmed without additional checking.
Each of these unreliability patterns generates a corresponding verification demand on the clinician: detection effort (noticing something may be wrong), evaluation effort (judging severity and scope) and correction effort (fixing or replacing the output). This is not routine quality assurance. It is additional labour created specifically by the introduction of AI into an accountable workflow. Qualitative research confirms that clinicians experience AI tools not as simple time-savers but as tools that demand active review, contextual correction and ongoing judgment about what the AI missed or misrepresented (Van Tiem et al. 2026).
In patient messaging, an inbox-drafting pilot found AI drafts generated for approximately 80% of messages that ultimately received no response, adding review burden without corresponding clinical value (Mandal et al. 2025).
When verification labour is not measured, productivity claims overstate gains. The Rotenstein finding, where 13.4 minutes of EHR time savings coexist with unchanged after-hours work, is consistent with verification burden absorbing a significant portion of apparent time savings (Rotenstein et al. 2026).
The Role Shift No One Is Naming
Verification burden does not only affect productivity accounting. It changes what clinical work is with artificial intelligence. When AI generates a draft note, the physician’s task shifts from authoring to supervised review. When AI drafts a patient message, the task shifts from composing to safety-editing. When AI suggests a triage pathway, the task shifts from clinical reasoning to override judgment.
These are not minor workflow adjustments. They represent a structural change in professional identity: from autonomous producer of clinical output to accountable verifier of machine-generated output. Research on physician responses to AI-based clinical decision support has identified professional identity threat as a significant factor in resistance, with clinicians perceiving AI as challenging their expertise, autonomy and professional standing (Jussupow et al. 2022). A scenario-based experimental study further linked the design of AI processes to perceived identity threat, finding that how AI is integrated into clinical workflows directly shapes whether clinicians experience the technology as supportive or threatening (Ackerhans et al. 2025).
Because verification labour is not measured, this identity shift is not acknowledged. Roles are not redesigned to reflect it. Training programmes do not address it. Workload models do not account for it. The professional absorbs the difference through discretionary effort, defensive checking and quiet frustration. Where organisations anticipate that AI will continue to reshape professional responsibilities, structured approaches to assessing and managing occupational identity transitions can help leadership plan role redesign before it becomes a silent resistance and retention crisis.
Bringing Shadow AI Out of the Shadows
Most governance articles recommend controlling unsanctioned AI use. That recommendation misreads the problem. Clinicians are not using external AI tools out of negligence. They are using them because those tools offer capabilities that sanctioned, narrower products do not match. Hospital survey data confirms that unsanctioned AI tools are present across health systems, adopted by clinicians seeking better functionality and faster workflows (Wolters Kluwer 2026). Shadow AI is not a compliance failure. It is an unrecognised R&D programme that the organisation is not yet benefiting from.
Rather than locking down what clinicians have learned, organisations should be capturing it. Identify the AI champions who are already using advanced models effectively. Establish reviewing structures that assess which workflows, models and tools should be brought into supported, governed use. Empower those champions to share methods and teach colleagues.
What Leaders Should Commission in the Next 90 Days
Measure verification burden as a cost of AI output. Choose one or two live workflows, starting with regular documentation and patient messaging. Map where AI-assisted work requires review, correction, override and escalation effort. Until this is measured, productivity claims remain incomplete.
Redesign roles to reflect verification labour and anticipate further identity shifts. Where AI is materially altering documentation, communication or oversight routines, job expectations and accountability boundaries should reflect that reality. Deloitte’s 2026 survey found 84% of organisations have not done this (Deloitte 2026). Build a two-to-three-year view of how professional roles will continue to evolve as AI penetration deepens.
Replace generic AI training with LLM literacy. Current programmes focus predominantly on prompting skills and AI fundamentals, which address the least safety-critical dimension of AI competence. Effective training must cover: critical evaluation (hallucination detection, trust calibration, error escalation), ethical-legal judgment (consent, privacy, liability, governance) in addition to operational use. This approach aligns with the EU AI Act's Article 4 AI-literacy obligations, which define literacy as competence for informed deployment and awareness of risks, and which enter enforcement from 2 August 2026 (European Commission 2025).
Bring shadow AI into governance. Identify clinicians already using advanced AI models effectively. Create structures for evaluating which external tools and workflows should be adopted organisationally. Treat these practitioners as assets, not compliance risks. The NHS AI scribing guidance provides a directly usable template (NHS England 2026).
Treat governance as implementation infrastructure. Require review-before-sign protocols for AI-assisted outputs. Build tool approval logic, monitoring requirements and auditability into procurement from day one. Test whether governance is translated into frontline practice in at least one workflow. Beware of bloated governance load.
The Questions Worth Answering
Healthcare leaders do not need more assurance that AI is transforming the sector. They need a clearer operating picture of what is already happening beneath adoption dashboards.
AI is not failing in healthcare. But it is being measured with the wrong instruments. When organisations track adoption rates and gross time savings without accounting for verification labour, they are mistaking apparent efficiency for realised value. When they overlook the professional role shift that verification burden creates, they are storing up clinician fatigue, resistance and retention risk. And when they govern the AI tools they provide while ignoring the superior tools their clinicians already use, they are securing the wrong perimeter.
The current evidence base gives leadership enough ground to act. Direct measurement of verification burden as a distinct operating cost in clinical settings has not yet been published, but the convergence of workflow evidence, error-rate data and clinician experience strongly supports its existence and significance. More relevant questions are: (1) How has the role changed for the professionals because AI is in the workflow? and (2) Is anyone measuring the human cost of AI?
This brief draws on randomised trials, large multi-site observational studies, regulatory texts, industry surveys and qualitative clinician research. Where evidence is cross-sector rather than healthcare-specific, this is noted in text. The concept of verification burden as a measurable operating cost represents the author’s theoretical framework, bridging Information Systems and economics; direct empirical measurement in clinical settings is the subject of ongoing research.
Conflict of interest
None.
References:
Ackerhans S, Wehkamp K, Petzina R et al. (2025) Perceived trust and professional identity threat in AI-based clinical decision support systems: Scenario-based experimental study on AI process design features. JMIR Formative Research, 9:e64266. doi.org/10.2196/64266
Adler-Milstein J, DeMasi O, Soleimani H et al. (2026) Subjective and objective impacts of ambulatory AI scribes. American Journal of Managed Care, 32(1):34–40. https://doi.org/10.37765/ajmc.2026.89869
American Medical Association (2026) 2026 physician survey on augmented intelligence. American Medical Association.
Asgari E, Montaña-Brown N, Dubois M et al. (2025) A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine, 8:274. doi.org/10.1038/s41746-025-01670-7
Deloitte (2026) The state of AI in the enterprise 2026: The untapped edge.
European Commission (2025) AI literacy: Questions and answers. Digital Strategy.
Ji Z, Lee N, Frieske R et al. (2023) Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38.
Jussupow E, Spohrer K & Heinzl A (2022) Identity threats as a reason for resistance to artificial intelligence. JMIR Formative Research, 6(3):e28750.
Kamara S (2026). From gross efficiency to realized productivity: Verification friction in AI-augmented knowledge work. SSRN Electronic Journal. Abstract ID: 6469865. papers.ssrn.com/sol3/papers.cfm?abstract_id=6469865
Kernberg AS, Gold JA & Mohan V (2024) Using ChatGPT-4 to create structured medical notes from audio recordings. Journal of Medical Internet Research, 26:e54419.
Lukac PJ, Turner W, Vangala S et al. (2025) Ambient AI scribes in clinical practice: A randomized trial. NEJM AI, 2(12). doi.org/10.1056/AIoa2501000
Mandal S, Wiesenfeld BM, Szerencsy AC et al. (2025) Utilization of generative AI-drafted responses for managing patient-provider communication. npj Digital Medicine, 8:591. doi.org/10.1038/s41746-025-01972-w
Microsoft (2024) 2024 Work Trend Index: AI at work is here. Now comes the hard part.
NHS England (2026) Using AI-enabled ambient scribing products in health and care settings. NHS Transformation Directorate.
Ramaswamy A, Tyagi A, Hugo H et al. (2026) ChatGPT Health performance in a structured test of triage recommendations. Nature Medicine, published online 23 February 2026. doi.org/10.1038/s41591-026-04297-7
Rotenstein LS, Holmgren AJ, Thombley R et al. (2026) Changes in clinician time expenditure and visit quantity with adoption of AI-powered scribes. JAMA. Advance online publication.
Van Tiem J, Cramer E, Iverson C et al. (2026) Listening to the note: Clinician perspectives on ambient AI scribes in medical documentation. JAMIA, 33(2):255–262.
Wolters Kluwer (2026) Survey finds broad presence of unsanctioned AI tools in hospitals and health systems.
