AI Medical Scribes: What the Latest Studies Really Show

Randomized trials and large rollouts of ambient AI scribes report modest time savings and less burnout, plus omissions and hallucinations to watch for.

AI Medical Scribes: What the Latest Studies Really Show

Ambient AI scribes listen to the consultation, with the patient's consent, and produce a draft clinical note for the physician to review. In two years they have gone from pilot projects to large deployments, especially in the United States. Expectations are high: hours of documentation saved and doctors looking at patients instead of screens. The evidence published through 2025 and 2026 gives a more nuanced picture. Benefits are real but smaller than the marketing suggests, and quality control remains the physician's job.

The first randomized trial: real but modest time savings

The strongest evidence so far comes from a pragmatic randomized trial by Lukac and colleagues, published in NEJM AI in November 2025. At UCLA Health, 238 outpatient physicians across 14 specialties were randomized to one of two commercial ambient scribes (Microsoft DAX Copilot or Nabla) or to usual care. The trial ran from November 2024 to January 2025.

  • The primary outcome was time spent writing notes, measured from EHR logs.
  • Nabla users saw a 9.5% reduction in time-in-note compared with the control group, a statistically significant result.
  • DAX users showed no significant change (−1.7%).
  • According to UCLA's summary, the reduction for Nabla users amounted to roughly 41 seconds per note, compared with 18 seconds in the control arm.
  • The authors described the improvements in burnout, task load and work exhaustion as signals worth confirming rather than definitive findings.

Time savings measured in seconds per note can still add up over a full clinic. They are, however, far from claims that AI scribes eliminate documentation.

Burnout: encouraging before-and-after data

A multicenter quality-improvement study by Olson and colleagues, published in JAMA Network Open in 2025, followed 263 physicians and advanced practice clinicians in six US health systems. After 30 days of using an ambient scribe:

  • Self-reported burnout fell from 51.9% to 38.8%.
  • Clinicians reported lower cognitive load related to notes, more focused attention on patients and less after-hours documentation.

The design matters for interpretation. This was a pre-post study without a randomized control group, and the scribe vendor facilitated data collection, although the authors state it had no role in the analysis. Novelty effects and selection of enthusiastic users can inflate early results.

Large-scale deployment: The Permanente Medical Group

The largest real-world experience published so far comes from The Permanente Medical Group (Kaiser Permanente, Northern California), reported in NEJM Catalyst. Between October 2023 and December 2024, 7,260 physicians used ambient AI across about 2.5 million patient encounters. The group estimated about 15,791 hours of documentation time saved, roughly 1,794 eight-hour workdays. In a survey of physicians, most reported a positive experience, citing lower mental workload and better patient interaction.

Spread across thousands of physicians, the figure is meaningful but not dramatic per person. It also comes from an organization with strong IT infrastructure and English-language consultations.

The quality question: omissions and hallucinations

Time savings mean little if the notes are wrong. Several studies have looked at what AI scribes get wrong:

  • A 2025 study in Frontiers in Artificial Intelligence compared ambient AI notes with physician-written notes. AI notes were more thorough and better organized but less concise. Hallucinated content appeared in 31% of AI notes versus 20% of physician notes, according to the authors' assessment.
  • A 2026 pilot published in JMIR Medical Informatics reviewed 356 of 7,545 AI-generated notes from 31 physicians. The most common error was omission (18% of reviewed notes), followed by hallucination (11.5%) and accidental inclusion of irrelevant content (9.3%).

Omissions are particularly insidious. A missing negative finding or a dropped medication change does not look like an error on screen. Reviewers also disagree on how to classify these errors, and older note-quality scales were not designed for language-model output.

What this means outside the US

Most published studies involve English-speaking clinicians, US health systems and cloud-based products. Clinicians in the Middle East and Africa face extra questions:

  • Language. Consultations often mix Arabic dialects with French or English. Performance on such conversations is rarely reported.
  • Data location. Ambient scribes typically send audio of the whole consultation to cloud servers. In countries with health data localization rules, such as the UAE under Federal Law No. 2 of 2019, this needs legal review.
  • Consent. Recording a full consultation requires clear, documented patient consent and a policy for retaining the audio.

Dictation after the visit, where the physician speaks a structured summary, is a lighter alternative. It captures less ambient detail but keeps the physician in control of content, and it can run entirely offline.

Key takeaways

  • The best randomized evidence shows modest time savings (around 10% of note-writing time for one product), not a revolution.
  • Before-and-after studies suggest less burnout, but designs without control groups overstate effects.
  • Omissions are the most frequent AI scribe error. Review every draft against your own memory of the visit.
  • Check the language mix, the data location and the consent process before any deployment in the region.
  • Measure your own baseline and results; do not rely on vendor figures.

Frequently asked questions

Do AI scribes really save time for doctors?

Yes, but modestly. In a 2025 randomized trial, one ambient scribe cut note-writing time by 9.5% versus usual care, while another showed no significant change.

Are AI-generated clinical notes accurate?

Not always. Studies report omissions and hallucinated content in a notable share of AI drafts, so the physician must review and correct every note before signing it.

Do ambient AI scribes reduce burnout?

A 2025 multicenter pre-post study found burnout fell from 51.9% to 38.8% after 30 days, but it had no randomized control group.

Sources

  1. NEJM AI — Ambient AI scribes in clinical practice: a randomized trial (Lukac et al., 2025)
  2. UCLA Health — UCLA study finds AI scribes may reduce documentation time and improve physician well-being
  3. HealthDay — Ambient artificial intelligence scribe linked to reduction in burnout (Olson et al., JAMA Netw Open 2025)
  4. American Medical Association — AI scribes save 15,000 hours and restore the human side of medicine
  5. Frontiers in Artificial Intelligence — Ambient AI vs physician-written notes (2025)
  6. JMIR Medical Informatics — Pilot quality-improvement study of AI-generated notes (2026)
Dictate your reports, nothing leaves your computer

Nabady Whisper transcribes your voice offline in English, French or Arabic, with report templates for every specialty.

General information, checked at the publication date; it is neither medical nor legal advice.

Share LinkedIn WhatsApp X