The Promise—and the Problem—of Ambient AI Scribes
Ambient artificial intelligence is rapidly becoming one of the most talked-about technologies in healthcare.
The promise is compelling: an AI system listens to the conversation between a physician and patient, converts the conversation into a transcript, and then generates a clinical note. In theory, the physician spends less time typing, clicks fewer buttons, and can focus more attention on the patient.
And there is evidence that ambient AI can reduce documentation burden.
But there is another question that deserves far more attention:
What happens when the AI gets the medical record wrong?
Recent research suggests that this is not simply a theoretical concern. Studies of commercial ambient AI scribes have identified hallucinations, omissions, incorrect information, and other clinically significant errors. At the same time, litigation involving ambient AI has raised questions about recording, consent, privacy, and the accuracy of AI-generated documentation.
The issue is not whether AI can make physicians more efficient.
The issue is whether an AI-generated version of a clinical encounter should become part of the permanent medical record.
A 26.3% Error Rate Is Not a Minor Problem
One of the most significant recent studies was published in Mayo Clinic Proceedings: Digital Health in 2025. Researchers evaluated five commercial ambient digital scribe platforms using 14 simulated clinical encounters.
The results deserve attention.
Across the platforms, the mean percentage of errors in generated clinical notes was 26.3%. Only 35.8% of correctly reported elements were consistently correct across all platforms.
Even more concerning, researchers identified an average of 3.0 errors per case with potential for moderate-to-severe harm, with individual cases ranging from zero to 21 potentially harmful errors.
The authors concluded that the platforms showed important variability in accuracy and quality and called for standardized, objective evaluation and reporting.
This deserves careful consideration.
A 26.3% error rate in an email or meeting summary might be inconvenient.
A 26.3% error rate in a medical record is a very different proposition.
A medical record can influence diagnoses, prescriptions, referrals, future treatment decisions, insurance determinations, quality reporting, and subsequent physicians’ understanding of the patient’s history.
An error can therefore continue to influence care long after the original encounter.
The Most Dangerous Error May Be Something That Is Missing
When people think about AI hallucinations, they often imagine an AI inventing something obviously absurd.
That is certainly a concern.
But healthcare documentation presents another problem:
Omission.
A missing medication, allergy, symptom, examination finding, clinical observation, or element of medical decision-making may not look obviously wrong when a physician quickly reviews a generated note.
The 2025 Mayo study specifically categorized errors as omissions, commissions, or partially correct information, demonstrating that the problem extends well beyond simple speech-to-text transcription mistakes.
This creates an uncomfortable paradox.
The more fluent and professional an AI-generated note looks, the easier it may be to assume that it is complete.
But polished language does not necessarily mean clinical accuracy.
Real-World Notes Are Not Immune
Laboratory simulations are important, but the next question is what happens in actual clinical practice.
A 2026 study published in JMIR Medical Informatics examined 7,545 AI-generated clinical notes produced during a real-world pilot involving 31 physicians.
Physicians conducted detailed quality assessments on a sample of 356 notes.
Among those assessed notes, accidental omissions were identified in 18%, hallucinations in 11.5%, and accidental inclusions in 9.3%. The researchers also recorded errors that physicians considered potentially serious or capable of creating imminent risk of harm if they were not corrected.
This is particularly important because these were not simply laboratory transcripts.
They were clinical notes produced during actual healthcare delivery.
The findings reinforce a fundamental principle:
AI-generated documentation is a draft that requires clinical verification—not an unquestionable record of what happened.
Physicians Don’t Automatically Trust AI-Generated Notes
Another 2026 study, published in JMIR Formative Research, surveyed emergency and pediatric emergency physicians using an AI-powered ambient scribe across four emergency departments.
Only 42.9% of respondents said they believed they could trust the ambient scribe’s documentation to be accurate, compared with 75% among respondents with previous experience using in-person scribes.
The study also found that only 23.1% of respondents considered the ambient scribe helpful for documenting the physical examination, while 35.7% considered it helpful for medical decision-making documentation.
Importantly, the study was small—only 14 physicians completed the survey—so its results should not be generalized to all physicians or all ambient-scribe systems.
But it provides an important real-world signal:
Physician trust in AI-generated documentation is not automatic, particularly for clinically important portions of the medical record.
A Transcription Can Be Accurate and Still Be Clinically Unsafe
There is another emerging concern.
Ambient clinical scribes typically combine automatic speech recognition with large language models. That means the system does not merely transcribe speech. It interprets the conversation and transforms it into clinical documentation.
A June 2026 preprint, Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes, examined what happened when background noise was introduced into otherwise identical clinical dialogues.
The researchers found that stationary ambient noise increased conventional Word Error Rate by only 0.71 percentage points, yet nearly doubled the rate of unsafe outputs. The authors described a disconnect between conventional speech-recognition accuracy and clinical safety.
This is a particularly important finding.
It suggests that measuring whether an AI system correctly recognizes words may not be enough.
A system can have relatively good transcription metrics while still making a clinically consequential interpretation.
In other words:
A transcription can be linguistically accurate while still being clinically unsafe.
Because this particular study is currently a preprint rather than a peer-reviewed journal publication, it should be viewed as emerging evidence rather than established clinical consensus.
The Legal Risk Is No Longer Entirely Hypothetical
The scientific questions are being accompanied by another development:
litigation.
A proposed class action involving Sharp HealthCare has raised allegations concerning the use of ambient AI technology to record patient encounters without appropriate consent.
The allegations reportedly include claims that AI-generated medical records stated that patients had been informed about recording and had consented—even though the plaintiff alleges that such conversations did not occur.
The case reportedly concerns potentially more than 100,000 patient encounters.
These are allegations in pending litigation, not findings that have been proven in court.
But the case illustrates an important category of risk.
The problem is no longer simply:
“What if the AI makes a mistake?”
It becomes:
“What if the AI creates a statement in the medical record saying that something happened when it did not?”
That is a fundamentally different risk.
More Litigation Over Ambient Recording
Other litigation has also raised questions about the use of ambient AI technology by healthcare organizations.
Patients have filed a proposed class action involving Sutter Health and Memorial Care alleging that Abridge’s ambient AI technology recorded and transcribed confidential physician-patient conversations without adequate consent.
The allegations involve issues including recording, transmission of audio to third-party systems, disclosure, consent, privacy, and state wiretap laws.
Again, these are allegations, not adjudicated findings.
But the larger point is significant:
Healthcare organizations deploying ambient AI are beginning to face legal challenges associated with the technology’s operation.
Privacy and consent are no longer abstract concerns that belong only in a technology white paper.
The Malpractice Question
There is an even more important question for physicians:
Who is responsible when an AI-generated error contributes to patient harm?
The physician ultimately reviews and signs the medical record.
That creates an uncomfortable situation.
An AI system may generate the note.
The physician may have only seconds or minutes to review it.
The physician may trust the system because it has performed correctly many times before.
And yet, if an important error remains in the signed record, the physician—not the algorithm—may be the person expected to explain the clinical decision that followed.
This does not mean that every physician using an ambient scribe is automatically exposed to malpractice liability.
It means that AI-generated documentation introduces another potential source of error that physicians and healthcare organizations need to understand and manage.
The legal question is therefore not simply whether an AI vendor says its system is accurate.
It is whether the healthcare organization has appropriate processes for consent, oversight, review, correction, auditability, data handling, and accountability.
The Problem With Calling It “Just Transcription”
One of the most important distinctions is that modern ambient AI scribes are not simply recording conversations.
The process is more like:
Conversation > speech recognition > transcription > AI interpretation > summarization > generated clinical note
At each step, information can potentially be lost, misunderstood, transformed, or generated.
A system can therefore produce a note that sounds completely coherent while still being clinically inaccurate.
That leads to an important principle:
A clinically fluent note is not necessarily a clinically accurate note.
Traditional transcription errors are relatively easy to conceptualize.
Generative AI introduces a different problem: the system can produce information that was never actually stated.
Ambient AI Can Still Be Useful
It is important not to overstate the evidence.
The research does not demonstrate that all ambient AI scribes are inherently unsafe.
In fact, the 2026 emergency-physician study found that many physicians were satisfied with the technology, and a majority reported that it improved documentation efficiency or reduced documentation burden.
That is why the debate should not be framed as:
AI versus no AI.
The better question is:
What role should AI play in clinical documentation?
AI can absolutely help physicians.
But helping a physician document medicine is different from allowing AI to independently reconstruct what happened during a medical encounter and then making that reconstruction the foundation of the medical record.
A Different Model of Clinical AI
This distinction is particularly important as healthcare moves beyond the first generation of AI scribes.
One model is:
Patient + physician conversation ( ambient recording ( transcription ( generative AI ( AI-written note ( physician review
Another model is:
Physician’s clinical knowledge + patient history + physician-generated concepts ( AI assists the physician’s documentation process
The difference is subtle but fundamental.
In the first model, AI is effectively attempting to determine what happened and communicate that determination through a generated narrative.
In the second, AI assists the physician while keeping the physician’s own clinical knowledge and judgment at the center of the documentation process.
The goal should not be to remove the physician from documentation.
The goal should be to remove unnecessary work without removing physician control.
The Future of AI in Medicine Should Be Physician-Centered
The lesson from the emerging evidence is not that healthcare should reject artificial intelligence.
Healthcare should demand better AI.
AI should reduce documentation burden without introducing unacceptable clinical risk.
It should help physicians work faster without requiring them to blindly trust generated narratives.
It should incorporate longitudinal patient information rather than treating every encounter as an isolated conversation.
And most importantly, it should recognize that the medical record is not simply a summary of a conversation.
It is a clinical and legal record of what the physician knows, observes, evaluates, and decides.
That distinction matters.
The future of medical AI should therefore not be about replacing physician-generated clinical knowledge with AI-generated prose.
It should be about building intelligent systems that amplify the physician’s expertise while keeping the physician firmly in control.
AI should make the doctor more powerful—not make the doctor responsible for whatever the AI happened to write.
Sources
- Anderson TN, et al. (Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters.( Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. DOI: 10.1016/j.mcpdig.2025.100292.
- (Error Frequency and Severity( in the 2026 JMIR Medical Informatics prospective ambient-scribe evaluation, involving 7,545 notes and 31 physicians.
- Marquis T, Kopp M, Anderson JS, Napoli AM, Brown LL, Berlyand Y. (AI-Powered Ambient Scribe Technology Experiences Among Emergency Physicians: Cross-Sectional, Mixed Methods Pilot Survey Study.( JMIR Formative Research. 2026;10. DOI: 10.2196/80401.
- Jiang X-H, et al. (Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes.( arXiv preprint, June 4, 2026.
Important: Legal matters referenced in this article involve allegations and should not be presented as established findings unless and until a court makes such findings. Likewise, individual studies have limitations, including small samples and simulated encounters in some cases. The evidence supports caution and appropriate physician oversight; it does not establish that every ambient AI system is inherently unsafe.


