What does the best evidence show about AI scribes?
The most rigorous study to date is a 2025 randomized trial of 238 physicians across 14 specialties [1]. It found that one AI scribe (Nabla) cut documentation time by 9.5% compared to usual care, while another (DAX Copilot) showed no significant reduction. Both scribes improved burnout scores on the Mini-Z scale (DAX: +2.83 points, Nabla: +2.69 points on a 10-50 scale) and reduced physician task load. However, clinicians reported occasional clinically significant inaccuracies (average score of 2.7-2.8 out of 5 on a frequency scale, where 1 is 'never' and 5 is 'always'). This means the technology can help, but errors still require vigilance.
A separate 2025 retrospective cohort study of 125 AI scribe users [3] found that users spent 8.5% less time in the electronic health record (about 2.4 minutes per appointment) and 15.9% less time on notes (about 1.8 minutes per appointment) compared to non-users. This study used a rigorous propensity-score matching method to reduce bias, strengthening the finding that AI scribes can improve efficiency. However, the effect sizes are modest—saving a few minutes per visit—which may not dramatically change a clinician's day.
What are the main limitations and risks?
Accuracy is the biggest concern. In the randomized trial [1], both scribes scored around 2.7-2.8 on a 5-point scale for 'clinically significant inaccuracies,' meaning errors occurred occasionally. A qualitative study of 22 physicians [2] found that while most felt AI scribes reduced cognitive load (100% of comments positive) and improved work-life integration (91% positive), perspectives on accuracy and note style were 'largely negative.' Physicians complained about overly long notes and the need for heavy editing, which can offset time savings.
The technology also struggles with non-English speakers and complex cases. The same qualitative study [2] noted limited functionality with non-English-speaking patients as a key barrier. A pilot in palliative medicine [5] showed mixed results: one resident saved significant time, while another saw no improvement, and the AI's organization and usefulness only improved for one user over time. This suggests that individual user skill and the clinical context heavily influence outcomes.
Even in a controlled simulation of 25 inpatient surgical scenarios [6], DAX Copilot scored 46.9 out of 50 on accuracy and completeness, but still required basic editing to fit surgical note style. The authors rated it 4 out of 5 for organization and usefulness, noting it is 'poised to revolutionize' documentation but needs more testing in real patient encounters.
So, are AI scribes ready for clinical deployment?
The answer is a qualified 'yes' for some settings, but 'not yet' for widespread, independent use. A 2025 systematic review [4] of eight studies concluded that AI scribes show promise in improving documentation efficiency and clinician workflow, but the evidence is 'limited and heterogeneous.' Most studies had small samples and specific settings, limiting generalizability. The review emphasized that accuracy varies significantly by technology and implementation approach.
The bottom line: AI scribes can reduce documentation time and burnout in outpatient primary care and some specialties, especially when used by motivated clinicians. But they are not a plug-and-play solution. They require ongoing human oversight, work best with English-speaking patients, and may need customization for different note styles. Clinics considering deployment should pilot the technology, train users, and monitor for errors—especially in complex cases. The technology is advancing rapidly, but as of 2025, it is a helpful assistant, not a replacement for the clinician's judgment.
About These Sources
This answer is built on 6 peer-reviewed studies — published in 2025, 6 from 2024 or later, 4 in Q1 journals, collectively cited 95 times — selected as the most relevant from 6 studies that passed quality screening, drawn from 51 papers retrieved from a database of over 500 million.
Sources used in this answer
Ambient AI Scribes in Clinical Practice: A Randomized Trial
In a randomized trial of 238 physicians, Nabla reduced documentation time by 9.5% while DAX showed no significant reduction; both improved burnout scores, but occasional inaccuracies were reported (score ~2.7-2.8/5).
Physician Perspectives on Ambient AI Scribes
In a qualitative study of 22 physicians, AI scribes reduced cognitive load (100% of comments positive) and improved work-life integration (91%), but accuracy and note style were criticized, especially for non-English speakers.
Use of an AI Scribe and Electronic Health Record Efficiency.
In a retrospective cohort of 125 AI scribe users, use was associated with 8.5% less EHR time and 15.9% less note time per appointment compared to non-users, with no significant change in after-hours work.
The Impact of AI Scribes on Streamlining Clinical Documentation: A Systematic Review
A systematic review of 8 studies found AI scribes improve documentation efficiency and clinician workflow, but evidence is limited, heterogeneous, and accuracy varies by technology and setting.
Ambient Artificial Intelligence Scribes: A Pilot Survey of Perspectives on the Utility and Documentation Burden in Palliative Medicine
A pilot in palliative medicine with 2 residents showed mixed results: one saved significant time, the other did not; AI note organization improved for only one user over time.
DAX Copilot: ambient AI scribe may help reduce surgical resident clinical documentation burden
In 25 simulated inpatient scenarios, DAX Copilot scored 46.9/50 on accuracy and completeness but required editing to fit surgical note style; rated 4/5 for organization and usefulness.
