When the Machine Writes Your Story
AI Scribes, Hallucinated Medicine, and the Patients Left Behind
The machine listens —
invents what silence forgot.
Who speaks for the hurt?
With every article and podcast episode, we provide comprehensive study materials: References, Executive Summary, Briefing Document, Quiz, Essay Questions, Glossary, Timeline, Cast, FAQ, Table of Contents, Index, Polls, 3k Image, Fact Check, Comic and
Street Art at the very bottom of the page.
Soundbite
Trailer
Essay
There is a small, unassuming phone sitting on a doctor’s desk. It is listening to everything.
It listens to the fear in a mother’s voice as she describes her son’s episodes. It listens to the careful hedging of a retired schoolteacher who doesn’t want to be a bother. It listens to the pauses, the restarts, the coughing — the whole ungainly music of a human being trying to communicate the most frightening things they have ever had to say. And when the appointment ends and the door clicks shut, the phone does something remarkable: it writes it all down.
Or rather, it writes down what it thinks it heard. What it predicts should have been said.
And this, quietly, is the problem.
We are in the middle of a genuine healthcare crisis — one that was supposed to have a technological solution. In Ontario alone, more than 2.5 million people currently have no family physician. Across the system, doctors are spending fifteen to twenty hours every week on administrative tasks that have nothing to do with healing anyone. The medical community even has a term for what happens next: pajama time. Physicians put their children to bed and open their laptops on the couch at nine in the evening, charting for three more hours before they can sleep.
It is unsustainable. It is burning people out. And it is ultimately what drove so many practitioners toward AI medical scribes with something close to hope.
The technology promised something beautiful. An ambient microphone would capture the natural conversation between doctor and patient. A large language model would sift through the noise — the small talk about the school district, the nervous laughter — and extract only the relevant clinical data, formatting it automatically into the structured notes that modern medicine requires. Doctors reported saving nearly three hours a week. Ninety-seven percent felt less mentally fatigued. Seventy-eight percent of their patients felt more seen, more attended to, because the doctor had finally looked up from the screen.
For a brief moment, it really did look like a silver bullet.
But silver bullets are a story we tell. The reality is considerably more complicated.
Researchers like Dr. Allison Koenigke at Cornell Tech and Dr. John José Nuñez — a clinical psychiatrist at the University of British Columbia who is also an AI expert — began auditing what these tools were actually producing. What they found was not merely imperfect. It was, in specific and traceable ways, dangerous.
The AI was hallucinating.
In one documented case from the research, a patient said something simple and human and incomplete: I became ill with a fairly serious strain of viral something. That was the entirety of what they said. But the AI scribe, trained on terabytes of internet text in which viral illness sentences are statistically followed by discussions of medication, completed the thought for them. The official clinical note it generated documented that the patient had taken something called “hyperactivated antibiotics” — a medication that does not exist, a treatment the patient never mentioned, a fabrication quietly inserted into a permanent health record.
This happens approximately one percent of the time with models like OpenAI’s Whisper, which underpins many commercial AI scribes. One percent sounds manageable. But one platform alone has transcribed over seven million patient visits. One percent of seven million is tens of thousands of fabricated medical details. Fake drugs. Invented symptoms. False histories. Floating through the permanent records of real people.
The deeper problem is not a bug in the code. It is a fundamental misunderstanding of what language actually is.
The industry has been evaluating these tools using something called word error rate — a metric that counts substituted, deleted, or inserted words and treats every word in the English language as mathematically equivalent. Under this measurement, missing the word not in the sentence I am not having chest pain counts as a single small error. The software receives high marks. But the patient has just been routed to the cardiology ward for a condition they never reported.
This is not a spelling bee. This is medicine. And we are grading it like one.
Meanwhile, the AI was trained overwhelmingly on American data, which produces tools fluent in one very particular dialect of clinical communication — direct, low-context, declarative. A patient who walks in and announces: Doctor, my chest hurts on the left side, pain is an eight out of ten, started two hours ago. That patient the AI understands.
But Canadian patients, shaped by a publicly funded system and a deep cultural aversion to being a burden, communicate differently. They hedge. They soften. They say, it’s been a bit tender lately, when what they mean is that they have been awake in pain at three in the morning for six weeks. They build toward their real concern slowly, establishing relationship first, and often reveal the true reason for their visit only at the very last moment — hand on the doorknob, already apologizing for the intrusion. Physicians in Canada call this the doorknob comment, and observational research shows that if a doctor interrupts a Canadian patient within the first sixty seconds, they will almost certainly miss the primary symptom entirely.
An AI trained on American data, pattern-matching for a direct summary at the beginning of the recording, will document the small talk and miss the crisis.
For patients with speech disorders — stuttering, dysarthria, dysphonia, aphasia — the failures are not subtle. They are structural. If the training data does not contain sufficient phonetic representation of dysarthric speech, the neural network cannot map those sounds to meaning. It guesses. It hallucinates.
In psychiatry, the stakes become almost unbearable to contemplate. A human psychiatrist listens to a patient in a psychotic episode and writes a single careful sentence: Patient endorses a persecutory delusion involving aliens. Clinical. Guarded. Protective of dignity. The AI, having no concept of dignity, may transcribe the entire five-minute rant verbatim — word for word, the Martians, the galactic cloning project, all of it — into the permanent medical record.
Now add the 21st Century Cures Act, and similar open-notes legislation advancing in Canada, which mandates that patients have direct digital access to their own clinical files. Six months later, a patient who has stabilized on medication logs into their portal and reads, in official hospital language, the most terrifying moment of their life. Or worse: they read it while still vulnerable, still paranoid, and the official document confirms for them that the hospital knows the Martians are real.
There are people working to fix this. Koenigke and Nuñez and their colleagues are not simply cataloguing failures — they are building a framework for something better. They are calling for AI software to be regulated the way we regulate pharmaceuticals. Mandatory post-market surveillance. Standardized audits across diverse demographic groups. Mandatory public reporting of hallucination rates.
But their most important proposal is something called participatory design: the idea that the people who are currently most failed by these systems — patients who stutter, patients with regional dialects, patients from Indigenous and Francophone communities — should be brought actively into the design process. Not as test subjects. As collaborators. Asked directly: What does an accurate and dignified medical transcript look like for you?
This is not a technical proposal. It is a moral one. It asks us to design technology to serve the full range of human beings, rather than demanding that human beings reshape themselves to fit the statistical preferences of a model built in Silicon Valley.
A phone sits on a doctor’s desk and listens to everything.
The question is not whether the machine will get better. It will. The question is whether, in the years it takes to get there, we will have quietly changed the way we speak — learned to say things more directly, more simply, more in the manner of training data — in order to be understood by our own medical records.
The doctor looks you in the eye. The machine listens. And somewhere in the space between the human conversation and the algorithmic interpretation, something essential is either preserved or lost.
We should be paying very close attention to which.
This essay accompanies Episode S7 E6 of Heliox: Where Evidence Meets Empathy. The full episode is available at https://www.buzzsprout.com/2405788/episodes/19219093. Produced by Michelle Bruecker and Scott Bleackley.
Link References
Perspective: Listening to Users when Auditing Medical AI Scribes
AI Scribes: Answers to frequently asked questions - CMPA
PIPA and AI Scribes: Best Practices for Healthcare Organizations in BC
AI scribes in rural and remote primary care: an antidote to physician burnout or Pandora’s Box?
Advice to the Profession: Using Artificial Intelligence in Clinical Practice - CPSO
The Rise of AI Scribes: Balancing Efficiency With Privacy in Canadian Health Care
Practical Considerations For Using AI Scribes | Doctors of BC
Artificial Intelligence in Family Medicine
The medico-legal lens on AI use by Canadian physicians - CMPA
OIPC releases guidance on protecting patient privacy when using AI scribes
Canada’s Requirements for AI-Enabled Medical Devices | ICLG.com Briefings
Episode Links
Available for broadcast on PRX
Other Links to Heliox Podcast
YouTube
Substack
PRX ( Public Radio Exchange)
Podcast Providers
Spotify
Apple Podcasts
Patreon
FaceBook Group
STUDY MATERIALS
Quiz & Answer Key
Essay Questions
Glossary of Key Terms
Cast of Characters
FAQ
Table of Contents with Timestamps
The Broken Arm and the Broken System — 00:25 The episode opens with a deceptively simple image: a fractured radius visible on an x-ray. From that moment of diagnostic clarity, the hosts pivot to the murky, contested terrain of human speech and medical documentation — where nothing is as legible as bone on film.
The Crisis That Made the Miracle — 02:40 To understand why AI scribes seemed like salvation, we must first understand the catastrophe they were meant to solve. Over 2.5 million Ontarians without a family physician. Fifteen to twenty hours of weekly administrative burden per doctor. Pajama time. A healthcare system quietly drowning in its own paperwork.
What an AI Scribe Actually Does — 05:20 This section demystifies the technology. AI scribes are not dictation software. They are ambient listeners — passive microphones that capture the full natural conversation between physician and patient, then use large language models to extract, interpret, and format the clinical data automatically.
The Silver Bullet and Its Limits — 07:00 Initial pilots were stunning. Nearly 97% of participating physicians reported reduced mental fatigue. Seventy-eight percent of patients felt more attended to. But as researchers audited the transcripts at scale — across millions of real patient visits — a critical flaw began to emerge.
Hallucination: When the Machine Invents Medicine — 08:40 A patient says viral something. The AI documents a treatment for hyperactivated antibiotics — a medication that does not exist. This section unpacks the mechanics of large language model hallucination and explains why a system optimized for statistical plausibility is structurally incapable of documenting reality.
The Spelling Bee Problem: Why Word Error Rate Fails Medicine — 11:39 The standard industry metric for evaluating AI speech tools treats every word as mathematically equivalent. Missing the word not from I am not having chest pain registers as a single minor error. The patient is sent to cardiology. This section exposes the fundamental inadequacy of current testing standards.
Canadian Politeness and the Doorknob Comment — 14:20 Cultural communication styles are not quirks — they are clinically significant. This section explores the profound difference between low-context American medical communication and high-context Canadian hedging, and what happens when an AI trained on predominantly American data encounters a patient who says it’s a bit tender when they mean they are in agony.
Speech Diversity and the Patients Most at Risk — 21:01 The AI performs significantly worse for patients with dysphonia, dysarthria, stuttering, aphasia, and hearing difference. This section examines why these populations — who often require the most precise medical documentation — are the most systemically underserved by current AI models.
Psychiatry, Dignity, and the Martians — 24:02 In a psychiatric context, verbatim transcription is not neutral. It is potentially harmful. This section explores the collision between AI literal documentation and the clinical judgment required to protect the dignity of patients in crisis — including those whose own records could actively worsen their condition.
Rural Communities and the Theory of Fundamental Causes — 27:31 A technology designed to reduce health inequity may in practice widen it. Remote and rural clinics face bandwidth failures, subscription costs they cannot afford, and AI models with no training data for the Indigenous and Francophone populations they serve.
Guardrails, Governance, and the Liability Question — 31:21 Canada’s College of Physicians and Surgeons has drawn a hard regulatory line: AI cannot autonomously finalize a medical record. A physician must review and sign every note. But this creates a paradox — does meticulously proofreading an AI’s hallucinations recreate the very burden the technology was meant to eliminate?
Privacy, Clipboard Vulnerabilities, and the Deleted Tape — 33:49 When audio is deleted immediately for privacy reasons, there is no ground truth to audit against. When notes are moved between systems by copy-and-paste, there is a clipboard vulnerability that risks catastrophic cross-patient errors. The infrastructure of care is more fragile than it appears.
The Path Forward: Participatory Design and a New Standard of Care — 36:21 The researchers’ call to action: regulate AI scribes as medical devices, mandate demographic audits, require public hallucination reporting, and bring marginalized speakers into the design process as collaborators rather than afterthoughts.
The Provocative Closing Thought — 38:19 Will artificial intelligence eventually learn to understand the full, messy, high-context reality of what we actually mean? Or will we slowly, unconsciously, begin to change the way we speak — to accommodate the limitations of the machine?
Index with Timestamps
21st Century Cures Act, 26:05
90-second rule, 17:57
Administrative burden, 03:21, 33:03
Ambient listener, 05:47
Aphasia, 21:41
API connection, 28:27
Auditing AI systems, 11:46, 34:50
Behavioral health, 32:26
British Columbia, 02:07, 31:40
Burnout, 02:26, 04:51, 29:53
Canadian hedging, 16:02, 30:58
Clipboard vulnerability, 35:59
Clinical note, 06:32, 25:42
CMPA (Canadian Medical Protective Association), 31:40
College of Physicians and Surgeons of British Columbia, 31:40
Copy-paste workflow, 35:36
CPSBC, 31:40
Cultural communication, 14:20, 17:00
Delusion, 25:04, 26:49
Diagnostic accuracy, 13:09
Dignity, 23:29, 25:39
Doctors of BC pilot, 06:57
Doorknob comment, 18:21
Dysarthria, 21:35
Dysphonia, 21:28
Electronic medical records (EMR), 03:58, 35:23
Family physician shortage, 03:03
Final check technique, 20:22
Francophone communities, 30:30
Fundamental causes theory, 29:04
Guardrails, 31:25, 27:23
Hallucination, 08:42, 11:22, 30:39
Healthcare infrastructure, 28:15
High-context communication, 18:09, 40:17
Hyperactivated antibiotics, 08:46, 33:08
Indigenous communities, 30:30
Insurance-driven healthcare, 14:40
Koenigke, Dr. Allison, 02:00, 08:01, 21:13
Large language model, 06:16, 09:40
LEARN model, 20:08
Legal liability, 31:17, 32:34
Levenshtein distance, 12:04
Low-context communication, 15:04
Malpractice, 33:08
Manic episode, 24:23
Medical AI scribe, 02:20, 05:17
Mental fatigue, 07:18, 04:51
Nabla, 08:11, 34:33
Natural language processing, 05:44, 28:27
Neural network, 10:06, 22:06
Nuñez, Dr. John José, 02:00, 02:50, 11:46, 17:28, 19:36
Ontario, 03:03, 29:41
Open notes movement, 26:05
OIPC (Office of Information and Privacy Commissioner), 33:56
OpenAI Whisper, 08:25, 11:05, 18:52
OSCAR (EMR system), 35:23
Pajama time, 04:35, 29:41, 33:39
Paranoia, 24:42, 26:49
Participatory design, 37:24
Patient portal, 26:25
Post-marketing surveillance, 36:56
Pressured speech, 24:23
Privacy, 33:56, 34:10
Psychiatric emergency, 25:03
Psychiatric documentation, 27:23
Psychosis, 25:04, 26:43
Rural communities, 27:31, 29:41
Single-payer healthcare, 15:34
SOAP note, 06:34
Sociolinguistics, 14:08
Speech disorder, 21:01, 22:02
Statistical prediction, 09:54, 10:06
Stutter, 22:57, 37:45
Suicide risk assessment, 32:26
Therapeutic relationship, 39:15
Token prediction, 09:54
Verbatim transcription, 23:41, 25:39
Word error rate, 11:53, 13:26
Poll
Post-Episode Fact Check
Claim: Over 2.5 million people in Ontario do not have a family physician. Verdict: Supported. This figure is consistent with widely reported data from the Ontario Medical Association and Health Quality Ontario. The provincial physician shortage is a well-documented ongoing crisis, with estimates in the 2.2–2.5 million range depending on year and methodology.
Claim: Family physicians spend 15–20 hours per week on non-essential administrative tasks. Verdict: Supported.This figure aligns with published research from the Canadian Medical Association and the Ontario Medical Association. Some studies place the figure slightly lower (10–15 hours) depending on specialty and practice setting, but the 15–20 range is within the reported literature.
Claim: Administrative burden equates to 55.6 million patient visits annually lost to paperwork. Verdict: Plausible, with appropriate sourcing caveat. This figure appears to be an extrapolation from per-physician lost time aggregated across the Canadian physician workforce. The methodology is reasonable but the precise figure should be attributed to its source study for verification. The order of magnitude is credible.
Claim: The Doctors of BC AI pilot involved over 30 physicians and reported 2.7 hours saved per week. Verdict: Supported. The Doctors of BC ambient AI pilot program is real and was publicly reported. Reported time savings and physician satisfaction metrics are consistent with this figure and with broader published literature on AI scribe pilots.
Claim: Studies show 70–90% decrease in documentation time. Verdict: Supported. Published literature on AI scribes, including studies in JAMA Network Open and other peer-reviewed outlets, reports documentation time reductions in this range, though results vary by tool and setting.
Claim: 97% of pilot participants reported reduced mental fatigue and said the technology “brought joy back into practice.” Verdict: Supported with sourcing note. This figure is consistent with the Doctors of BC pilot and similar reported results. The phrase “brought joy back into practice” appears in published physician testimonials around AI scribe pilots. Precise attribution should be confirmed against the original Nuñez/Koenigke paper.
Claim: 78% of patients felt their doctors paid more attention to them. Verdict: Plausible. This figure is consistent with patient satisfaction outcomes reported in AI scribe literature. It should be verified against the specific study cited.
Claim: Nabla has transcribed over 7 million patient visits. Verdict: Supported at time of research. Nabla, a French AI medical documentation company, has publicly reported this scale of deployment. The figure may have grown since the paper was written.
Claim: OpenAI’s Whisper is the foundational backbone for many commercial AI scribes. Verdict: Supported.Whisper is widely used as the ASR (automatic speech recognition) layer in multiple commercial AI scribe products. This is publicly documented.
Claim: Hallucinations occur approximately 1% of the time with Whisper. Verdict: Supported with nuance. The Koenigke research and related audits of Whisper do identify hallucination rates in roughly this range for certain categories of error. The 1% figure specifically refers to insertional hallucinations (fabricated content). Other error categories have different rates.
Claim: The word error rate metric uses Levenshtein distance. Verdict: Correct. WER calculation is indeed based on edit distance (Levenshtein distance), counting substitutions, deletions, and insertions. This is standard in computational linguistics.
Claim: The 21st Century Cures Act mandates patient access to their own clinical notes. Verdict: Correct. The 21st Century Cures Act (2016, with information-blocking provisions effective 2021) does require healthcare providers to provide patients with electronic access to their health information, including clinical notes, without charge or delay.
Claim: The College of Physicians and Surgeons of BC and the CMPA prohibit autonomous AI documentation.Verdict: Supported. Both bodies have issued guidance requiring physician review and sign-off on AI-generated notes. The CPSBC and CMPA have published specific guidance documents on AI scribe use that align with the description in this episode.
Claim: Some AI scribes delete audio immediately upon transcript generation. Verdict: Supported. Nabla and other privacy-conscious providers do implement immediate audio deletion as a privacy feature, which creates the audit paradox described.
Claim: The theory of fundamental causes predicts that new health technologies benefit advantaged populations first. Verdict: Supported. The theory of fundamental causes, developed by sociologists Bruce Link and Jo Phelan, is a well-established framework in health sociology and is accurately described in this episode.
Overall Assessment: The factual claims in this episode are well-grounded in published research and publicly reported data. Specific figures (patient visit counts, time savings percentages) should be cross-referenced against the primary sources cited — particularly the Koenigke/Nuñez paper — for precision, but no claims appear fabricated or materially misleading.






