
A patient in England was told they had demyelination, the nerve damage that underlies conditions including multiple sclerosis. The test result had actually read “null demyelination,” and an AI scribe had dropped the word that reversed the meaning.
That case appears in a warning from Healthwatch England, the statutory patient watchdog, reported by the Guardian today. Its finding is that the tools now transcribing consultations for GPs and hospital doctors are getting drug names and diagnoses wrong, and that patients are often the ones who notice.
Twenty-seven different AI scribes are in use across the health service in England. They sit in the consulting room, listen, and produce the note that goes into the record and the letter that goes to the patient.
The errors collected by the watchdog are mundane in form and serious in effect. One scribe swapped a prescribed drug for a different one with a similar name, the kind of confusion that pharmacology spends considerable effort trying to design out.
Another summary omitted a consultant’s instruction that the patient should seek a repeat prescription for migraine medication. A third recorded a doctor telling a patient to continue their Prozac, when that doctor had neither prescribed it nor discussed it.
The common thread is that nothing looked broken. A fluent, plausible note is exactly what these systems produce, which is why an error survives the glance a busy clinician gives it before signing off.
“These inaccuracies may persist in their records if the patient doesn’t catch them,” the watchdog warned. That places the last line of defence on the person least equipped to know what the note should have said.
Rachel Power, chief executive of the Patients Association, is among those raising concerns, alongside clinicians including the London GP Shier Ziser Dawood and Charlotte Blease of Uppsala University in Sweden. The objection is not to the technology so much as to its arrival without a safety net.
Because there is no England-wide oversight of these tools. The Medicines and Healthcare products Regulatory Agency has not classified AI scribes as medical devices, which leaves them outside the regime that would test them for safety and effectiveness before deployment.
The regulator did publish guidance in August clarifying where the line sits. A system that only transcribes what was said is not a device, while one that suggests a diagnosis or a treatment may well be, which puts a great deal of weight on how each product is described by its vendor.
The incentive that follows is obvious enough. A scribe marketed as a passive transcriber avoids a regulatory process that a scribe marketed as a clinical assistant would have to complete.
Transcription error is also a different failure from the one most AI safety work anticipates. Nobody here was misled by a hallucinated fact; a real sentence was rendered slightly wrong, and slightly wrong is sufficient when the sentence names a drug.
None of which addresses why these tools spread so quickly. Clinical documentation is the administrative burden doctors complain about most, and a system that reliably removes an hour of typing a day will be adopted whether or not anyone has assessed it.
The problem is what happens in between those two facts. A tool adopted for its speed, unassessed because of how it is categorised, producing a document that becomes the permanent clinical record, is a chain in which no single link is obviously anyone’s responsibility.
The remedy the watchdog points to is unglamorous and probably right. Patients should be told when a scribe is being used and given their notes to check, which turns an accidental safety mechanism into a deliberate one.
Twenty-seven products, no device classification, and a check performed by whoever happens to read their letter carefully: that is the current arrangement in the NHS in England.
How AI scribes work and why they are attractive
AI scribes are ambient documentation tools that use automatic speech recognition and natural language processing to convert the dialogue between a clinician and a patient into structured clinical notes. They are designed to relieve doctors of the burden of typing notes during a consultation, allowing them to focus on the patient. The technology has been promoted as a way to reduce burnout, improve the patient-clinician interaction, and increase efficiency. In the UK, the adoption of such tools has been rapid, especially after the COVID-19 pandemic accelerated the digital transformation of healthcare services.
Many AI scribes are built on large language models that can summarise long conversations into a concise summary, capture key clinical details, and even generate patient-facing letters. Some products are integrated directly into electronic health record systems, while others stand alone and produce text that clinicians must copy and paste. The vendors behind these tools often claim high levels of accuracy, but the evidence base for their safety and effectiveness remains thin, and there have been no robust independent evaluations published in peer-reviewed journals.
The appeal to clinicians is tangible. In the NHS, doctors spend an estimated one-third of their time on documentation. A tool that can halve that time could free up valuable hours for patient care. Furthermore, many GPs and hospital consultants report that administrative workload is a major driver of stress and early retirement. The promise of AI scribes is not just about convenience; it is about retaining clinicians in a workforce that is increasingly stretched.
However, the speed of adoption has outpaced the development of appropriate safeguards. The Department of Health and Social Care has not mandated a specific framework for AI scribes, and individual NHS trusts are free to choose which products to deploy. This has led to a fragmented landscape in which different trusts use different tools, with varying levels of training for staff and often no formal evaluation of the risk. As a result, the safety of these systems is being tested in real time, in live consultations, with patients as the eventual beneficiaries or victims of the outcomes.
Consequences of transcription errors in clinical settings
Medical transcription errors can have profound consequences. The case of the missing “null” in “null demyelination” is a stark example: the patient was presumably much more alarmed than they needed to be, and the erroneous note could have led to unnecessary follow-up tests, referrals, or even treatment. In other cases, transcription errors can lead to the wrong medication being prescribed, an allergy being missed, or a critical instruction being omitted. For instance, if an AI scribe fails to note that a patient is to stop a particular drug, the patient might continue taking it indefinitely, potentially causing adverse effects.
Even seemingly minor errors can erode trust between patients and clinicians. When a patient reads a letter that contains an inaccurate description of their condition or the treatment plan, they may become confused or anxious. They might question the competence of their doctor, or worse, they might act on the wrong information. In the NHS, where patients are encouraged to participate in their own care, such errors undermine the principles of shared decision-making and informed consent.
Moreover, the clinical record is a legal document. It is used for continuity of care, medico-legal investigations, and research. An error that goes uncorrected could persist for years and influence future decisions by other healthcare professionals. The Healthwatch England report highlights that patients are often the first to spot these errors, which suggests that clinicians are not adequately checking the output of AI scribes. This raises questions about professional accountability and the delegation of clinical documentation to automated systems.
Regulatory gaps and potential solutions
The current regulatory status of AI scribes is ambiguous. The MHRA’s guidance from August was intended to clarify the boundary between a simple transcription tool and a medical device that requires rigorous scrutiny. However, critics say the guidance is too permissive and creates a loophole that vendors can exploit. By changing the marketing language, a product can be positioned as a “passive transcriber” rather than a “clinical decision support tool,” thus escaping regulation even if it uses the same underlying algorithms and capabilities.
This is not a hypothetical concern. There have been instances where AI systems initially cleared as low-risk have later been found to have harmful biases or errors, but because they were not regulated, there was no mechanism for recalling or updating them. The lack of a central registry for AI scribes means that when an error is detected, it is shared informally among clinicians rather than systematically reported to a national body. This prevents the healthcare system from learning from mistakes and improving safety.
Some experts suggest that AI scribes should be classified as medical devices by default, given that they generate content that feeds into clinical decision-making. Others argue that a lighter-touch approach, such as a voluntary certification scheme or mandatory external audit, could be more appropriate. In any case, there is a clear need for a robust evaluation framework that includes testing on real-world consultations, monitoring for errors, and feedback mechanisms for clinicians and patients.
The General Medical Council and other professional bodies have issued guidance on the use of AI in practice, but these documents are often generic and fail to address the specific challenges of AI scribes. For example, they do not specify who is responsible when an AI scribe makes an error, or what steps clinicians should take to verify the accuracy of the generated notes. The Healthwatch England report recommends that patients be alerted when an AI scribe is in use and be given a copy of the notes to review. This is a simple, low-cost intervention that could mitigate the risk of errors going unnoticed.
The broader context of AI in healthcare
The issues highlighted in this article are not unique to AI scribes. The NHS has been aggressively pursuing the use of artificial intelligence across a range of applications, from diagnostic imaging to predictive analytics to patient triage. While these technologies hold great promise, they also introduce new risks that traditional healthcare governance structures are ill-equipped to handle. The case of AI scribes is a cautionary tale: a technology that is adopted for its immediate benefits, without adequate oversight, can create unintended harms that undermine patient safety.
One of the challenges is that AI systems are often opaque. Clinicians may not understand how a system arrived at a particular output, and they may over-trust the system based on the vendor’s claims. This is known as automation bias, a phenomenon in which humans are more likely to trust a machine’s suggestion than the same suggestion from a human. In the case of AI scribes, the notes generated by the system may appear authoritative and complete, leading clinicians to sign off without careful reading.
Another issue is the potential for bias in natural language processing models. These models are trained on large datasets, which may not be representative of the diverse patient population that the NHS serves. For example, the AI scribe might have more difficulty understanding accents or dialects, leading to higher error rates for patients from certain regions or ethnic backgrounds. This could exacerbate health inequalities, a concern that has been raised by digital health researchers but has not been adequately addressed by vendors or regulators.
Despite these concerns, the momentum behind AI scribes is unlikely to slow down. The pressure on the NHS to increase efficiency is intense, and the government has set ambitious targets for the adoption of AI technologies. The key is to find a balance between innovation and safety, so that the potential benefits of AI scribes can be realised without compromising the quality of care. This will require a collaborative effort between regulators, clinicians, patients, and technology developers to establish standards that are both rigorous and practical.
For now, the message from Healthwatch England is clear: caution is needed. The technology is not yet mature enough to be deployed without safeguards, and the responsibility for ensuring safety should not rest solely on patients. Clinicians must be trained to critically evaluate AI-generated notes, and the use of AI scribes should be transparent to patients. As the NHS moves further into the age of artificial intelligence, the lessons learned from AI scribes will be invaluable in shaping the future of healthcare.
