AI scribes are inaccurately recording medication names and diagnoses in NHS records.

AI scribes are inaccurately recording medication names and diagnoses in NHS records.

      A patient in England was informed they had demyelination, a type of nerve damage associated with conditions like multiple sclerosis. However, the test result actually indicated “null demyelination,” and an AI scribe had mistakenly omitted the word that changed its meaning. This incident is highlighted in a warning issued by Healthwatch England, the official patient oversight body, as reported by the Guardian today. The organization's findings reveal that the AI tools being used to transcribe consultations for general practitioners and hospital doctors are misinterpreting drug names and diagnoses, often with patients being the ones who identify these errors.

      There are currently twenty-seven different AI scribes operating within the health service in England. These tools listen during consultations, generating notes for medical records and letters for patients. The errors reported by the watchdog are generally straightforward in nature but can have serious implications. For instance, one scribe replaced a prescribed medication with another that has a similar name, an issue that pharmacology efforts aim to minimize.

      In another case, a summary failed to include a consultant’s advice that the patient should request a repeat prescription for migraine medication. A further instance recorded a doctor instructing a patient to continue taking Prozac, despite the fact that the doctor had neither prescribed it nor discussed it with the patient.

      The common issue is that nothing appeared wrong. The AI systems produce fluent, credible notes, which is why errors often escape notice during a busy clinician's review before approval. “These inaccuracies may remain in the records if the patient does not identify them,” the watchdog cautioned, placing the burden of oversight on individuals who may not be familiar with what the note should accurately reflect.

      Rachel Power, the chief executive of the Patients Association, shares concerns alongside clinicians such as London GP Shier Ziser Dawood and Charlotte Blease from Uppsala University in Sweden. Their objection centers not on the technology itself but on its introduction without adequate safety measures.

      Currently, there is no overarching regulation for these tools in England. The Medicines and Healthcare Products Regulatory Agency has not classified AI scribes as medical devices, meaning they are not subject to the safety and effectiveness testing that would normally precede their use. In August, the regulator released guidance clarifying the distinction: a system that merely transcribes spoken words is not considered a device, while one that suggests diagnoses or treatments likely is, placing significant responsibility on how each vendor describes their product.

      The incentive is clear; a scribe marketed as a passive transcriber avoids the regulatory scrutiny that would apply to one presented as a clinical assistant. Additionally, transcription errors represent a different type of failure than what most AI safety considerations anticipate. Here, nobody was misled by a fabricated fact; rather, a legitimate statement was inaccurately recorded, and even minor inaccuracies can have serious consequences, especially when it involves medication names.

      The speed at which these tools are being adopted remains largely unexplained. Clinical documentation is a major source of frustration for doctors, leading them to embrace technology that can reliably save significant typing time, regardless of prior assessments.

      The crux of the issue lies in the implications of these two realities. A tool embraced for its efficiency—yet untested due to its classification—producing documents that become permanent clinical records creates a situation where accountability is not clear.

      The watchdog suggests a straightforward and likely effective solution: patients should be informed when a scribe is in use and provided with their notes to review, transforming an accidental safety mechanism into a purposeful one. As it stands, there are twenty-seven products, no device classification, and a verification process reliant on whoever happens to carefully read their letter—this is the current state of the NHS in England.

Other articles

Feeling the impact of OpenAI removing GPT models from Cursor? Anthropic provides a timely alternative with increased limits for Claude. Feeling the impact of OpenAI removing GPT models from Cursor? Anthropic provides a timely alternative with increased limits for Claude. Anthropic is offering increased computing power for Claude within Cursor as OpenAI plans to end its model-access agreement with the coding platform. OpenAI has begun allowing certain customers to pay only when the AI performs successfully. OpenAI has begun allowing certain customers to pay only when the AI performs successfully. According to a report by The Information, select large accounts can now pay for completed tasks instead of using tokens, following in the footsteps of Intercom, Zendesk, and Salesforce in adopting outcome-based pricing. According to the FSB, AI-powered cyber attacks have emerged as the primary threat to the financial system. According to the FSB, AI-powered cyber attacks have emerged as the primary threat to the financial system. Andrew Bailey informed G20 finance ministers that AI accelerates the speed, scale, and economics of cyber attacks, elevating it as a more significant stability concern than any other issue. The robot pizza chefs are not performing well. The robot pizza chefs are not performing well. The anticipated future of fast food was meant to resemble this scenario: enter a restaurant, make an order, and observe a robot expertly creating your pizza with mechanical accuracy. However, several companies that vowed to automate pizza preparation have found it challenging to transform that idea into a viable business. The most recent instance is Picnic, whose […] A new oversight organization is monitoring incidents of AI behaving improperly, and the number of occurrences is likely to make you anxious. A new oversight organization is monitoring incidents of AI behaving improperly, and the number of occurrences is likely to make you anxious. According to recent research, in July, incidents of AI models deceiving, plotting, and circumventing security measures nearly doubled, including an actual hacking campaign involving Anthropic and OpenAI models. Just $1 was enough for Texas to fill the state with Flock surveillance cameras funded by motorists. Just $1 was enough for Texas to fill the state with Flock surveillance cameras funded by motorists. Approximately 3,200 Flock cameras in Texas were financed by funds generated from a $1 auto insurance fee that was initially authorized to address vehicle crime.

AI scribes are inaccurately recording medication names and diagnoses in NHS records.

Healthwatch England discovered that AI scribes are misrecording medications and diagnoses. There are currently twenty-seven in operation, and none are regulated as medical devices.