

A consultant in an outpatient clinic finishes seeing a patient, and somewhere in the room a microphone has already turned the conversation into a draft note. Nobody typed it. Nobody is checking, in real time, whether it captured the right allergy, the right dose, the right nuance of what the patient actually said. That note will sit in the record, reviewed later, trusted mostly by default. This is what ambient voice technology now does across a growing share of NHS consultations, and this week the Health Services Safety Investigations Body confirmed what many close to the rollout have quietly suspected: nobody yet has a reliable way of knowing when it goes wrong.
HSSIB's new investigation into ambient voice technology follows engagement with national bodies who flagged safety concerns that adoption had outpaced. The body's own language is notable for its restraint rather than alarm. It describes routes for recognising and reporting AI-related incidents as not yet mature enough to give confidence that emerging risks are being identified. That is a careful way of saying the NHS does not currently know how often these tools mishear, mistranscribe or misattribute clinical information, because the mechanisms for finding out are still being built.
The timing exposes an awkward mismatch. HSSIB expects to publish its findings in summer 2027. In the eighteen months before that report lands, ambient voice technology is not pausing for evaluation. NHS Midlands has just completed a region-wide procurement of one supplier's ambient scribe covering more than a thousand GP practices and fifteen acute and community trusts, deployment already live in several. Nineteen suppliers sit on NHS England's self-certified registry, a mechanism built explicitly to accelerate adoption by letting vendors evidence their own compliance rather than undergo centralised assessment. Other trusts are running their own tenders in parallel. The infrastructure for scaling this technology is considerably further advanced than the infrastructure for knowing whether it is safe.
Complicating matters further, the Medicines and Healthcare products Regulatory Agency published guidance earlier this month clarifying that some ambient voice products count as medical devices and others do not, depending on what they are used for. Tools limited to transcription, summarisation or letter drafting fall outside device regulation altogether, on the reasoning that a clinician reviews the output before it is acted upon. That distinction may hold up in principle. In a system where documentation time is chronically scarce and review is itself under pressure, it asks a great deal of individual clinicians to catch what an ambient system gets wrong, consistently, appointment after appointment, without the incident data to tell them what kinds of errors to watch for.
None of this means the technology is unsafe, and the early evidence on time saved is genuinely encouraging. A national evaluation published earlier this year found consistent reductions in documentation time, alongside a frank admission that nobody yet understands whether those minutes translate into better patient outcomes or simply get absorbed into an already stretched system. That gap between measured outputs and understood outcomes is precisely the space HSSIB's investigation now has to work in, and it is a familiar one. The Federated Data Platform arrived under similar terms: a technology procured and scaled first, with governance assembled around it after the fact rather than ahead of it.
What distinguishes this case is the self-certification model itself. Suppliers evidencing their own safety criteria is a reasonable way to move quickly in a system desperate for capacity. It is a poor substitute for independent assurance once deployment reaches tens of thousands of clinicians and over a million patient consultations. NHS leaders who signed procurement deals this year are not wrong to have done so, but they are now operating years ahead of the evidence base that would tell them what they have actually bought. For an investigation body whose findings will not land until 2027, the honest task is not to slow adoption that has already happened, but to build, quickly, the reporting architecture that should have existed before it did.