The most widely deployed medical artificial intelligence in 2026 is not diagnosing anything. It is writing notes.
Ambient documentation systems — software that listens to a consultation and drafts the clinical note — have spread through hospitals and clinics faster than any diagnostic tool ever has. Not because note-writing is medically interesting, but because clinicians were spending a substantial portion of every working day on documentation, and burnout had become a workforce crisis in its own right.
That is a useful corrective to how this subject is usually presented. AI in healthcare entered mainstream practice through the administrative door, not the diagnostic one, and the reasons are worth understanding before examining anything else.
Where It Actually Sits in the Care Pathway
It helps to map the applications against the journey a patient actually takes, because the maturity varies enormously between stages.
- Access and triage. Symptom assessment, appointment routing, and prioritising who needs to be seen urgently. Widely deployed, variable quality.
- Screening. Population-level detection in people without symptoms. Strong evidence in a few specific areas, weak in most.
- Diagnostic support. Assisting a clinician in interpreting an image, a trace, or a set of results. The most regulated and best-validated category.
- Treatment planning. Dosing, radiotherapy planning, drug interaction checking, protocol matching.
- Monitoring. Tracking deterioration in hospital and chronic conditions at home.
- Administration. Documentation, coding, scheduling, prior authorisation, and discharge summaries.
The public conversation focuses on diagnosis. The actual deployment weight sits at the two ends — triage and administration — where the regulatory burden is lower and the operational pain is higher.
What "Cleared" Actually Means
This is the single most misunderstood point in the entire field, and it matters for anyone evaluating a claim.
Regulators have authorised a large number of AI-enabled medical devices, the majority of them in radiology. But regulatory clearance in most jurisdictions establishes that a device is substantially equivalent to something already on the market and performs as described. It does not establish that using it improves patient outcomes.
Those are very different bars. A tool can accurately detect a finding and still fail to help, because:
- The finding was already being caught reliably by existing practice.
- Detecting it earlier does not change what happens to the patient.
- The additional alerts consume clinician attention that was better spent elsewhere.
- The workflow around it was never redesigned, so the output is ignored.
When you read that a system "detects condition X with high accuracy," the useful follow-up question is not about the accuracy figure. It is whether anyone has shown that patients treated in departments using it do better than patients in departments that do not.
Diagnosis: Assistive, With One Notable Exception
Almost all diagnostic AI in clinical use operates as a second opinion. It flags, highlights, quantifies, or prioritises — and a clinician makes the decision and carries the responsibility.
The notable exception is autonomous diabetic retinopathy screening. Systems authorised to produce a screening result without a specialist reviewing the image have been in use since 2018, when the first such authorisation was granted in the United States. The reason this particular application went first is instructive: the task is narrowly defined, the imaging is standardised, the disease is common, and specialists to read the images are scarce in exactly the places the screening is needed.
Those four conditions rarely occur together. Where they do, autonomous operation becomes defensible. Where they do not — which is most of medicine — the assistive model remains the standard, and will for the foreseeable future.
Treatment: Narrow Wins, Broad Disappointments
Treatment planning has produced the field's most instructive failure and some of its quietest successes.
The failure is well documented. An ambitious effort to build a system that would recommend cancer treatments across the full breadth of oncology was scaled back after struggling to deliver clinical value. The problem was not computing power. It was that oncology treatment decisions depend on patient preference, comorbidity, local drug availability, trial access, and clinical judgement — most of which is not in any dataset.
The quiet successes are narrower:
- Radiotherapy planning, where contouring organs at risk is time-consuming and geometrically well-defined.
- Dosing support for medications with narrow therapeutic windows.
- Drug interaction and allergy checking, unglamorous and genuinely protective.
- Protocol and trial matching, finding the small number of patients eligible for a specific study.
- Sepsis and deterioration alerting, with important caveats discussed below.
The pattern holds across the field: narrow, well-bounded tasks succeed; broad clinical reasoning does not.
The Failure That Should Be Taught Everywhere
In 2021, researchers published an external validation of a widely deployed proprietary sepsis prediction model used across many American hospitals. The finding was uncomfortable: the model performed substantially worse in practice than the developer's own reported figures, missing a large share of sepsis cases while generating a considerable volume of alerts.
The lesson is not that sepsis prediction is impossible. It is about what happens when a model is deployed at scale on the basis of internal validation, without independent evaluation in the settings where it is used.
Three practical implications for anyone assessing a clinical tool:
- Internal validation predicts very little. Performance on the developer's own data routinely fails to transfer.
- Population shift matters enormously. A model built on one health system's patients may perform poorly on another's.
- Alert burden is a clinical harm, not just an annoyance. Clinicians who receive too many alerts stop reading them, including the correct ones.
The Bias Problem Is Not Hypothetical
A landmark 2019 study published in Science examined an algorithm used across the American health system to identify patients needing extra care. It found that the algorithm systematically under-referred Black patients relative to their actual level of illness.
The mechanism is worth understanding because it recurs. The algorithm did not use race as an input. It used healthcare spending as a proxy for healthcare need — and because less money had historically been spent on Black patients at equivalent levels of illness, the model learned to score them as healthier.
The general form of this failure: a model trained on a proxy for the thing you care about will faithfully reproduce whatever inequity is embedded in that proxy. It will do so invisibly, at scale, and with the appearance of objectivity.
Related documented issues include devices and datasets that perform differently across skin tones, and models trained overwhelmingly on populations from a small number of wealthy countries then deployed globally.
Back to the Scribe: Why It Won First
It is worth returning to where this article started, because the reasons ambient documentation spread so quickly explain a great deal about what succeeds in healthcare technology generally.
The application had an unusual combination of properties:
- The pain was acute and universal. Documentation burden affects every clinician in every specialty, and it was measurably driving people out of the profession.
- The regulatory path was light. A tool that drafts a note a clinician then reviews and signs is not making a clinical decision, which places it outside the strictest device categories in most jurisdictions.
- Failure is visible and recoverable. A wrong sentence in a draft note is caught by the clinician signing it. A wrong diagnostic flag may not be caught at all.
- The benefit accrues to the person using it. Most clinical AI benefits an institution while costing a clinician time. This inverted that.
That last point is the one healthcare technology has historically got wrong. Systems that generate value for administrators by adding work for clinicians get resisted, quietly and effectively, regardless of how well they perform.
The caution that comes with it: a drafted note is a document with legal and clinical weight. Errors introduced by transcription or summarisation become part of a permanent record, and the clinician who signs it owns them. Reviewing rather than skimming is not optional, and the time pressure that made the tool attractive is exactly what discourages careful review.
What Patients Should Realistically Expect
If you are receiving care in 2026, here is roughly where things stand:
- Your consultation may be recorded and transcribed by an ambient documentation system. You should be told, and you can ask.
- Your scans may be pre-screened by software that flags urgent findings for faster review. A radiologist still reads them.
- Your triage may be partly automated, particularly for appointment routing.
- Your diagnosis is made by a clinician. Anything you are told about your condition came through a human decision.
- Your data may be used for research, subject to the consent framework where you live, and you generally have the right to ask about this.
- You can ask questions. Whether a tool was involved, what it indicated, and what your clinician thought about it are reasonable things to ask.
The Honest Position
The gap between capability and clinical benefit remains wide, and it is closing more slowly than the announcements suggest.
What has clearly worked: narrow detection tasks with standardised inputs, workflow prioritisation, documentation relief, and screening in settings where the alternative was no screening at all.
What has repeatedly disappointed: broad clinical reasoning, systems deployed without independent validation, and tools bolted onto workflows nobody redesigned.
The most reliable predictor of whether a hospital gets value from AI in healthcare is not which system it purchased. It is whether the organisation was willing to change how work is done around it — and whether anyone measured what happened to patients afterwards.
That measurement, more than any model architecture, is what separates the deployments that helped from the ones that merely happened.

0 Comments