In 2016, one of the most influential figures in modern machine learning suggested publicly that medical schools should stop training radiologists, because the profession would be obsolete within a few years.
A decade later, radiology has a workforce shortage in most developed countries. Imaging volumes have grown faster than the number of people qualified to read them. Departments are running backlogs. And the software that was supposed to replace radiologists is instead being used to decide which scan they should look at first.
That gap between prediction and outcome is the most useful thing to understand about AI in medical imaging, because the reasons for it explain what the technology is actually good at.
Why the Prediction Failed
The forecast assumed the radiologist's job was image interpretation. It is a portion of the job, and not the portion that resisted automation.
What the role actually involves:
- Deciding which examination is appropriate for a clinical question.
- Setting the acquisition protocol and adjusting it for the individual patient.
- Interpreting findings in the context of the patient's history, previous imaging, and current clinical picture.
- Communicating uncertainty to the referring clinician, often in conversation.
- Performing image-guided procedures.
- Taking responsibility for the conclusion.
A model trained to detect a specific finding on a specific type of image addresses one narrow slice of the second-to-last item. It performs that slice well. It does not touch the rest, and the rest is where most of the clinical value and all of the legal responsibility sits.
What It Actually Does in a Department
The applications in routine use are less dramatic and more useful than the original promise.
- Worklist prioritisation. Scanning incoming studies and moving likely-urgent cases to the front of the queue.
- Detection assistance. Marking regions for the radiologist's attention — nodules, fractures, haemorrhages, lesions.
- Quantification. Measuring volumes, dimensions, and changes over time, consistently.
- Image quality and reconstruction. Producing usable images from faster acquisitions or lower radiation doses.
- Protocol optimisation, selecting sequences based on the clinical question.
- Reporting support, drafting structured reports from findings for the radiologist to edit and sign.
- Comparison with priors, automatically registering a current study against previous ones.
Note how many of these are workflow rather than interpretation. That is where the demonstrated value has accumulated.
Triage Is the Application That Works
If one use case has clearly justified itself, it is urgent case prioritisation.
The clearest example is acute stroke. Software that analyses a scan immediately after acquisition, identifies a suspected large vessel occlusion, and alerts the intervention team directly — rather than waiting for the study to reach a radiologist in queue order — compresses the interval before treatment. In a condition where treatment effectiveness declines sharply by the hour, that compression is meaningful in a way that a marginal accuracy improvement is not.
The general principle applies beyond stroke: the value is not in reading the scan better. It is in reading the right scan sooner.
Departments using this well have restructured around it — defining who receives the alert, what they do, and what happens if the flag turns out to be wrong. Departments that installed the software without changing the process report little benefit, which is not a surprise.
Quantification: The Boring Superpower
The least discussed advantage is consistency in measurement.
Human measurement of a lesion, a volume, or a change over time varies — between radiologists, and for the same radiologist on different days. That variability matters enormously when the clinical decision depends on whether something grew.
Automated quantification is not necessarily more accurate in an absolute sense. It is more reproducible, which for tracking change is often the more valuable property. Applications include tumour response assessment, cardiac function measurement, brain volume tracking in neurological conditions, and bone density.
This is unglamorous, well-validated, and steadily changing how follow-up imaging is interpreted.
MRI Specifically: Speed Is the Real Gain
Magnetic resonance imaging deserves separate attention, because the most valuable application there is not interpretation at all. It is acquisition.
An MRI examination is slow, and the slowness is a clinical problem rather than a scheduling inconvenience. Patients must remain still for extended periods; children and people in pain frequently cannot, which means sedation or a failed study. Scanner time is scarce, which means waiting lists.
Reconstruction methods that produce diagnostic-quality images from less acquired data attack this directly. The scan captures fewer measurements, and the reconstruction fills in the rest. The practical consequences:
- Shorter examinations, which reduces motion artefact and improves image quality in practice.
- Fewer sedated paediatric studies.
- More patients scanned per day on the same machine.
- Studies becoming feasible for patients who could not previously tolerate them.
The equivalent gain in computed tomography is dose reduction — reconstructing diagnostic images from lower radiation exposure, which matters most for patients requiring repeated scanning.
The caution attached to both: a reconstruction algorithm is generating plausible image content from incomplete data. Done well, the result is diagnostically equivalent. Done badly, or applied outside the conditions it was validated for, it can produce images that look clean while omitting or inventing fine detail. This is an active area of scrutiny, and departments adopting it need to understand what the method was validated on.
The Reading Room Reality
Two documented human factors determine whether any of this helps, and both cut against the technology.
Automation bias. When a system marks a region, radiologists are more likely to attend to it — and, more concerningly, less likely to scrutinise regions it did not mark. A tool with high sensitivity for one finding can reduce detection of findings it was never trained on.
Alert fatigue. A system producing frequent low-value flags trains its users to dismiss them reflexively, including the correct ones. This is a well-established pattern across clinical decision support generally.
Mitigations that departments have found necessary:
- Setting operating thresholds for clinical utility rather than for benchmark performance.
- Displaying the tool's output after an initial unaided read, rather than before.
- Auditing disagreements between the system and the radiologist, in both directions.
- Making clear what the tool was not trained to detect.
- Monitoring whether detection of other findings declines after deployment.
That last measure is rarely taken and is arguably the most important.
Clearance Is Not Evidence of Benefit
Regulators have authorised a large number of imaging AI products, and radiology accounts for the majority of authorised medical AI devices overall.
But most of those authorisations rest on demonstrating that the device performs comparably to a predicate or to a reference standard on retrospective data. Relatively few rest on prospective evidence that patients managed with the tool have better outcomes than patients managed without it.
Prospective trials do exist in some areas. Studies of AI-supported mammography screening in European settings have examined whether the approach can maintain cancer detection while reducing reading workload, with published results that have been encouraging in specific screening contexts. That kind of evidence — prospective, in a real programme, measuring what happened — is the standard worth looking for, and it remains the exception rather than the rule.
Where It Fails
The failure modes are consistent and predictable:
- Distribution shift. A model trained on images from one scanner manufacturer, one protocol, one population frequently degrades on another. This is the single most common cause of disappointing real-world performance.
- Population differences. Disease prevalence and presentation vary between populations, and a model calibrated for one may be poorly calibrated for another.
- Rare findings. Models are trained on what is common. The unusual case — which is precisely where a radiologist's judgement matters most — is where they are weakest.
- Incidental findings outside the model's task, which it will not mention.
- Silent degradation. Performance can decline gradually as equipment, protocols, and populations change, with nothing in the output indicating that it has.
That final point argues for ongoing monitoring rather than one-time validation, and few institutions have the infrastructure for it.
What This Means for Radiologists
The realistic picture, based on how the past decade actually unfolded:
- Volumes are rising faster than workforce, and the tools are absorbing part of the gap rather than replacing anyone.
- Time is shifting from measurement and search toward interpretation, communication, and procedures.
- Familiarity with these systems — including their failure modes — is becoming part of professional competence.
- Responsibility has not moved. The signature on the report belongs to a person.
The role most affected is not the radiologist but the technologist and the reporting workflow around them, where protocol selection, image quality checking, and structured reporting are all being partially automated.
The Reasonable Summary
AI in medical imaging has delivered real, measurable value in a narrower band than promised: prioritising urgent studies, measuring consistently, improving image quality at lower dose, and reducing the administrative weight of reporting.
It has not delivered autonomous interpretation, and the obstacles to that are not primarily technical. They are about context, responsibility, rare cases, and the fact that a scan is one input into a clinical decision rather than the decision itself.
Ten years on from a confident prediction of obsolescence, the profession is short-staffed and the software is helping it cope. That is a less exciting story than the one originally told, and a considerably more useful one.
This article is general information about healthcare technology, not medical advice. For questions about a scan or your own care, speak to a qualified clinician.

0 Comments