Blog
From Distant Glow to Floodlight: The Long Arrival of AI in Healthcare Part 4 of 5
Artificial intelligence spent fifty years approaching the clinic. Here is what changed when it finally walked in.
· Charles Faul

Every previous chapter of this series has really been about preparation. The records went digital, the departments acquired their tools, the pipes slowly connected, and all the while a light waited on healthcare’s horizon: artificial intelligence, forever described as five years away. This instalment follows that light as it actually behaved: the glow, the stalls, the false dawn, and an arrival that looked nothing like the prophecy.
“It’s just completely obvious that within five years deep learning is going to do better than radiologists. … It might be 10 years, but we’ve got plenty of radiologists already.”
Geoffrey Hinton, deep learning pioneer and Nobel laureate, speaking in 2016 [1]
A Machine That Was Right, and Never Used
In the early 1970s at Stanford, a program called MYCIN questioned doctors about their patients and recommended antibiotic therapy. In a blinded evaluation published in JAMA, infectious disease experts rated its prescriptions favourably against those of human specialists [2]. And then, nothing. In the words of its own creators, “MYCIN was never used routinely in patient-care settings” [3]. The reasons were not technical. A consultation required the doctor to find a terminal, log in and answer a long interrogation, much of it about results already sitting in other hospital computers; the developers estimated that a full session “could have required as long as 30 minutes or an hour”, which they conceded was “clearly unacceptable” [3]. Readers of part two will recognise the pattern half a century before it had a name: the technology was capable, but the workflow could not carry it.
Winters and a False Dawn
Funding froze in the periods now remembered as AI winters, and healthcare busied itself with the digitisation projects of parts one and two. When artificial intelligence returned to medicine’s front pages, it did so as a marketing campaign. IBM’s Watson, fresh from winning the quiz show Jeopardy! in 2011, was promoted as a revolution in cancer care. The reckoning arrived in stages. By late 2016, the University of Texas MD Anderson Cancer Center had shelved its Watson-based Oncology Expert Advisor project, an audit recording roughly 62 million US dollars spent on a system that was still not ready for clinical use [4]. In 2018, internal IBM documents obtained by STAT News showed the flagship oncology product had recommended “unsafe and incorrect” cancer treatments [5]. IEEE Spectrum’s post-mortem of the whole programme was titled, simply, How IBM Watson Overpromised and Underdelivered on AI Health Care [6], and in 2022 IBM sold Watson Health’s data assets to a private equity firm, unwinding a business it had reportedly spent four billion dollars assembling [7].
Hinton’s 2016 prediction about radiologists belongs to the same era of confidence, and its fate is instructive in a subtler way. The machines did become very good at reading images, yet radiologists were not replaced, because performing a task is not the same as doing a job. The prediction failed not on capability but on category, and that distinction, between what a model can do and what a health system can safely absorb, is the most reliable lesson this half-century offers.
What Quietly Worked
While the headlines chased Watson, quieter work was compounding, and it looked deliberately unglamorous. In 2016, Google researchers showed in JAMA that a deep learning system could detect referable diabetic retinopathy in retinal photographs with sensitivity of 87% to 90% and specificity above 98%, judged against a panel of board-certified ophthalmologists [8]. Two years later a system called IDx-DR screened for the same disease in ordinary primary care offices, achieving 87.2% sensitivity and 90.7% specificity in its pivotal trial and becoming what its study authors describe as “the first FDA authorized autonomous AI diagnostic system in any field of medicine” [9]. The trickle became a current: the US regulator’s list of AI-enabled medical devices now runs past 1,400 entries, roughly three quarters of them in radiology [10, 11]. Notice the shape of the success. These tools passed trials before they reached patients, took on narrow and checkable tasks, and slotted in beside the specialists rather than replacing them. Evidence first, deployment second: the exact reverse of the Watson sequence. The light was learning manners.
The Language Models Arrive
The next shift was of kind, not degree. Large language models, trained on text at unprecedented scale, proved able to handle medical knowledge itself. In 2023 a Google research system published in Nature scored 67.6% on questions styled after the US medical licensing examination, the first convincing crossing of that bar, and within months its successor reached 86.5% [12, 13]. Yet answering examination questions is not practising medicine, and a model that knows facts is not yet a system that can be trusted with a live consultation. What the benchmarks did establish is that machines could now work in ordinary clinical language, and that opened a different door: not diagnosis, but documentation.
The most consequential early evidence comes from The Permanente Medical Group in California, where 7,260 doctors used ambient AI scribes across more than 2.5 million patient encounters [14, 15]. The headline estimate, 15,791 hours of documentation time returned in 63 weeks, sounds enormous, until you do the arithmetic: spread across millions of visits it amounts to well under a minute per encounter, measured observationally in a single health system, with every note still reviewed and signed by the doctor [15]. The more consistent finding, there and in the wider literature, is not minutes but attention. 84% of the doctors reported a positive effect on communication with their patients, and a 2026 review concluded that “while time savings vary, the most consistent value proposition for ambient scribes lies in reduced cognitive load, enhanced patient engagement, and clinician well-being” [15, 16]. On the scale this series has been applying to every technology since the mainframe: promising, not yet proven. That is already more than most of this chapter’s history can claim.
A View from the South
For African health systems, the story looks different, because the scarcity is different. Radiologists are rare across much of the continent, which is why the World Health Organization has recommended computer-aided detection for tuberculosis screening since 2021; in June 2025 it went further, announcing that six software products had met its performance standards for people aged 15 and older, after independent validation by FIND [17, 18]. This is global health policy formally putting algorithms to work at population scale, and it happened for the high-burden countries of the South first.
The same region supplied the decade’s most instructive failure, and it deserves more than a footnote. Babyl, Rwanda’s AI-assisted telemedicine service, was by most clinical measures a success. It registered more than 2.5 million users, nearly a fifth of the country, integrated with the national insurance system, and held a ten-year partnership with the Rwandan government; research on its consultations suggested healthcare professionals asked more questions in less time than comparable in-person care [19]. None of that mattered in August 2023, when its London-based parent collapsed into bankruptcy and the service simply stopped. A fifth of a nation lost its digital front door to healthcare overnight, because of decisions taken in boardrooms an ocean away [19]. The light can arrive from far away; owned from far away, it can be switched off from far away too.
What the Floodlight Shows
Fifty years separate MYCIN’s quiet shelving from the ambient scribes, and the difference between them is not intelligence so much as fit and proof. MYCIN demanded the doctor’s time; the scribe is designed to return attention. Watson was sold before it was proven; the systems that lasted were proven before they were sold. The light on the horizon has finally switched on, and like any floodlight it shows everything: the patchwork, the disconnected pipes, the tired workforce, and also how far the new tools still have to go.
So the question this series must answer in closing is not whether the technology has arrived. It is whether it has earned the right to be trusted, and by what standard we should judge it. It might not be different this time. That is precisely why the standard of evidence has to be higher. Part five is about that standard: what repayment would actually look like, and the guardrails that decide whether trust is stolen or returned.
References
1. The New Republic. The “Godfather of AI” Predicted I Wouldn’t Have a Job. He Was Wrong. 2024. https://newrepublic.com/article/187203/ai-radiology-geoffrey-hinton-nobel-prediction
2. Yu VL, Fagan LM, Wraith SM, et al. Antimicrobial Selection by a Computer: A Blinded Evaluation by Infectious Diseases Experts. JAMA. 1979;242(12):1279-1282.
3. Buchanan BG, Shortliffe EH (eds). Rule-Based Expert Systems: The MYCIN Experiments of the Stanford Heuristic Programming Project. Addison-Wesley; 1984. Chapter 32. https://people.dbmi.columbia.edu/shortliffe/Buchanan-Shortliffe-1984/Chapter-32.pdf
4. Engadget. IBM’s Watson AI runs into trouble fighting cancer. 20 February 2017. https://www.engadget.com/2017-02-20-watson-cancer-research-runs-into-trouble.html
5. Ross C, Swetlitz I. IBM’s Watson supercomputer recommended “unsafe and incorrect” cancer treatments, internal documents show. STAT News, 25 July 2018. https://www.statnews.com/2018/07/25/ibm-watson-recommended-unsafe-incorrect-treatments/
6. Strickland E. How IBM Watson Overpromised and Underdelivered on AI Health Care. IEEE Spectrum, 2019. https://spectrum.ieee.org/how-ibm-watson-overpromised-and-underdelivered-on-ai-health-care
7. Fierce Healthcare. IBM sells Watson Health assets to investment firm Francisco Partners. 21 January 2022. https://www.fiercehealthcare.com/tech/ibm-sells-watson-health-assets-to-investment-firm-francisco-partners
8. Gulshan V, Peng L, Coram M, et al. Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs. JAMA. 2016;316(22):2402-2410.
9. Abràmoff MD, Lavin PT, Birch M, Shah N, Folk JC. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digital Medicine. 2018;1:39. doi:10.1038/s41746-018-0040-6
10. US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices (device list). https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
11. The Medical Futurist. The Current State of FDA-Approved AI-Enabled Medical Devices. March 2026. https://medicalfuturist.com/the-current-state-of-fda-approved-ai-based-medical-devices/
12. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620:172-180. doi:10.1038/s41586-023-06291-2
13. Singhal K, Tu T, Gottweis J, et al. Towards Expert-Level Medical Question Answering with Large Language Models (Med-PaLM 2). arXiv:2305.09617, 2023. https://arxiv.org/abs/2305.09617
14. Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst. 2025. doi:10.1056/CAT.25.0040
15. American Medical Association. AI scribes save 15,000 hours, and restore the human side of medicine. https://www.ama-assn.org/practice-management/digital-health/ai-scribes-save-15000-hours-and-restore-human-side-medicine
16. Ohde JW, Thompson A, Liu Z, et al. Barriers and opportunities of scaling ambient AI scribes for clinical documentation across diverse healthcare settings. npj Digital Medicine. 2026;9:369. doi:10.1038/s41746-026-02554-0
17. World Health Organization. Use of computer-aided detection software for tuberculosis screening. https://www.who.int/publications/b/79103
18. World Health Organization. WHO approves six software products for computer-aided detection of TB on chest X-ray. 11 June 2025. https://www.who.int/news/item/11-06-2025-who-approves-six-software-products-for-computer-aided-detection-of-tb-on-chest-x-ray
19. ICTworks. Babyl Paradox: When Evidence-Based Digital Success Cannot Beat Corporate Failure. https://www.ictworks.org/digital-success-cannot-beat-corporate-failure/
