Blog
From Stolen Hours to Human Touch: What AI Must Prove to Medicine Part 5 of 5
Real-world evidence, real guardrails, and the standard the next generation of healthcare software has to meet
· Charles Faul

We began this series in a room full of filing cabinets, and we arrive now at its final question. Software digitised medicine’s memory, then specialised, then fragmented, then quietly claimed the clinical hour; artificial intelligence approached for fifty years and finally arrived. The theft was measured in hours, but what the hours contained was attention: the looking, listening and noticing that medicine is actually made of. Will the new intelligence give that attention back, or simply become the next thief with better manners? This series ends by asking what proof would look like, and who gets to set the standard. Because our team builds this technology ourselves, we have less licence than anyone to confuse promising results with proven benefit. This closing instalment applies the same suspicion to AI that the earlier parts applied to everything else.
“The secret of the care of the patient is in caring for the patient.”
Francis W. Peabody, MD, address to Harvard medical students, 1926, published in JAMA in 1927 [1]
Attention, Measured
Start with what can already be observed. At the deployment we met in part four, 84% of doctors using ambient scribes reported a positive effect on communication with their patients, and the patients noticed the same thing from the other chair: 47% saw their doctor spending less time looking at the computer, and 39% said their doctor spent more time speaking directly to them [2]. Those are the numbers to hold on to, more than any tally of minutes, because they measure the very thing part two showed being lost. The time estimates are earlier, softer evidence: observational, aggregated, drawn so far from single health systems rather than randomised comparisons, and every AI-drafted note still needs the doctor’s review and signature [2, 3]. The wider literature is converging on the same emphasis, finding that the most consistent benefits are reduced cognitive load, better patient engagement and improved wellbeing, with time savings varying from setting to setting [4]. For a generation, the person in the consulting room who received the least attention was the one on the examination couch. The early evidence says that is beginning to reverse. Proving it durably, across systems, languages and populations, is the work still ahead.
A Colleague in the Corner
A different class of AI, and it is important to keep the classes distinct, assists with the decisions themselves. In Nairobi, the primary care network Penda Health ran an AI clinical copilot, developed with OpenAI, across 39,849 patient visits in 15 clinics. The clinical teams using it made 16% fewer diagnostic errors and 13% fewer treatment errors, with the largest gains in the visits that triggered its most serious safety alerts [5]. Two things temper the result. Although clinicians were randomised, the study was funded and published by OpenAI, which also took part in the analysis, and it remains a preprint; when an independent randomised trial of the same tool, funded by the Gates Foundation, followed in 2026, it found no significant reduction in treatment failure [12]. And it is evidence about decision support, not about documentation tools like ambient scribes; the categories should not be allowed to borrow each other’s results. What makes the study encouraging is its design philosophy. The tool watches quietly, speaks rarely and escalates only when it must, and the human makes the final call. But part two taught this series what oversight is worth when it is badly designed: people learn to swat away noisy help. The mirror-image failure, automation bias, can quietly turn “the human decides” into “the human agrees”. A safety net is only as good as the attention it leaves the human free to give, so the guardrail must be built to make oversight real rather than ceremonial.
A View from the South: Where the Light Lands
South Africa carries one of the world’s heaviest tuberculosis burdens, and several of the screening products endorsed by the WHO, six of which now meet its performance standards, have been evaluated against films from South Africa’s own national TB prevalence survey [6, 7]. What is new is the reach. Paired with the latest ultraportable X-ray machines, screening can now travel by bakkie to communities no radiologist will ever visit [8]. And remember the nurse from part two, writing every consultation into a paper register for a clerk to retype. Everything this series has examined converges on her: intelligence that documents, retrieves and reports quietly in the background would return more attention in a two-nurse clinic in Limpopo than in any Californian medical centre. Whether it actually reaches her, in her language, on her clinic’s connectivity and under her country’s law, is the real test of this technology.
The Guardrails
None of this is guaranteed, and this is where the series must apply its own standard hardest. Start with bias. In 2019, researchers showed in Science that a widely used American care management algorithm, built on past healthcare costs as its proxy for medical need, had systematically under-served Black patients: correcting it would have raised the share of Black patients flagged for additional help from 17.7% to 46.5% [9]. Babyl’s overnight disappearance from Rwanda, told in part four, carries the matching lesson about ownership: a clinically successful service is not safe if its off-switch lives on another continent. And the WHO has responded to the newest generation of systems with dedicated guidance on the ethics and governance of large multi-modal models in health [10].
For documentation AI specifically, the checklist is longer, and a company that builds these tools should be the first to write it down. What are the omission and hallucination rates, and who measures them? When a note attributes words to the wrong speaker, or drops a negation, who catches it? Has the patient consented to the recording; where does the audio live, for how long, and under whose law? South Africa’s POPIA makes purpose limitation a legal obligation, not a preference, and bars moving patient data offshore unless one of its section 72 conditions is met. What is the medico-legal status of an AI-drafted note, and can every note be traced to the model version that produced it and to the correction that followed? How does accuracy hold up across accents and languages in a country with twelve official ones? And what does the system do when it is uncertain? These are not reasons to refuse the technology. They are the price of trusting it, and any vendor unwilling to answer them in public has answered them already.
Full Circle
In 1926, before the first electronic computer existed, Francis Peabody told a hall of Harvard medical students the thing this whole series has been circling: the secret of the care of the patient is in caring for the patient [1]. A century of technology has neither disproved him nor, until recently, served him particularly well. Paper buried the caring under filing cabinets. The first software wave freed the record and captured the doctor. The great digitisation connected the billing before it connected the care. And then, slowly and at last, machines arrived that could listen, summarise, flag and fetch, and attention began to flow back towards the bedside.
Which returns us to where part one began, with Eric Topol’s conviction that AI’s greatest opportunity “is the opportunity to restore the precious and time-honored connection and trust - the human touch - between patients and doctors” [11]. Five instalments later, the verdict I can stand behind is this: the opportunity is real, and the outcome is not yet earned. Every earlier generation of healthcare software solved one problem and created another, and this one will be no exception unless it is held to a harder standard than its predecessors, by the doctors, nurses and patients it claims to serve. The purpose of healthcare software was never the software. It was always the two people in the room.
A note on vantage point: Akili AI, where I work, builds clinical documentation and workflow technology for healthcare teams in Africa and beyond. This series is written from inside that effort, which is exactly why its standard of evidence needs to be higher than its enthusiasm.
References
1. Peabody FW. The Care of the Patient. JAMA. 1927;88:877-882. See also: Hektoen International, Francis Peabody: caring for the patient. https://hekint.org/2017/02/01/francis-peabody-caring-for-the-patient/
2. American Medical Association. AI scribes save 15,000 hours, and restore the human side of medicine. https://www.ama-assn.org/practice-management/digital-health/ai-scribes-save-15000-hours-and-restore-human-side-medicine
3. Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst. 2025;6(5). doi:10.1056/CAT.25.0040
4. Ohde JW, Thompson A, Liu Z, et al. Barriers and opportunities of scaling ambient AI scribes for clinical documentation across diverse healthcare settings. npj Digital Medicine. 2026;9:369. doi:10.1038/s41746-026-02554-0
5. OpenAI and Penda Health. Pioneering an AI clinical copilot with Penda Health. 2025. https://openai.com/index/ai-clinical-copilot-penda-health/
6. World Health Organization. WHO approves six software products for computer-aided detection of TB on chest X-ray. 11 June 2025. https://www.who.int/news/item/11-06-2025-who-approves-six-software-products-for-computer-aided-detection-of-tb-on-chest-x-ray
7. Computer-aided detection of tuberculosis from chest radiographs in a tuberculosis prevalence survey in South Africa: external validation and modelled impacts of commercially available artificial intelligence software. The Lancet Digital Health. 2024. doi:10.1016/S2589-7500(24)00118-3
8. Min J, Halton J, Villegas C, Chua A, Hewison C, Deborggraeve S. Scanned: The global investments in computer-aided detection and ultraportable X-ray for tuberculosis. PLOS Global Public Health. 2025;5(3):e0004232. doi:10.1371/journal.pgph.0004232
9. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342
10. World Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. https://www.who.int/publications/i/item/9789240084759
11. Topol E, quoted in part one of this series. See: Topol E. Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again. Basic Books; 2019.
12. Agweyu A, Mwaniki P, Menon V, et al. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial. Nature Medicine. 2026;32:3032-3039. doi:10.1038/s41591-026-04503-6
