Back to the blog
Physical AI for critical environments July 13, 2026 13 min read

AI in infection prevention: what actually works and what is hype

Every vendor this year calls their product AI-driven, and most of the time you have no way to check. Here is the honest sort: what in AI and automation for infection prevention is proven, what is emerging, and what is only marketing, and how to tell them apart.

AI in infection prevention: what actually works and what is hype — ROZOR
Quick answer

AI in infection prevention sorts into proven, emerging and hype. Proven means automation and sequencing, not machine learning: electronic surveillance that counts cases accurately, genomic sequencing that finds outbreaks conventional investigation misses, and automated measurement of hand hygiene. Emerging means risk prediction, which discriminates well but has never been shown to prevent an infection.

Every vendor who walks into your office this year will use the word. The surveillance module is AI-driven, the badges are AI-powered, the disinfection robot has an AI brain. You have to decide which of those sentences means anything, usually with no way to check. There is a way, and it is simpler than the marketing would like: sort every claim by what it was actually measured against.

Does AI actually prevent infections in hospitals?

Three technologies in this field earn their place in an infection prevention programme today. None of them is artificial intelligence, and the rest of this article is the evidence.

A 2025 scoping review led by WHO's infection prevention team catalogued 100 real-world studies of AI in IPC published between 2014 and 2024. More than half, 53 percent, were predictive analytics, and machine learning dominated the methods at 75 percent. The reviewers found the field held back by "limited prospective validation or implementation science."1 Across all of medicine, one systematic review could find just 41 randomised trials of any machine-learning intervention, half of them single-site.2 No published system has yet been shown to lower a healthcare-associated infection rate.

Four different things get called "it works," which is how a demo can be impressive and useless at once. Discrimination is how well a model sorts patients, the area under the curve on the slide, and it is not an outcome. Detection is what a system saw: clusters found, cases flagged. Detection is not prevention. Process is what people did: compliance rates, cleaning scores. Compliance is not infection. Outcome is what happened to patients, and only that tier earns the word "works." Hold the four apart and you can read almost any vendor's evidence in a minute.

Three-column sort of AI and automation in infection prevention. Proven: automated electronic surveillance, whole-genome-sequencing outbreak detection, and automated measurement of hand hygiene, none of which is machine learning. Emerging: machine-learning risk prediction for C. difficile and sepsis, which discriminates well but has never been shown to prevent an infection. Hype: claims that AI reduces infections, AI environmental monitoring that automates an ATP proxy, and AI-powered disinfection, where the kill is physics.
Figure 1. What actually works in AI and automation for infection prevention, sorted by what each claim was measured against. Proven means automation and sequencing, not machine learning. Emerging means risk prediction that discriminates well but has never prevented an infection. Hype includes "AI-powered disinfection," where the germicidal action is physics.

What in AI and automation is actually proven?

Start with automated electronic surveillance, which beats manual chart review at what surveillance is for: counting correctly. Sepsis surveillance built on clinical criteria pulled from the electronic record was more than twice as sensitive as surveillance built on administrative claims, 69.7 percent against 32.3 percent, and the two told opposite stories about the same six years: by claims, incidence looked to be climbing 10.3 percent a year; by clinical criteria, it was flat.3 Note the endpoint. That is surveillance accuracy, not an infection reduction. And note the word it does not contain: pulling structured criteria out of a record is automation, not learning.

Europe's PRAISE roadmap, drafted by 30 IPC experts, says the same: manual chart review is "resource intensive and limited by concerns regarding interrater reliability."4 There is a ceiling, though. A systematic review of electronically assisted HAI surveillance found sensitivity generally above 0.8 but specificity ranging from 0.37 to 1.0, and a system at the low end buries your team in false positives. Those reviewers concluded the technology "has yet to reach a mature stage."5 So ask a surveillance vendor for the specificity in a hospital like yours, and what that adds to your team's chart review.

The second proven technology is genomic, and it earns its place by telling you when you are wrong. At one US academic centre, whole-genome sequencing surveillance with machine-learning mining of the record ran alongside conventional practice. Sequencing identified 99 transmission clusters, with a plausible route found for 65.7 percent of them, while traditional investigation over the same period prompted sequencing for 15 suspected outbreaks and confirmed transmission in just five patients, and also "misidentified outbreaks for which transmission did not occur."6 Genomics overturns assumptions too: over three years in Oxfordshire, 45 percent of C. difficile cases proved genetically distinct from every previous case, so they could not have come from a known hospital case.7

Notice where the machine learning sits. Sequencing found the clusters; the model searched the record for what those patients had in common. The intelligence explained the outbreak. The sequencing found it.

The third proven item is the narrowest: automated hand-hygiene monitoring measures better than a clipboard. Events at dispensers visible to a human auditor ran at 3.75 per dispenser per hour; out of that auditor's sight at the same moment, 1.48.8 Your observed compliance figure is, in part, a measurement of the auditor, a fault in the method and not in the people being watched. It establishes measurement validity, not compliance, and certainly not infection.

So look at what is in the proven column. Automated extraction from a record. Genomic sequencing. An electronic counter on a dispenser. Two contain no artificial intelligence at all, and the third contains some, but not in the part that finds the outbreak. The applications in infection control with the strongest evidence behind them are not, in the marketing sense, AI.

Can machine learning predict which patients will get an infection?

It can sort them, and sorting is genuinely difficult. Whether sorting them changes what happens to them is the question nobody has answered.

HAI prediction is where more than half the field's real-world studies sit,1 and the models are not bad: two groups built daily C. difficile risk models at their own hospitals, reaching areas under the curve of 0.82 at the University of Michigan and 0.75 at Massachusetts General,9 and in sepsis, 130 published models reported areas under the curve as high as 0.99.10

Two facts sit underneath those figures. The first, buried in that C. difficile paper: the models were built separately, and "many of the top predictive factors differed between facilities."9 A risk model is partly a description of the disease and partly a description of your hospital, its patient mix, its testing habits. A vendor's "it scored 0.82" stays a fact about somewhere else until the model is tested on your data. The second: of the 28 papers behind those 130 sepsis models, three had clinically implemented one, with mixed results.10

There is an honest exception, and burying it would be exactly the selectivity this article criticises. Two of those C. difficile models were later evaluated prospectively, running live on data they had not seen, and they held: 0.744 to 0.748 at Massachusetts General, 0.778 to 0.767 at Michigan Medicine. The models "were robust to dataset shift," the authors concluded, naming the failure mode in which a deployed model decays as the data around it changes.11 Dataset shift is therefore a risk you monitor for, not a law that models rot. Those models have also never prevented an infection, because nobody has run that trial.

The strongest outcome signal in clinical machine learning is not in infection prevention at all. In a sepsis early-warning system studied prospectively across five hospitals, patients whose alert a clinician confirmed within three hours had lower in-hospital mortality, an adjusted absolute reduction of 3.3 percent (CI 1.7 to 5.1).12 Quote the design alongside it every time: prospective but not randomised, comparing clinicians who confirmed the alert quickly against those who did not. It is the best card in the deck and still not proof, which makes it the reason to demand the trial, not to skip it.

What does "validated" mean when an AI vendor says it?

Almost nothing, until you ask which validation. Internal validation is cross-validation on the model's own data, and it is nearly worthless as evidence that the model travels. External validation tests it on another site's data. Prospective evaluation runs it live, before the outcome is known. Outcome evaluation asks whether patients were better off. When a vendor says "validated," they almost always mean the first.

Here is what the first rung is worth. The Epic Sepsis Model was deployed at hundreds of US hospitals and sold as validated. In 2021 an academic group externally validated it for the first time, across 38,455 hospitalisations: area under the curve 0.63, against the 0.76 to 0.83 reported internally, 67 percent of sepsis patients missed, and alerts on 18 percent of everyone admitted, which is alert fatigue with a number attached.13 Epic has since revised the model. Carry the bigger finding into your next meeting: a model deployed across an entire industry had never been externally validated by anyone, not by the vendor and not by any of the hundreds of hospitals running it, until outsiders did it unbidden.

Nor is that isolated. Only 4 percent of those 100 real-world IPC AI studies carried out independent validation before wider deployment, and only 15 percent were integrated into existing hospital systems.1 And when a large model literature is finally audited, the result can be worse than thin: of 232 prediction models published for COVID-19, every one was rated at high or unclear risk of bias.14

Be precise about what that adds up to, because the cynical reading is wrong too. The case against clinical AI is not that the models are broken. It is that they are unvalidated and unevaluated, which is a different and far more fixable problem. The one infection prevention study that ran the prospective test found the models held.11 Hardly anyone runs the test, and that is the thing to fix, starting inside your own procurement.

A second failure mode would survive all four rungs. An algorithm used to select patients for extra care, running on roughly 200 million people a year, systematically under-referred Black patients: trained to predict healthcare cost as a proxy for need, it learned the spending pattern rather than the illness.15 The mathematics were sound; the target was not. So ask what a model actually predicts, and whether that is what you care about. An HAI model trained on who got tested has learned your testing habits.

Rung four is rare because it is expensive. A cluster-randomised, multicentre, crossover trial of an automated disinfection technology in hospital rooms found that adding UV-C to terminal cleaning lowered acquisition of four multidrug-resistant organisms combined by about 30 percent among patients later admitted to those rooms, a risk ratio of 0.70 (95% CI 0.50 to 0.98), while the same trial found no further reduction in C. difficile infection when UV-C was added on top of an already-sporicidal bleach protocol, a null stratum at a risk ratio of 1.00 (p=0.997).16 Multi-site, randomised, and willing to publish the answer it did not want. The evidence for UV-C disinfection sets out that trial and what it does and does not show. No AI system in infection prevention has produced one.

Does automated hand-hygiene monitoring reduce infection rates?

Not demonstrably, and the useful part of the answer is what comes after that.

These systems are 13 percent of the field's real-world studies,1 so after risk prediction this is the product you will be shown most. A 2024 systematic review of 43 publications found that the systems alone do little: combined "with additional intervention (visual or auditory cue, performance feedback)," they "could increase hand hygiene compliance in the short term," but "impact on infection rates was difficult to determine," and too little is known to recommend routine uptake.17 The causal chain the category rests on is weak before any technology is bolted to it, since Cochrane rates multimodal hand-hygiene interventions as ones that "may slightly improve" compliance and "may slightly reduce infection rates," at low to very low certainty.18 Automating a weak chain does not strengthen it.

So here is what the evidence does tell you to do. The measurement is real: the threefold observer gap8 means a true compliance baseline cannot come from a clipboard. And what moves the number is the cue and the feedback, not the sensing, which infection prevention has known for twenty years. Across 36 acute-care hospitals, only 48 percent of standardised high-touch surfaces (9,910 of 20,646) were adequately cleaned at terminal cleaning; structured feedback raised it to 77 percent.19 The instrument that measured the gap did not close it. The feedback loop closed it, and running that loop took people, protected time and someone who owned it, which is a finding about workload, never a verdict on the people doing the work.

An automated hand-hygiene system, read the same way, is a feedback loop with a sensor attached. Buy the sensor, leave the loop unstaffed, and you have bought a very expensive counter. Automation has always belonged exactly there: the best-practice bundle for environmental surfaces names five components, including compliance monitoring with feedback and no-touch technology as an adjunct to manual cleaning.20 Both slots existed a decade before anyone put "AI" in front of them.

Is "AI-powered disinfection" a real category?

Not in the way the phrase implies, and this deserves bluntness, because we sell a disinfection robot.

Germicidal UV-C works by physics. Ultraviolet light in the germicidal band damages the nucleic acids of an organism it reaches, and the reaction proceeds identically whether the emitter arrived at that spot by following a route a technician taught it or by a learned navigation policy. Photons do not become more germicidal because there is a model upstream of them. So if there is any artificial intelligence inside a disinfection robot, it lives in the navigation, the mapping, the cycle decisions and the record. It is never in the kill. That is true of the category, not of any one manufacturer, and it decides how you should buy: demand the UV-C evidence base, not the AI evidence base. What dose reached the surface, at what distance, for how long, and with what in the way. How far those variables differ between machines is set out in not all UV-C devices are equal.

That applies to our own machine, so we will be precise about ours. The ROZOR Disinfection Robot is an autonomous, mobile UV-C device. Its AI is built to run the cycle: where to go, where to stop, how long to dwell, which surfaces to prioritise, and when a room is done. The AI is also built to analyse that cycle afterwards and report it back, naming the room, when the cycle ran, how long it ran and a measure of coverage, which is the subject of the disinfection audit trail. None of that inactivates an organism. The kill is photons, on our machine and on every other.

Now hold that list against the physics, because dwell time is dose. Dose is irradiance multiplied by exposure time, so an AI built to decide how long to dwell is an AI built to set a dose, which places it inside the chain that ends at a log reduction rather than safely outside it. We are not going to pretend otherwise, and it is why this article keeps pointing you at the output instead of the algorithm: a dose decision is the last thing a buyer should take on a vendor's word, and that includes ours. So what earns any device like ours a place in your bundle is the dose it delivers to the surfaces it reaches, and whether it can show you afterwards what that cycle actually did. Both of those are things the machine's behaviour can be checked against, and both stay checkable whatever sits upstream of them. Ask us for those two things. Ask every vendor for them.

AI-assisted environmental monitoring belongs in the hype column too, and it is subtler. It is 3 percent of the field's real-world studies,1 and what most of these products automate is an ATP reading, which measures organic residue: a proxy for how well a surface was cleaned, never a measure of whether the organisms left on it are alive, a distinction taken apart in visual cleanliness versus disinfection. Automating a proxy does not upgrade the proxy; it delivers the wrong number faster. To know whether a surface is safe, the question is still what dose reached it.

One piece of vocabulary is left, and it is ours. ROZOR describes itself as Physical AI for critical environments, and the phrase means something narrow: a machine that moves through a real room and acts in it, with the intelligence in the route it takes, the time it dwells, and the record it is built to hand back. It does not license the other reading, the one in which the disinfection itself is performed by a model. The disinfecting is still done by photons reaching a surface. That is the reading to hold us to, and to hold every other vendor who reaches for the word.

How do you appraise an AI claim in infection prevention?

The most useful document in this subject is free. That WHO-led scoping review ships a 41-item decision-support checklist across six domains: governance, data quality, technical infrastructure, human and workflow fit, risk and compliance, and economics.1 It was written by infection prevention people for infection prevention people. Read it once and keep it on the desk.

The second tool is a filter you can apply in a single email. TRIPOD+AI is a 27-item reporting standard for clinical prediction models, written to make studies "complete, accurate, and transparent" so that they can be appraised at all, and it covers regression and machine learning alike, puncturing the notion that machine-learning models are exempt from ordinary evidence standards.21 Ask a vendor for their TRIPOD+AI-compliant report. It can still describe a useless model, but it will describe it honestly, which is the whole job of a filter.

Then five questions. Which rung of validation is this, and if the answer is not clear, assume the first. What exactly does the model predict, and is that what you care about?15 What is the specificity in a hospital like yours, and who does the extra chart review?5 Who staffs the feedback loop?17 And what happens when your electronic record changes?11

The frame around all five is older than the technology. WHO's core components for infection prevention treat any single intervention as one part of a multimodal programme, never a substitute for one,22 which is the sentence to hold when a product is presented as a solution rather than a component. Honest automation buys you a better count, a faster genome and a real feedback loop, and those are worth having. What it does not buy is a shortcut around the programme, and any vendor implying otherwise, including one selling a robot, is selling you the adjective. The vendors worth your time will tell you which column their product sits in.

See how the ROZOR Disinfection Robot fits your prevention bundle. The ROZOR Disinfection Robot delivers no-touch UV-C disinfection as an adjunct to your cleaning programme, physical AI for critical environments. Learn more about the ROZOR Disinfection Robot.

Frequently asked questions

Does AI reduce hospital-acquired infections?

No published study has shown an AI system lowering a healthcare-associated infection rate. The WHO-led review of 100 real-world IPC AI studies found only 4 percent had been independently validated before wider deployment, and concluded the field is constrained by limited prospective validation or implementation science. The strongest outcome evidence anywhere in clinical machine learning is a non-randomised sepsis mortality signal, and sepsis is not a healthcare-associated infection.

What AI tools actually work in infection prevention today?

Three, and two of them contain no machine learning. Automated electronic surveillance counts cases far more accurately than administrative claims data, 69.7 percent sensitivity against 32.3 percent. Whole-genome sequencing finds transmission clusters that conventional investigation misses, and clears the ones it wrongly suspected. Automated hand-hygiene monitoring measures compliance without the roughly threefold observer effect a human auditor introduces.

Can machine learning predict which patients will get an infection?

It can discriminate between them. Daily C. difficile risk models reached areas under the curve of 0.82 and 0.75 at two academic centres, and two such models held their discrimination when tested prospectively. But an area under the curve is not a prevented infection, and no infection prevention model has been shown in a trial to lower an infection rate.

What does it mean when a vendor says their AI is "validated"?

Ask which validation. Internal validation on the model's own data is nearly worthless as evidence that it will work in your hospital, while external, prospective and outcome validation get progressively harder and rarer. The Epic Sepsis Model was sold as validated and deployed at hundreds of hospitals, yet its first external validation (area under the curve 0.63, 67 percent of sepsis cases missed, alerts on 18 percent of all patients) was carried out by an academic group, not by the vendor.

Does automated hand-hygiene monitoring reduce infection rates?

Not demonstrably. A 2024 systematic review of 43 publications found these systems can raise compliance in the short term when paired with a cue or performance feedback, that "impact on infection rates was difficult to determine," and that too little is known to recommend routine uptake. Cochrane rates the compliance-to-infection link itself at low to very low certainty. You are buying a feedback loop, so budget for the people who will run it.

Is "AI-powered disinfection" a real category?

The germicidal action of UV-C is physics, not software. Any artificial intelligence in a disinfection robot sits in its navigation, mapping, cycle decisions and record-keeping, never in the inactivation of an organism. That still puts it in the dose path, because dwell time is dose, which is why a disinfection robot is appraised on the UV-C evidence base, meaning the dose delivered to a surface at a stated distance and time, rather than on the AI evidence base.

Sources

  1. Gastaldi S, Tartari E, Satta G, Allegranzi B. "Advancing infection prevention and control through artificial intelligence: a scoping review of applications, barriers, and a decision-support checklist." Antimicrobial Stewardship & Healthcare Epidemiology 2025;5(1):e317. https://pmc.ncbi.nlm.nih.gov/articles/PMC12722576/
  2. Plana D, Shung DL, Grimshaw AA, Saraf A, Sung JJY, Kann BH. "Randomized Clinical Trials of Machine Learning Interventions in Health Care: A Systematic Review." JAMA Network Open 2022;5(9):e2233946. https://pmc.ncbi.nlm.nih.gov/articles/PMC9523495/
  3. Rhee C, Dantes R, Epstein L, et al. "Incidence and Trends of Sepsis in US Hospitals Using Clinical vs Claims Data, 2009-2014." JAMA 2017;318(13):1241-1249. https://pubmed.ncbi.nlm.nih.gov/28903154/
  4. van Mourik MSM, van Rooden SM, Abbas M, et al. "PRAISE: providing a roadmap for automated infection surveillance in Europe." Clinical Microbiology and Infection 2021;27(Suppl 1):S3-S19. https://pubmed.ncbi.nlm.nih.gov/34217466/
  5. Streefkerk HRA, Verkooijen RPAJ, Bramer WM, Verbrugh HA. "Electronically assisted surveillance systems of healthcare-associated infections: a systematic review." Eurosurveillance 2020;25(2):1900321. https://pubmed.ncbi.nlm.nih.gov/31964462/
  6. Sundermann AJ, Chen J, Kumar P, et al. "Whole-Genome Sequencing Surveillance and Machine Learning of the Electronic Health Record for Enhanced Healthcare Outbreak Detection." Clinical Infectious Diseases 2022;75(3):476-482. https://pubmed.ncbi.nlm.nih.gov/34791136/
  7. Eyre DW, Cule ML, Wilson DJ, et al. "Diverse sources of C. difficile infection identified on whole-genome sequencing." New England Journal of Medicine 2013;369(13):1195-1205. https://pubmed.ncbi.nlm.nih.gov/24066741/
  8. Srigley JA, Furness CD, Baker GR, Gardam M. "Quantification of the Hawthorne effect in hand hygiene compliance monitoring using an electronic monitoring system: a retrospective cohort study." BMJ Quality & Safety 2014;23(12):974-980. https://pubmed.ncbi.nlm.nih.gov/25002555/
  9. Oh J, Makar M, Fusco C, et al. "A Generalizable, Data-Driven Approach to Predict Daily Risk of Clostridium difficile Infection at Two Large Academic Health Centers." Infection Control & Hospital Epidemiology 2018;39(4):425-433. https://pmc.ncbi.nlm.nih.gov/articles/PMC6421072/
  10. Fleuren LM, Klausch TLT, Zwager CL, et al. "Machine learning for the prediction of sepsis: a systematic review and meta-analysis of diagnostic test accuracy." Intensive Care Medicine 2020;46(3):383-400. https://pubmed.ncbi.nlm.nih.gov/31965266/
  11. Kamineni M, Otles E, Oh J, et al. "Prospective evaluation of data-driven models to predict daily risk of Clostridioides difficile infection at two large academic health centers." Infection Control & Hospital Epidemiology 2023;44(7):1163-1166. https://pmc.ncbi.nlm.nih.gov/articles/PMC10024639/
  12. Adams R, Henry KE, Sridharan A, et al. "Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis." Nature Medicine 2022;28(7):1455-1460. https://pubmed.ncbi.nlm.nih.gov/35864252/
  13. Wong A, Otles E, Donnelly JP, et al. "External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients." JAMA Internal Medicine 2021;181(8):1065-1070. https://pubmed.ncbi.nlm.nih.gov/34152373/
  14. Wynants L, Van Calster B, Collins GS, et al. "Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal." BMJ 2020;369:m1328. https://pubmed.ncbi.nlm.nih.gov/32265220/
  15. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. "Dissecting racial bias in an algorithm used to manage the health of populations." Science 2019;366(6464):447-453. https://pubmed.ncbi.nlm.nih.gov/31649194/
  16. Anderson DJ, Chen LF, Weber DJ, et al. "Enhanced terminal room disinfection and acquisition and infection caused by multidrug-resistant organisms and Clostridium difficile (the Benefits of Enhanced Terminal Room Disinfection study): a cluster-randomised, multicentre, crossover study." The Lancet 2017;389(10071):805-814. https://pubmed.ncbi.nlm.nih.gov/28104287/
  17. Gould D, Hawker C, Drey N, Purssell E. "Should automated electronic hand-hygiene monitoring systems be implemented in routine patient care? Systematic review and appraisal with Medical Research Council Framework for Complex Interventions." Journal of Hospital Infection 2024;147:180-187. https://pubmed.ncbi.nlm.nih.gov/38554805/
  18. Gould DJ, Moralejo D, Drey N, Chudleigh JH, Taljaard M. "Interventions to improve hand hygiene compliance in patient care." Cochrane Database of Systematic Reviews 2017;9(9):CD005186. https://pubmed.ncbi.nlm.nih.gov/28862335/
  19. Carling PC, Parry MF, Rupp ME, Po JL, Dick B, Von Beheren S. "Improving cleaning of the environment surrounding patients in 36 acute care hospitals." Infection Control & Hospital Epidemiology 2008;29(11):1035-1041. https://pubmed.ncbi.nlm.nih.gov/18851687/
  20. Rutala WA, Weber DJ. "Best practices for disinfection of noncritical environmental surfaces and equipment in health care facilities: A bundle approach." American Journal of Infection Control 2019;47(Suppl):A96-A105. https://doi.org/10.1016/j.ajic.2019.01.014
  21. Collins GS, Moons KGM, Dhiman P, et al. "TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods." BMJ 2024;385:e078378. https://pubmed.ncbi.nlm.nih.gov/38626948/
  22. World Health Organization. "Guidelines on core components of infection prevention and control programmes at the national and acute health care facility level." Geneva: WHO; 2016. https://www.who.int/publications/i/item/9789241549929