Can ChatGPT Perform HAI Surveillance? A New Study Offers Important Insights
Can ChatGPT Perform HAI Surveillance?
Not reliably — not yet.
In a study our team published in the American Journal of Infection Control, ChatGPT-4 correctly applied NHSN surveillance definitions to validated case scenarios 45% of the time. Across three identical runs of the same 20 scenarios, only 20% were answered correctly every time. Trained infection preventionists significantly outperformed the model on the same cases.
What the study did
Kelly Holmes, Mishga Moinuddin, and Sandi Steinfeld, with Xiaoyu Liu, ran a simulation-based evaluation using 20 de-identified surveillance vignettes — the same validated scenarios we use to build and test infection preventionist competency.
The AI took the same test our infection preventionists take. Each scenario was run three times under identical conditions, which is what surfaced the finding that matters most.
Accuracy was the smaller problem
ChatGPT-4 performed best on laboratory-identified events, where the determination follows fairly directly from a positive result and a date. It struggled with cases requiring temporal reasoning, interpretation of multiple criteria, and the layered logic that defines most real surveillance work — the questions where two experienced IPs might themselves disagree.
That pattern is not surprising, and by itself it might just mean the tool needs a better version.
Inconsistency is the real problem
The same scenario, run three times, produced the same correct answer only 20% of the time.
Consider what that means operationally. A surveillance method that returns different determinations for identical input cannot be validated. It cannot be audited. It cannot be reproduced when CMS selects your facility for HAI Validation and asks you to defend a case determination from fourteen months ago.
We test our own infection preventionists for interrater reliability precisely because consistency is a prerequisite for trustworthy data — not a bonus on top of accuracy. A tool that is right on average but unpredictable case-by-case fails that standard regardless of its average.
What about newer models?
A fair question, and the honest answer has two parts.
Newer models will almost certainly score higher on accuracy. That trajectory is real and it will likely continue. But the consistency finding is a property of how these systems work rather than a limitation of one version. Language models generate probabilistically; identical prompts can produce different outputs by design. That behavior is useful for drafting and unhelpful for a determination that has to be reproducible two years later.
The standard for a surveillance tool was never "better than the previous version." It should be accurate, validated, reproducible, and defensible.
Where AI can be useful in infection prevention right now
None of this means these tools have no place. Used carefully, they save real time:
Drafting and revising policies, competencies, and education materials for expert review
Summarizing lengthy guidance documents to orient yourself before reading the source
Producing first drafts of committee reports, staff communication, and outbreak notifications
Reformatting and restructuring data you've already validated
Literature searching as a starting point, with every citation verified independently
The common thread: these are tasks where a knowledgeable person reviews the output and remains accountable for it. That is a different category from case determination, where the output is the decision and becomes a reported number.
If someone is selling you an AI surveillance tool
Worth asking, before anything else: has it been validated against a known case set with published results? Does identical input produce identical output every time? Can it show its reasoning against specific NHSN criteria? And who is accountable for the determination — the vendor, or you?
Where we can help
IP&MA supports organizations on surveillance accuracy, competency validation, and NHSN data integrity — including evaluating where emerging tools fit into a surveillance program without compromising the reliability your reported data depends on.