"We've Never Disagreed" Is Not Evidence of Agreement (Why We Test Our Own Infection Preventionists Annually)

What Happens When Two Infection Preventionists Read the Same Chart

Every year, our IP&MA infection preventionists sit down and take a test they can fail.

We assign all members of our team approximately thirty case scenarios, drawn from a validated test bank, that require them to apply NHSN surveillance definitions to exactly the kind of real-world, ambiguous chart that is likely to be encountered in a patient's medical record. It is not easy, and it is not designed to be. People do fail it. Not often, and not always the people you'd expect. Year of IP experience protect you less than you'd think when a definition changed last January and you've been applying the old one from memory.

Most infection prevention programs don't do this. Not because they don't care about accuracy, but because nothing requires it, and nothing in the reporting process would reveal the problem if it existed.

‍ ‍

Why does agreement matter as much as accuracy?

Healthcare-associated infection data is not observed, it is interpreted based on published, standard definitions. Every number a hospital reports is the product of a person applying a complex, frequently updated definition to an incomplete medical record. Two experienced infection preventionists can read the same chart, apply the same definition in good faith, and reach different conclusions.

The disagreements are rarely about obvious cases. They cluster around the edges: whether a device was in place long enough, whether a symptom was present on admission or developed after, whether documentation four days apart describes one infection or two. Reasonable people read the same notes and land in different places.

When that happens, nothing catches it. The event is entered, submitted to NHSN, and travels onward into public reporting, quality benchmarking, and CMS payment determinations. There is no step in that chain where a second reviewer asks whether the first one was right. And at the front of that chain is a patient whose infection was either counted or wasn't.

If two IPs looking at the same chart can't agree on whether a case meets criteria, that is worth discovering internally … before it becomes a reported number.

‍ ‍

Confidence and consistency are different things

APIC's 2020 MegaSurvey found that 60% of infection preventionists rate themselves as proficient or expert in surveillance. Meanwhile, the largest study of NHSN surveillance accuracy found that participants miscoded more than a third of test cases. Both findings can be true simultaneously, and that is the entire point. Surveillance is consistently reported as the most time-consuming task in the role. People who do it constantly become genuinely skilled at it.

"We've never had a disagreement" is not evidence of agreement. In most programs, it means no two people have ever reviewed the same case.

‍ ‍

What our own test taught us about testing

When our team developed these case scenarios, we assumed the hard part would be writing questions difficult enough to be meaningful. It wasn't. When we applied item difficulty and item discrimination analysis to the scenarios, we found that case studies written by expert infection preventionists did not always elicit the response we intended from IPs with varying levels of experience. Some questions were ambiguous in ways their authors could not see. The test bank itself had to be validated before it could validate anyone else.

That finding has stayed with us. If experienced IPs writing deliberately clear scenarios can produce ambiguity without noticing, the charts we review every day are at least as open to interpretation as we think they are — probably more. We've published multiple papers and abstracts on this work: including building the test bank, on interrater reliability across experience levels, and on how coding accuracy changes with structured training over time. Measuring our own surveillance turned out to be a research question, not just an internal exercise.

‍ ‍

What annual testing surfaces

Results tell us where definitions are being applied inconsistently, which recent definitional changes didn't fully land, and where newer infection preventionists need support rather than correction. In our longitudinal analysis of structured surveillance training, we identified that Structured, repeated surveillance training and testing were associated with improved HAI coding accuracy. That pattern is the argument for testing annually rather than once at orientation. It also produces something increasingly useful during survey: documented, objective evidence of surveillance competency, rather than an attestation that training occurred.

‍ ‍

Trying a smaller version

You don't need a full, validated test bank to start. Select ten of your own recent cases. Have two infection preventionists review them independently, without discussion, and record their determinations. Then compare. The conversation that follows the disagreements is usually more valuable than the agreement rate itself.

We keep doing this because it is the way we know to be sure. Our IPs take a test every year that has no personal upside — a passing score changes nothing about their day, and a hard question just means an uncomfortable afternoon. They show up for it anyway and commit to closing knowledge gaps that the testing reveals. That willingness to be measured and commitment to improving skills is what separates a surveillance program you can trust from one you only hope is right.

‍ ‍

If you'd like the formal version, IP&MA's Infection Surveillance Interrater Reliability Assessment uses validated case scenarios to measure agreement across your team and identify where support would help most. Request a discovery call to talk through what that would look like in your program.

Next
Next

Construction, Water & Air: The Environment Is Not Neutral