Picture a city that spent ten years laying flawless pipes. Every home, clinic, school is plumbed, and the second you turn a tap, water comes rushing out. What’s missing in this picture? The treatment plant that ensures the water is filtered and safe. So whatever goes into the system is exactly what comes out of every faucet, contaminants and all.
You can see where this discussion is heading when discussing healthcare data. The pipes are largely built. Data moves efficiently between hospitals, labs, payers, and the app on your phone faster than ever, and interoperability, the plumbing under all of it, is well established.
But we’re splashing in uncertainty. Can we trust the data running through the infrastructure?
Connection Was Never the Finish Line
It’s not as if the great builders of the interoperability system considered their job complete when once the floodgates opened and the data was free to flow from place to place. That we are at this place in our journey, examining the quality, provenance and usefulness of the data is not surprising. It’s exactly where we should be.
There are four core domains of interoperableexchange: sending, receiving, finding, and integrating electronic health information. According to ONC's 2023 data brief, only about 43% of U.S.hospitals routinely engage in all.
So why such an alarmingly low percentage? It has to do with the deluge of information. All the data pours out of the same tap. Clean and contaminated, actionable and junk, everything arrives together, and the clinician is left with a bucketfull of meaningful data swimming around with contaminates. This makes the data itself untrustworthy and therefore, inactionable.
What “Dirty Data” Actually Means

In a health record,contamination usually shows up as one of these:
● Duplicates: One patient living in the system as two or more slightly different people.
● Missing pieces: a lab value with no unit attached, which is a riddle, not a result.
● Inconsistency: the same fact written three different ways across three systems.
● Entry errors: a stray “MI” that could mean myocardial iknfarction or the state of Michigan.
None of this is exotic. Research compiled by The Pew Charitable Trusts found that providers consider a match rate of 99 percent or higher the goal, yet the reality sits far below it: a Black Book survey cited in that work put the average duplicate rate inside a single organization at roughly 18 percent, and match rates between organizations can fall to 50 percent or lower.
Now connect a system like that into a national network, and every duplicate, blank, and typo travels to every partner who touches the EHR. The signal moves beautifully. It is just carrying nonsense.
So How Do We Clean It?

Obviously, clean data is the goal. We need a way to measure it by an industry accepted standard. Handily, one already exists. Data quality breaks down into six measurable dimensions determining what data is worth sharing.
● Accuracy: is the data actually correct?
● Completeness: is anything missing?
● Consistency: does the same fact match everywhere it lives?
● Timeliness: is it current, or are you treating last year's patient?
● Validity: is it in a format the receiving system can read?
● Uniqueness: one patient, one record.
How AI Adds Urgency to the Issue
An AI model drinks whatever is in the supply. Currently, AI can only amplify poor data quality, it cannot yet correct it. AI moves so quickly and with such reach that a single data quality issue that might have affected one decision can influence thousands of AI-assisted decisions. This increases risk. These records inform providers’ decisions. Let’s never forget that at the end of every line of data, there lives a patient.
The Rules Are Moving This Way Too
None of this is lost on the people writing the regulations. If you want a truly interesting read on regulated, standardized language, click on our May post, “Get Your Head in the Game: Understanding the HTI-5 Rulebook.” Data must share a common language to make it genuinely usable and not just technically shared.
Clean It at the Source
The pipes were the hard engineering concern of the last decade, and we did our job well. We all trust the connection points. The data flows easily and unencumbered thanks to the undergirding of interoperability systems.
So let’s now improve what flows through them. We must filter out incomplete, inconsistent, or inaccurate data by adopting standardized reference checks that ensure the data is accurate, complete, consistent, timely, valid and unique.

