hello@thompsoncode.com

Quality Data Matters

Data is everywhere. But the data that helps and organization to win, is high quality. And it doesn't happen by accident. It is an intentional act.


Data Standards

Organizations that win with data have data standards. These standard set the tone for the organization. It ensures that your organzaiton follows a consistent set of standards for data entry, data validation, data usage and so forth. Standards should be written and very specific with examples of both good and bad data and should be part of the on-going traing that everyone that works with data needs to understand.

An example of detailed data standard is an email field should be email@domain.com or blank. So in your QC process if the is a NA in an email field it violates the standard.

Without data standards data can break production systems, cause poor decisions form leadership and much more. Poor data costs your organizaiton. Do it right and have a data standard.

Data Quality Defined

There are multiple attributes that define data quality:

  • Complete
  • Consistent
  • Valid
  • Integrity
  • Timely
  • Reasonable
  • Unique
  • Accurate

Let's look at each of these in more detail.

Complete

Quality data is complete. It doesn't have missing values, unless the value should missing. An example of a field that should be missing would be a date of death for someone still living.

  • An example of a value that should be there is a date of birth. It should always exist. If it doesn't there is a problem with data entry or validation.
  • An example of value that might be missing that is acceptible is a date of death for someone still living. The value doesn't exist.

If we encounter a data set with missing fields, we follow a process called data cleaning to get it up to minimum standards. Also good quality control for data entry and data validation is very important. I have seen organizations that don't have these continually struggle with data.

Consistent

Quality data is consistent. It says the same thing consistently across rows and fields as the case may be. So you don't have multiple rows of dates that look like this:

  • Janurary 6, 2026
  • 1/6/2026
  • Jan. 6, 2026

This would require lost of cleaning and should be caught be the QC process if you have data standards.

Valid

Quality data is within approprate range. An easy example if a person's age is 0 or more. It cannot be nagative, a letter, etc. If it is you have a problem.

Integrity

Quality data has correct relationships. So if you have a database of students at a schoo and a student has no parents, it lacks integrity and there is data loss somewhere.

Timely

Quality data is timely. This means data gets to the person it needs to be used in a timely manner so it can actually be used. If department A updates data that then is made available for department B it needs to arrive at department B with enough time for them to do something with it. We ensure this with a mechanism called a data pipleline.

Reasonable

Quality data is reasonable. It lives with in acceptible ranges. For exmple, go back to the age example.

  • 10, 50 or even 100 are acceptible.
  • 10 trillion is not.

Unique

Quality data is unique. This means you don't have exact duplicates. There are some secenarios when a duplicate in some fields if fine. Such a case may be a customer makes mutiple orders. While many of the fields are duplicate (name, address, etc.) the order number is unique. If too many fields are duplicate, especially identifier fields (order number) you may have system issues and the data needs to be cleaned.

Accurate

Quality data is accurate. If the value is actually 5 you don't have 500. That is a data entry and validation issue. You may have cases where the data is rounded but that should be defined in data standards or the specific needs of your project.

Summary

In summary, if you want to win with data, your oragnizaiton needs to adapt data standards follow good data practices to ensure you have high quality data.