Data Matching
Data matching is the process of comparing two or more sets of records to identify those that refer to the same real-world entity, such as the same person or organization, whether within a single dataset or across separate sources. It helps organizations link, deduplicate, and connect related information that may be stored inconsistently. This entry describes the technique itself and does not address the lawful bases, safeguards, or restrictions that may apply when the data being matched relates to identified or identifiable individuals.
Data matching, also referred to as record linkage or entity resolution, is the task of identifying and correlating records across one or more datasets that correspond to the same underlying entity, using techniques that range from exact (deterministic) matching of data elements to probabilistic or similarity-based comparison of patterns and attributes. As a data governance and data quality capability, it supports deduplication, data integration, and the establishment of a consolidated view of an entity, and it typically depends on and informs related governance activities such as lineage, stewardship, and catalog management. Where matching operates on personal data, it constitutes processing that may attract data protection obligations; matching, correlating, or linking records does not by itself render data non-personal, and pseudonymized identifiers used to enable linkage generally remain personal data. This definition covers the matching technique and its governance context only; it does not address specific lawful bases, cross-border transfer mechanics, retention rules, or whether a data protection impact assessment is required, all of which depend on jurisdiction, purpose, and implementation.
Why it matters
Data matching underpins many core data governance and data quality objectives, including deduplication, data integration, and the creation of a consolidated view of an entity across systems that store information inconsistently. When records referring to the same person or organization are scattered across separate sources or duplicated within a single dataset, the resulting fragmentation undermines reporting accuracy, operational efficiency, and the reliability of downstream decisions. Effective matching helps organizations link related information and establish a trustworthy single view, which in turn supports stewardship, lineage, and catalog activities that depend on knowing which records correspond to which real-world entities.
Where data matching operates on personal data, it constitutes processing and may attract data protection obligations. A common expert-level error is to assume that correlating or linking records somehow diminishes the sensitivity of the data; it does not. Matching, correlating, or linking records does not by itself render data non-personal, and pseudonymized identifiers used to enable linkage generally remain personal data. Organizations that treat matching as a purely technical data quality exercise, detached from privacy accountability, risk overlooking that the act of linking may itself require a lawful basis and appropriate safeguards.
The governance implications extend beyond the technique itself. Because matching can consolidate previously separate information about an individual, it can increase the identifiability and richness of a profile, which is a governance and privacy consideration that must be assessed in context. This entry describes the technique and its governance framing only; whether a data protection impact assessment is required, what lawful basis applies, and how retention and cross-border transfer rules affect a given implementation all depend on jurisdiction, purpose, and design, and are out of scope here.
Who it's relevant to
Inside Data Matching
Common questions
Answers to the questions practitioners most commonly ask about Data Matching.