Skip to main content
Category: Data Quality

Data Matching

Also known as: Record Linkage, Entity Resolution, Data Linkage, Object Identification, Field Matching, Exact Data Matching (EDM)
Simply put

Data matching is the process of comparing two or more sets of records to identify those that refer to the same real-world entity, such as the same person or organization, whether within a single dataset or across separate sources. It helps organizations link, deduplicate, and connect related information that may be stored inconsistently. This entry describes the technique itself and does not address the lawful bases, safeguards, or restrictions that may apply when the data being matched relates to identified or identifiable individuals.

Formal definition

Data matching, also referred to as record linkage or entity resolution, is the task of identifying and correlating records across one or more datasets that correspond to the same underlying entity, using techniques that range from exact (deterministic) matching of data elements to probabilistic or similarity-based comparison of patterns and attributes. As a data governance and data quality capability, it supports deduplication, data integration, and the establishment of a consolidated view of an entity, and it typically depends on and informs related governance activities such as lineage, stewardship, and catalog management. Where matching operates on personal data, it constitutes processing that may attract data protection obligations; matching, correlating, or linking records does not by itself render data non-personal, and pseudonymized identifiers used to enable linkage generally remain personal data. This definition covers the matching technique and its governance context only; it does not address specific lawful bases, cross-border transfer mechanics, retention rules, or whether a data protection impact assessment is required, all of which depend on jurisdiction, purpose, and implementation.

Why it matters

Data matching underpins many core data governance and data quality objectives, including deduplication, data integration, and the creation of a consolidated view of an entity across systems that store information inconsistently. When records referring to the same person or organization are scattered across separate sources or duplicated within a single dataset, the resulting fragmentation undermines reporting accuracy, operational efficiency, and the reliability of downstream decisions. Effective matching helps organizations link related information and establish a trustworthy single view, which in turn supports stewardship, lineage, and catalog activities that depend on knowing which records correspond to which real-world entities.

Where data matching operates on personal data, it constitutes processing and may attract data protection obligations. A common expert-level error is to assume that correlating or linking records somehow diminishes the sensitivity of the data; it does not. Matching, correlating, or linking records does not by itself render data non-personal, and pseudonymized identifiers used to enable linkage generally remain personal data. Organizations that treat matching as a purely technical data quality exercise, detached from privacy accountability, risk overlooking that the act of linking may itself require a lawful basis and appropriate safeguards.

The governance implications extend beyond the technique itself. Because matching can consolidate previously separate information about an individual, it can increase the identifiability and richness of a profile, which is a governance and privacy consideration that must be assessed in context. This entry describes the technique and its governance framing only; whether a data protection impact assessment is required, what lawful basis applies, and how retention and cross-border transfer rules affect a given implementation all depend on jurisdiction, purpose, and design, and are out of scope here.

Who it's relevant to

Data governance and data quality leads
Those responsible for deduplication, data integration, and establishing a consolidated view of an entity rely on matching as a core capability. They must coordinate it with lineage, stewardship, and catalog management, and ensure that the accountability for match outcomes is demonstrable rather than merely asserted.
Data protection officers and privacy professionals
Where matching operates on personal data, it constitutes processing that may attract data protection obligations. These practitioners need to assess the privacy implications of linking records, recognizing that matching does not render data non-personal and that pseudonymized identifiers used to enable linkage generally remain personal data. Whether a lawful basis or an impact assessment is required depends on jurisdiction, purpose, and implementation.
Privacy engineers and data integration teams
Those implementing deterministic, probabilistic, or similarity-based matching must understand how technique choice affects both data quality outcomes and the identifiability of consolidated records, and should design in traceability that supports governance and privacy review.
Information governance and stewardship functions
Because matching consolidates and connects related information, stewards need to manage the resulting single view of an entity, maintain lineage over linked records, and ensure the catalog reflects how entities are resolved across sources.

Inside Data Matching

Record Linkage
The core process of comparing records across or within datasets to identify those that refer to the same individual or entity. Matching may be deterministic, relying on exact agreement of key identifiers, or probabilistic, assigning a likelihood score based on partial or fuzzy agreement across multiple attributes.
Matching Attributes
The fields used to establish a match, such as name, date of birth, address, or reference numbers. Where these attributes identify or can be reasonably linked to a living individual, they generally constitute personal data, and their use in matching is itself a processing activity subject to applicable data protection law.
Lawful Basis for Processing
Under the EU GDPR and UK GDPR, data matching as a processing activity requires an identified lawful basis. Consent is only one option among several; legitimate interests, legal obligation, or public task may apply depending on context. The appropriate basis depends on the purpose and the parties involved, and the position may differ under regimes such as the CCPA and CPRA.
Purpose and Compatibility
Matching that combines data collected for separate purposes raises questions of purpose limitation. In most jurisdictions the new use should be assessed for compatibility with the original collection purpose; this entry does not cover the detailed criteria for that assessment.
Governance Controls
Data governance elements relevant to matching include documented data lineage, stewardship over matching rules, and quality controls addressing false positives and false negatives. These are distinct from, though often paired with, the information security controls that protect the datasets during matching.
Risk and Assessment
Where matching is likely to result in high risk to individuals, a data protection impact assessment may be warranted under the EU GDPR or UK GDPR. Such an assessment is not automatically mandatory for every matching activity; the need depends on the nature, scope, context, and risk of the processing.

Common questions

Answers to the questions practitioners most commonly ask about Data Matching.

Does data matching produce anonymized data because it only links records rather than exposing raw identifiers?
No. Data matching links or reconciles records that relate to the same individual, and the linked output generally remains personal data. Because matching typically depends on identifiers or quasi-identifiers, the process itself does not sever the connection to a data subject. Anonymization requires irreversibility and would place the result outside the scope of most data protection regimes, whereas matched or linked datasets remain reversible and identifiable. This entry does not cover the specific technical thresholds for assessing whether a dataset has been effectively anonymized.
Is data matching purely a data quality or governance activity, separate from data protection obligations?
Not entirely. Data matching sits at the overlap between data governance and data protection. From a governance perspective it supports data quality, deduplication, and lineage, and these are governance concerns around ownership, stewardship, and policy. However, where matching involves personal data, it is also a processing activity and generally attracts data protection obligations, including the need for a lawful basis and accountability. Treating it as only a governance or data quality exercise risks overlooking those obligations. This entry does not detail specific lawful bases or retention rules that may apply.
Which party bears responsibility for data matching when a third party performs it on our behalf?
Accountability generally follows the roles defined under applicable law. The party that determines the purposes and means of the matching typically acts as the data controller and bears primary accountability, while a party matching data strictly on documented instructions typically acts as a data processor. Where two parties jointly determine purposes and means, a joint controller relationship may arise. These roles should be assessed against the facts of the arrangement rather than assumed from contract labels. This entry does not address the specific contractual terms that different regimes require between such parties.
Do we always need a data protection impact assessment before carrying out data matching?
Not always. A data protection impact assessment is not automatically mandatory for every matching activity. Under regimes such as the EU GDPR and UK GDPR, an assessment is generally expected where processing is likely to result in a high risk to individuals, and large-scale matching, profiling, or combining datasets from different sources can be indicators of such risk. Whether one is required depends on the context, scale, and nature of the matching. This entry does not set out the specific criteria each supervisory authority uses to determine when an assessment is required.
How should we record data matching in our records of processing activities?
Matching is generally a distinct processing operation and, where it involves personal data, would typically be reflected in records of processing activities where such records are required. It is important to remember that a records of processing activities obligation is a documentary requirement to describe processing, not the same thing as a data inventory or cataloguing tool, even though such tools may help populate those records. The records should reflect the actual matching activity, its purposes, and the categories of data and individuals involved. This entry does not prescribe the exact fields any specific regime requires.
Does applying encryption or tokenization to the datasets before matching remove them from data protection scope?
Generally no. Encryption and tokenization are security and pseudonymization measures that can reduce risk and support integrity and confidentiality, but they do not by themselves make data non-personal. Pseudonymized data that can be re-linked to individuals, including through matching keys, typically remains personal data. These controls address the security dimension of confidentiality and integrity, while the governance and lawful-basis questions around the matching itself remain. This entry does not cover the specific technical standards for implementing encryption or tokenization.

Common misconceptions

Data matching only involves personal data if it produces a confirmed match.
The comparison of identifying attributes is itself processing of personal data, regardless of whether a match is ultimately confirmed. The matching operation, including any intermediate scoring or attempted linkage, generally falls within the scope of applicable data protection law.
Consent is always required before matching records.
Consent is one lawful basis among several under the EU GDPR and UK GDPR. Depending on the purpose and the party carrying out the matching, legitimate interests, legal obligation, or public task may be more appropriate. Selecting a basis requires assessment of context rather than defaulting to consent.
Matching on pseudonymized or tokenized identifiers removes it from data protection scope.
Pseudonymization is reversible and pseudonymized data generally remains personal data. Matching that uses tokenized or pseudonymized keys does not make the activity fall outside data protection law, because the individuals can typically still be distinguished or re-identified.

Best practices

Identify and document a specific lawful basis for the matching activity before it begins, and record the reasoning where multiple bases could apply, recognizing that no single basis guarantees compliance.
Assess whether the matching purpose is compatible with the purposes for which the source data was originally collected, and document that assessment.
Determine whether the matching is likely to result in high risk to individuals and, where so, conduct a data protection impact assessment rather than assuming one is either always or never required.
Maintain documented data lineage and stewardship over matching rules, treating governance evidence as demonstrable accountability rather than stated intent.
Establish quality controls that address both false positives and false negatives, since matching errors carry direct consequences for the individuals concerned.
Apply appropriate security controls to the datasets during matching while keeping them distinct from the governance controls, and do not treat pseudonymization or tokenization as removing the data from protection scope.