Skip to main content
Category: Data Quality

Data Profiling

Also known as: Data Archeology
Simply put

Data profiling is the process of examining an organization's data to understand how it is structured, stored, and interconnected, and to assess its quality. It helps organizations identify issues such as incompleteness or inaccuracy so the data can be cleansed and used more reliably. It is generally a data governance and data quality activity rather than a security or privacy control.

Formal definition

Data profiling is the systematic examination of data from one or more sources to characterize its structure, content, and quality, typically measuring attributes such as completeness and accuracy and surfacing interconnections across datasets. It supports data quality management, cleansing, and informed data decisions, and is commonly used to prepare or summarize data for downstream use. Within a governance context it contributes to understanding data content and quality, but this definition does not address lawful basis, personal data handling obligations, retention, cross-border transfer, or security controls, which fall outside its scope. Note that data profiling in this data-quality sense is distinct from 'profiling' as defined under data protection regimes such as the EU GDPR or UK GDPR, which concerns automated processing to evaluate personal aspects of individuals; the evidence here supports only the data-quality meaning.

Why it matters

Data profiling matters because the reliability of nearly every downstream data activity depends on knowing what an organization actually holds, how it is structured, and where quality problems exist. Reviewing and cleansing data to understand its structure and maintain quality standards allows organizations to surface issues such as incompleteness or inaccuracy before that data is used for analytics, reporting, migration, or operational decisions. Without this understanding, decisions may be made on data that is inconsistent or unfit for purpose.

Within a data governance context, profiling contributes to a broader understanding of data content and quality, which supports stewardship, cataloguing, and informed data decisions. Because it characterizes existing information and the interconnections across datasets, it can help data teams make better, more defensible choices about how data should be handled and prepared for downstream use.

It is important not to overstate what data profiling in this data-quality sense achieves. This activity does not address lawful basis, personal data handling obligations, retention, cross-border transfer, or security controls, all of which fall outside its scope. It should also not be confused with 'profiling' as defined under data protection regimes such as the EU GDPR or UK GDPR, which concerns automated processing to evaluate personal aspects of individuals; the two are distinct concepts that happen to share a word.

Who it's relevant to

Data Governance and Stewardship Leads
Those responsible for data ownership, stewardship, cataloguing, and data quality use profiling to understand how data is structured and interconnected and to identify quality issues. It contributes to demonstrable understanding of data content and quality, though it does not, on its own, satisfy accountability obligations that require evidence beyond stated intent.
Data Quality and Data Management Teams
Teams focused on measuring completeness and accuracy and on cleansing data rely on profiling to surface incompleteness or inaccuracy so that data can be filtered, reconciled, and prepared for reliable downstream use.
Analytics and Data Engineering Practitioners
Those preparing or summarizing data for analytics, migration, or operational use benefit from profiling because it characterizes structure and quality across multiple sources before that data is consumed, supporting better, more informed data decisions.
Privacy and Compliance Professionals (with a caution)
Privacy and compliance practitioners should note that data profiling in this data-quality sense is distinct from 'profiling' under the EU GDPR or UK GDPR, which concerns automated processing to evaluate personal aspects of individuals. This activity does not address lawful basis, personal data handling, retention, cross-border transfer, or security controls, so it should not be relied upon as a privacy or security measure.

Inside Data Profiling

Structure Discovery
Examination of data to assess whether it is consistent and correctly formatted, including analysis of formats, patterns, and data types across a column or dataset. This is a data quality and governance activity, not a security control.
Content Discovery
Inspection of individual data values to identify errors, null values, ambiguities, and anomalies within records. In a data protection context, this may surface the presence of personal data or special category data, though profiling alone does not classify or assign a lawful basis for that data.
Relationship Discovery
Identification of how data elements relate across tables, sources, and systems, including keys, dependencies, and overlaps. This supports data lineage and cataloging efforts within a governance program.
Statistical Analysis
Generation of summary metrics such as counts, distributions, minimum and maximum values, uniqueness, and frequency to characterize a dataset and quantify quality issues.
Quality Metrics
Measures such as completeness, uniqueness, validity, and consistency that describe the fitness of data for a given purpose. These metrics inform stewardship decisions but do not by themselves establish regulatory compliance.
Distinction from GDPR Profiling
Data profiling as a data quality practice is distinct from 'profiling' as defined under the EU GDPR and UK GDPR, which refers to automated processing to evaluate personal aspects of an individual. The two terms share a word but address different concerns; treatment of automated profiling of individuals differs by jurisdiction and is out of scope for this entry.

Common questions

Answers to the questions practitioners most commonly ask about Data Profiling.

Is data profiling in the governance sense the same as profiling under the EU GDPR?
No, and conflating them is a common mistake. In a data governance context, data profiling generally refers to examining data to assess its structure, content, quality, completeness, and lineage, typically to support data quality management and cataloging. Under the EU GDPR, profiling is a defined term referring to automated processing of personal data to evaluate, analyze, or predict aspects of a natural person. These are distinct concepts: the governance activity is a technical quality exercise, while the regulatory term carries specific obligations tied to automated processing of individuals. Treatment of the regulatory concept also differs across regimes such as the UK GDPR, the CCPA and CPRA, and others, so scope your usage to the instrument in question.
Does running data profiling on a dataset satisfy a records of processing activities obligation or serve as a data inventory?
No. Data profiling assesses the quality and characteristics of data itself, whereas a records of processing activities obligation, where applicable, requires documenting processing operations, purposes, categories of data and data subjects, recipients, and related details. A profiling exercise or its tooling may produce useful inputs, but it does not by itself constitute a records of processing activities record or a complete data inventory. These are separate governance artifacts serving different accountability purposes, and demonstrable evidence is generally required rather than the mere existence of a profiling tool.
How does data profiling relate to a data catalog and data lineage?
Data profiling typically feeds a data catalog by supplying observed metadata about content, structure, and quality, and it can help establish or validate lineage by revealing where values originate and how they propagate. In practice, profiling results are often published into the catalog so stewards and data owners can act on them. Profiling is a governance activity focused on understanding data characteristics; it complements but does not replace catalog and lineage capabilities, and it does not on its own address security controls such as access management.
Who is accountable for acting on data profiling findings?
Accountability generally rests with the assigned data owners and data stewards defined in the governance framework, rather than with the profiling tool or the analyst who runs it. Under governance frameworks, accountability typically requires demonstrable evidence that findings were reviewed, prioritized, and remediated, not merely stated intent. Where profiling touches personal data, obligations may also attach to the controller in the relevant regime; this entry does not cover the specific allocation of statutory obligations.
Does profiling personal data require any additional safeguards?
When profiling operates on personal data, practitioners generally apply the same protections owed to that data throughout its lifecycle, and should be aware that profiling may itself constitute processing. Applying pseudonymization does not remove data from the scope of most regulation, and encryption or tokenization does not make data non-personal, so those measures do not exempt a profiling activity from applicable obligations. Whether a data protection impact assessment is warranted depends on context and is not automatically required. This entry does not cover lawful basis selection, retention, or cross-border transfer mechanics.
Where in a data lifecycle is profiling typically performed?
Profiling is commonly performed at ingestion or onboarding of a new source, during migration or integration projects, and on an ongoing basis to monitor data quality over time. It is often used to establish a baseline of quality metrics against which subsequent changes can be measured. The appropriate cadence and placement depend on the organization's data quality objectives and governance policies; this entry does not prescribe specific tooling, thresholds, or retention periods for profiling results.

Common misconceptions

Data profiling is the same as the 'profiling' regulated under the GDPR.
Under the EU GDPR and UK GDPR, profiling generally refers to automated processing used to evaluate personal aspects of a natural person, which triggers specific obligations. Data profiling as a data quality and governance technique is a separate activity focused on assessing the structure, content, and quality of datasets. Sharing terminology does not make them interchangeable, and regulatory treatment differs across regimes.
Data profiling is a security or data protection control that reduces regulatory risk on its own.
Data profiling is primarily a data governance and data quality activity. While profiling may help surface where personal data resides, it does not classify data, assign a lawful basis, implement confidentiality or integrity controls, or by itself satisfy any compliance obligation. Accountability under governance frameworks generally requires demonstrable evidence and controls beyond profiling output.
Profiling output can be treated as a complete records of processing activities or data inventory.
Profiling describes the state and quality of data; it does not constitute a records of processing activities obligation, which is a distinct regulatory record. A profiling tool's output is not the same as a maintained data inventory or processing register, and it typically does not cover processing purposes, lawful bases, recipients, or retention rules.

Best practices

Scope the exercise explicitly, distinguishing profiling for data quality purposes from any assessment of automated profiling of individuals under applicable privacy regimes, and document which you are performing.
Treat profiling as a governance activity that feeds data catalogs, lineage, and stewardship decisions rather than as a standalone compliance or security measure.
Where profiling surfaces likely personal data or special category data, route findings to appropriate classification and lawful-basis review processes rather than assuming the data is handled correctly.
Retain profiling results and methodology as demonstrable evidence to support accountability, since stated intent alone is generally insufficient under governance frameworks.
Define and agree on quality metrics such as completeness, uniqueness, validity, and consistency with data owners and stewards before profiling, so results map to actionable ownership.
Recognize the limits of the exercise and address retention rules, cross-border transfer mechanics, and data classification through separate dedicated processes rather than expecting profiling to cover them.