Data Aggregation
Data aggregation is the process of gathering data from one or more sources and combining or summarizing it into a unified, report-based form for analysis. While this can produce useful insights and a holistic view of information, bringing data together can also increase risk, because the combined result may reveal more than any individual data point on its own.
Data aggregation is the process by which raw data is searched, gathered, combined, and expressed in a summarized form to support statistical analysis, reporting, or decision-making. From a risk and governance perspective, the aggregation of otherwise discrete data elements can produce a holistic view whose sensitivity or re-identification potential exceeds that of the constituent parts, an effect sometimes described as the aggregation problem. This entry addresses the concept and its associated risk profile only; it does not cover specific lawful bases for aggregating personal data, cross-border transfer mechanics, retention obligations, or whether a particular aggregated output constitutes personal data, special category data, or effectively anonymized data. Note in particular that aggregation or summarization does not by itself render data non-personal; whether an aggregated dataset falls outside data protection obligations depends on the residual re-identification risk and is a jurisdiction- and context-specific determination.
Why it matters
Data aggregation matters because the risk profile of combined data frequently exceeds the sum of its parts. Individual data elements may appear innocuous in isolation, but when compiled together they can provide a holistic view that reveals sensitive attributes, patterns, or identities not discernible from any single point. This effect, sometimes described as the aggregation problem, means that governance and risk assessments performed at the level of discrete data fields can understate the exposure created once those fields are brought together into a unified, report-based form.
For data protection and governance professionals, this has a direct practical consequence: aggregation or summarization does not by itself render data non-personal. Whether an aggregated output falls outside data protection obligations depends on the residual re-identification risk, which is a jurisdiction- and context-specific determination rather than an automatic outcome of the aggregation process. Treating an aggregated dataset as anonymized without assessing that residual risk is a common and consequential error, because an insufficiently protected aggregate may remain personal data and continue to attract the associated obligations.
Aggregation also sits at the intersection of governance and security. Governance concerns such as ownership, stewardship, data quality, and lineage determine what sources are combined and for what purpose, while security controls address how the combined result is protected. Neither discipline alone addresses the aggregation problem, and accountability for aggregated outputs generally requires demonstrable evidence of how sensitivity was assessed, not merely a stated intent to summarize responsibly.
Who it's relevant to
Inside Data Aggregation
Common questions
Answers to the questions practitioners most commonly ask about Data Aggregation.