A dataset does not become anonymous simply because the obvious identifiers have disappeared.
Italy’s privacy regulator has made that distinction expensive.
On 2 October, the Garante announced a €7 million fine against IQVIA Solutions Italy after concluding that a large health-data database the company regarded as anonymous still contained personal data.
The case concerns sensitive health information, not an advertising audience or a marketing clean room.
But the underlying data question matters far beyond healthcare:
When does pseudonymisation reduce risk, and when does a dataset actually stop being personal data?
For measurement teams, CRM operators and advertisers working with persistent IDs, that is a distinction worth auditing.
What the Italian regulator found
According to the Garante, IQVIA had built a database containing health information relating to around one million patients from 800 general practitioners.
The database was used for research commissioned in part by pharmaceutical companies.
IQVIA’s position was that the patient information had been anonymised.
The regulator disagreed.
One element was especially important: each patient was associated with a code that allowed the same person to be followed over time.
That persistent code sat alongside detailed information including year of birth, sex, diagnoses, symptoms, prescriptions, examinations, vaccinations and location information.
The Garante said that combination made it possible to isolate individual patients and, using reasonably available means, potentially reidentify them.
The authority therefore treated the dataset as personal data rather than anonymous information.
A random-looking identifier is not the same as anonymity
This is the part marketers should pay attention to.
A database can remove:
- names;
- email addresses;
- telephone numbers;
- account numbers;
and still remain personal data.
If a stable identifier continues to distinguish one person from another, that dataset can preserve a person’s history across time.
Add enough attributes and the possibility of singling someone out can increase.
That is pseudonymisation territory, not necessarily anonymisation.
The practical distinction is important because pseudonymised personal data remains inside the GDPR framework.
Anonymised information that genuinely falls outside the ability to identify a person is different.
The IQVIA decision did not invent that distinction.
It shows how a regulator can apply it to a real, high-dimensional dataset.
Why this matters for marketing data
Most marketing teams are not processing patient diagnoses.
But modern advertising and analytics systems regularly use persistent identifiers.
Examples include:
- hashed email addresses;
- CRM customer IDs;
- device or browser identifiers;
- clean-room join keys;
- loyalty IDs;
- pseudonymous conversion records;
- long-lived audience identifiers.
Hashing an email address does not magically transform the underlying person into anonymous data.
Neither does replacing a customer number with another stable code.
The question is what the organisation and other reasonably capable parties can still do with the dataset.
Can records be linked over time?
Can they be joined to another table?
Can a small group be isolated using location, age, purchase behaviour or other attributes?
Is there another party that holds the lookup information?
Those questions are more useful than asking whether the identifier looks unreadable.
Clean rooms do not remove the classification question
Data clean rooms can reduce unnecessary exposure.
They can constrain which fields parties exchange, limit raw-data access and control what queries or outputs are permitted.
Those are meaningful protections.
But the label clean room does not automatically mean the underlying information is anonymous.
A clean room may intentionally operate on pseudonymous personal data so that two parties can match records without openly exchanging direct identifiers.
That can still be a legitimate architecture.
It simply means the privacy analysis does not end when the identifier is hashed.
Teams still need to understand legal basis, purpose limitation, data minimisation, retention, access and the risk of reidentification.
The Italian case is useful precisely because the regulator looked at the dataset as a system rather than focusing on whether names had been stripped out.
Run this six-question identifier audit
European marketing and analytics teams can use the ruling as a trigger for a short data audit.
1. Is the identifier persistent?
Can the same user, household or customer be followed across sessions, purchases or months?
If yes, document why that persistence is necessary.
2. What attributes sit beside it?
List location, age bands, purchase history, browsing behaviour, interests and any other fields.
Risk comes from combinations, not just individual columns.
3. What can the dataset be joined to?
Map internal and external tables that can connect through the identifier or overlapping attributes.
4. Who holds the lookup information?
A dataset may look anonymous to one processor while another party can reconnect the identifier to an individual.
Record that relationship.
5. How small can segments become?
Highly granular combinations can make people easier to isolate even without a name.
Review minimum audience and reporting thresholds.
6. What would have to be true for you to call it anonymous?
Write that test down.
If the organisation cannot explain why reidentification is no longer reasonably possible, “anonymous” may be too strong a description.
What the IQVIA ruling does not say
The decision should not be stretched into a claim that hashing is unlawful.
It does not ban pseudonymisation.
It does not say every persistent identifier is automatically illegal.
It does not ban data clean rooms.
And it does not mean every pseudonymous advertising dataset creates the same risk as a longitudinal health database.
The regulator’s decision involved unusually sensitive and detailed information.
The useful lesson is narrower:
Removing direct identifiers is not enough, by itself, to prove anonymity.
That is a better standard for marketing teams because it forces the organisation to examine what the data can still reveal and how it can still be connected.
The marketing takeaway
For years, digital measurement language has used terms such as “anonymous”, “hashed”, “de-identified” and “privacy-safe” too loosely.
The IQVIA case is a reason to clean up that vocabulary.
If a dataset still uses stable identifiers and allows a person’s activity to be connected over time, call it what the evidence supports.
Then apply the controls appropriate to that classification.
That is more useful than relying on a reassuring label.
NEMO will keep tracking these distinctions in Data & Measurement, especially where clean rooms, conversion APIs, first-party audiences and AI-driven measurement create new combinations of persistent data.