Personally Identifiable Information (PII)

Personally identifiable information (PII) is data that can identify a specific person, either on its own or when combined with other available data. The second half of that definition is where the difficulty lives: whether a piece of data is PII depends on what else exists alongside it, which means PII is a property of a situation rather than a fixed list of fields.

Direct identifiers Name, government identification number, passport number, email address, biometric identifiers
Indirect identifiers Date of birth, postcode, gender, job title, device identifiers — identifying in combination
U.S. reference definition NIST SP 800-122
EU equivalent ‘Personal data’ under the GDPR — broader than PII, and it includes pseudonymized data
Sensitive categories Health, biometric, financial, government identifiers, and in the EU a defined set of special categories
Biometric status Increasingly regulated as its own category rather than as ordinary PII
Key operational risk Re-identification — stripping names is not the same as anonymizing
Governing principle Data minimization — the safest record is the one never collected

Why the definition resists a fixed list

A postcode identifies nobody. A date of birth identifies nobody. Gender identifies nobody. Combine the three and you have narrowed a national population to a handful of individuals — in many cases to one. Research on re-identification has repeatedly demonstrated that small numbers of ordinary attributes are sufficient to single someone out of a large dataset.

This is why treating PII as a checklist of protected fields fails. An organization that carefully encrypts names while leaving a rich set of quasi-identifiers in the clear has protected the field that looks sensitive and left the ones that do the identifying.

The two frameworks most teams work under also do not line up:

PII (U.S. usage) Personal data (GDPR)
Scope Information that identifies or can be used to identify an individual Any information relating to an identified or identifiable natural person
Pseudonymized data Often treated as outside scope once identifiers are removed Expressly still personal data
Online identifiers Treated inconsistently Expressly included — IP addresses, cookie identifiers, device IDs
Practical effect Narrower Broader — a dataset outside U.S. PII definitions may still be in scope

Anonymization, pseudonymization, and the gap between them

These get used interchangeably and mean different things, with different legal consequences.

Pseudonymization replaces identifiers with a token while a key exists somewhere that reverses it. The data still relates to an identifiable person, and under the GDPR it remains personal data with the full set of obligations attached.

Anonymization means the link cannot be restored by anyone, including the original holder. Achieved properly it takes the data outside these regimes entirely. The difficulty is that it is much harder than it appears, and a dataset described as anonymized is frequently pseudonymized with the key retained somewhere in the organization — a distinction that only becomes visible during a breach.

Why PII matters for identity verification

Identity verification is unavoidably a PII-handling activity. Establishing who someone is means processing the most sensitive categories that exist: a government document, a facial image, a biometric template. That creates an obligation, and it creates a tension worth naming directly.

The obligation is minimization. The purpose of a verification check is a decision — is this person who they claim to be, are they old enough, do they match this document. The purpose is not to accumulate a durable archive of customer identity documents. Retaining images and biometric data beyond the point they are needed creates a liability that outlasts any benefit, and it is exactly the store that data breaches convert into fullz for sale.

The tension is that breached PII is what makes verification necessary in the first place. Knowledge-based authentication collapsed because the answers became public. Every organization holding more PII than it needs is contributing to the erosion of the checks everyone else relies on. Identity document verification that returns a decision and retains the minimum is the design that resolves it, and identity verification should be judged partly on what it declines to keep.

What PII controls can’t do

Removing names is not anonymizing. Quasi-identifiers re-identify people routinely, and a de-identified dataset is often nothing of the kind.

Encryption does not reduce scope. Encrypted personal data is still personal data, with the same obligations attached.

Consent does not remove the duty of care. A person agreeing to provide information does not transfer the risk of holding it.

Biometric data cannot be reissued. A compromised password is replaced; a compromised face is not, which is why biometric identifiers are increasingly regulated separately.

Frequently asked questions

What counts as personally identifiable information?

Direct identifiers such as name, government identification number and biometric data always qualify. Indirect identifiers — date of birth, postcode, gender, device identifiers — qualify when they can be combined to single out an individual, which small combinations frequently can. Whether data is PII depends on what else is available alongside it.

Is PII the same as personal data under the GDPR?

No. ‘Personal data’ is broader. It expressly includes online identifiers such as IP addresses and cookie identifiers, and it continues to cover pseudonymized data. A dataset that falls outside a narrow U.S. definition of PII may still be personal data and fully in scope.

What is the difference between anonymization and pseudonymization?

Pseudonymization replaces identifiers with tokens while a key exists that can reverse the process, so the data remains personal data. Anonymization means the link cannot be restored by anyone, which takes the data outside these regimes. True anonymization is considerably harder than it is usually assumed to be.

How long should identity verification data be retained?

Only as long as the purpose requires and any applicable regulation mandates. Verification exists to produce a decision, not an archive. Where record-keeping rules apply — as under Customer Identification Program requirements — those set the period; beyond them, retaining document images and biometric data creates liability without benefit.

Related reading

Discover Our Solutions

Exploring our solutions is just a click away. Try our products or have a chat with one of our experts to delve deeper into what we offer.

Report
Mapping the Rise of AI-Powered Identity Fraud

AI didn't just make fraud faster. It made it a system. We analyzed millions of identity interactions to map how identity attacks are evolving across regions, attack types, and sophistication levels — and what organizations need to rethink to keep pace.

See the Data