Data Extraction

Data extraction is the conversion of information held in an image or unstructured file into structured fields a system can act on. In identity verification it means turning a photograph of a document into a name, a date of birth, a document number and an expiry date. It is the first step in a verification flow and, on its own, it verifies nothing.

Input A photograph, scan or PDF — unstructured, variable quality
Output Named fields with values, usually with per-field confidence scores
Techniques Optical character recognition, barcode decoding, MRZ parsing, chip reading, layout analysis
Identity documents Highly structured — field positions are known once the document type is classified
Non-identity documents Far less structured — bank statements and pay stubs vary by issuer
Precision requirement Unusually high — one wrong character in a document number is a failure
Strongest quality check Cross-referencing the same value from two independent encodings
What it does not do Establish that the document or the data is genuine

How extraction works on an identity document

The sequence is more constrained than general document processing, because identity documents are designed to be machine-read.

The document is located in the frame, then classified — which country, which document type, which version. That classification is what makes everything after it tractable, because once the template is known, field positions are known. A U.S. driver’s license and a German identity card share almost nothing structurally, and neither is read by the same model.

Then the encodings are read. Printed fields go through optical character recognition. The machine-readable zone is parsed according to the ICAO specification, with check digits that validate the read. A PDF417 barcode on the back of a U.S. license decodes to a defined field set. Where a chip is present, it is read over NFC and carries a signature.

The valuable step is the last one: the same fact is now available from several sources, and they should agree. A date of birth appears in the printed text, the MRZ and the barcode. When those disagree, either the read failed or the document was altered — and either way the system knows something is wrong without any forensic analysis at all.

Extraction is not verification

This is the distinction the term itself invites people to blur, and it is worth being blunt about.

Operation The question it answers
Data extraction What does this document say?
Security features check Is this document genuine?
Tamper detection Has this image been altered?
Biometric comparison Is the presenter the person the document describes?

Only the first is extraction, and a process that stops there has converted an unverified image into structured data — which then flows downstream looking authoritative because it is well-formatted. A tampered document extracts cleanly. A photograph of a screen extracts cleanly. A high-quality counterfeit extracts perfectly, because it was made to.

The other failure mode is quieter and more common: confident output on an input that should have been rejected. A blurred or partially obscured field can still produce a plausible-looking value, and a downstream system with no visibility into capture quality treats it as fact. Systems that score their own input and request a retake avoid creating a problem that is invisible afterwards.

Why this matters for identity verification

Extraction quality determines what every later step has to work with. A misread document number fails a database check that would otherwise have passed. A misread date of birth produces a false age result or an unresolvable sanctions match. The cost of an extraction error is rarely borne by the extraction step — it surfaces later as a false decline or an unexplained review.

It also carries the largest single share of recoverable onboarding conversion. Applicants who retry capture three times and abandon were lost to extraction and capture quality, not to a fraud decision.

The honest framing is that extraction is necessary infrastructure and never sufficient evidence. Identity document verification is extraction plus authentication plus binding, and for the non-identity document stack — pay stubs, bank statements, proof of address — non-ID document verification faces the same gap with less structure to lean on.

What data extraction can’t do

It does not authenticate anything. It reads what is present, whether or not what is present is genuine.

It cannot recover what the capture lost. Enhancement improves appearance without restoring detail that was never recorded. Past a threshold the only correct action is a retake.

Confidence is not accuracy. A high score on a poor input is the most dangerous output in the pipeline.

Coverage is per document type and version. Extraction depends on classification, and an unrecognized document version degrades to generic reading at best.

Frequently asked questions

What is data extraction in identity verification?

Converting a photograph or scan of a document into structured fields — name, date of birth, document number, expiry date — that a system can act on. It combines optical character recognition, barcode decoding, machine-readable zone parsing and, where available, chip reading.

Is data extraction the same as document verification?

No, and conflating them is a common and consequential error. Extraction reads what the document says. Verification establishes that the document is genuine, that the image has not been altered, and that the person presenting it is its holder. A tampered document or a good counterfeit extracts perfectly.

How accurate does extraction need to be?

More accurate than most document processing applications, because a single wrong character in a document number causes a downstream check to fail. This is why cross-referencing matters — the same value read independently from printed text, the machine-readable zone and a barcode should agree, and disagreement signals a problem without any further analysis.

Why do extraction errors cause customer drop-off?

Because the cost surfaces later. A misread field produces a failed check or an unexplained review rather than an obvious error, and repeated capture attempts cause abandonment. A large share of recoverable onboarding conversion is decided at capture and extraction rather than at the fraud decision.

Related reading

Discover Our Solutions

Exploring our solutions is just a click away. Try our products or have a chat with one of our experts to delve deeper into what we offer.

Report
Mapping the Rise of AI-Powered Identity Fraud

AI didn't just make fraud faster. It made it a system. We analyzed millions of identity interactions to map how identity attacks are evolving across regions, attack types, and sophistication levels — and what organizations need to rethink to keep pace.

See the Data