PDF417

The two-dimensional barcode many ID cards and driver's licenses in the region carry on the back. What it holds, why it is there, and what it lets you check that the front of the document does not.

In short

PDF417 is a stacked two-dimensional barcode format: instead of a single row of bars, it groups several short rows one on top of the other, which lets it store far more information than a linear code in a similar space. It is defined in the ISO/IEC 15438 standard.

On identity documents it serves a precise function. The front of the document shows the printed data for a person to read. The PDF417 on the back holds that same data in structured format, so a machine can read it without depending on print quality or on character recognition.

That duplication is what makes the code interesting for verification: there are two representations of the same data on the same document, and they can be cross-checked.

Why the structured data is on the back

In the region, the back of citizenship cards and driver's licenses is the usual spot for the barcode, and it is not an arbitrary convention.

The front has to be legible to a person: photograph, name, number, date, visible security features. All of that takes up space and is designed for the human eye. The back is left available for what only a machine has to read.

A two-dimensional code in that space solves three things at once. It stores more fields than would fit in a text strip. It includes error correction, so it stays legible even if part of the print is scratched, worn or poorly captured. And it avoids the most fragile step of reading the front, which is interpreting printed characters with OCR over a background of security patterns.

There is a direct operational consequence: a flow that captures only the front of the document does not access any of that data. It is reading the less reliable of the two available representations.

What a PDF417 on an identity document contains

The standard defines how information is encoded, not what information is encoded. Each issuing authority decides that, which is why the content varies between countries and between versions of the same document.

Even so, the pattern repeats. It is common to find the holder's biographical data and the document's registry data: first and last names, document number, date of birth, sex, issue and expiration dates, and often some internal identifier from the registry that issued the credential.

Two things are worth knowing before designing a process around this.

The first is that the structure is not universal. There is no single regional format: decoding the code returns a string, and knowing which field each part of that string is requires knowing the specification of the specific document and its issuing version.

The second is that some issuers encrypt the code's content. The data is there, but only whoever holds the key, usually the authority itself, can read it. Facing a document like that, the barcode contributes nothing to the cross-check, and verification relies on what is printed and on the security features of the physical document.

PDF417, MRZ and chip

All three are ways for a document to carry its data in machine-readable form, and each is responsible for something different.

  • PDF417 — greater capacity and tolerance for print damage. Its content and structure are defined by each issuer, and it can be encrypted. It needs decoding.
  • **MRZ** the strip of characters at the bottom, with an identical structure across countries because an international standard defines it, and with check digits that allow verifying consistency on the spot.
  • **NFC chip** the data is signed by the issuing authority. It is the only one of the three that allows checking authenticity cryptographically.

The underlying difference: PDF417 and the MRZ say what is written on the document. The chip lets you check who wrote it.

A document may carry all three, two or just one. That is why a document verification process that assumes only one of them runs out of data as soon as the type of document presented changes.

What reading it is good for in a verification

Extracting the data is the means, not the end. What the code enables are checks.

  • Cross-checking sides — the code's fields are compared against what the front shows printed. It is the basis of the front-and-back cross-check, and it is where an alteration made on just one side shows up.
  • Extraction quality — when the same field arrives through two paths, a discrepancy flags a reading problem before it reaches the case file.
  • Fields the front does not show — some issuers include registry data in the code that is not printed, and it is useful for tracing the document.
  • Less capture friction — a code with error correction decodes well under photo conditions where character recognition on the front starts to fail.

What it does not do, and it is worth saying with the same clarity: the barcode does not prove the document is authentic. A full back can be reproduced. What it proves is what data that back carries, which is exactly what allows cross-checking it against the other side.

Frequently asked questions

It is a stacked two-dimensional barcode format, defined in the ISO/IEC 15438 standard, that organizes information into several stacked rows instead of a single line of bars. That lets it store considerably more information than a linear code and survive partial print damage, because it includes error correction. On identity documents it is used to carry, on the back and in structured format, the same data the front shows printed.

It depends on the country and the document version, because the standard defines how it is encoded but not what is encoded. It is common to find the holder's biographical data and the document's registry data: name, number, date of birth, sex and issue and expiration dates, often together with some internal registry identifier. Some issuers encrypt that content, in which case the code is only legible to whoever holds the key.

No. Reading it returns data, not proof of authenticity, and a back can be reproduced. Its value lies in enabling checks: cross-checking those fields against what is printed on the front, catching reading discrepancies, and accessing fields the front does not show. Full verification adds to that the analysis of the document's security features and, when it exists, reading the chip, which does allow checking the issuing authority's signature.

One identity, one SDK

VU ONE brings identity verification, authentication and fraud protection together on a single identity graph.

The verification you run at signup stays available to authentication and to your fraud rules, with no repeated processes and no duplicated data.

Verify, Authenticate and Protect, consolidated in one place.

Request a demo