TL;DR: PDF metadata analysis reads the parts of a file the page never shows: the Info dictionary, the XMP packet, creation and modification timestamps, producer and creator strings, and the incremental update sections that record every save. Read against the stated date and claimed issuer, those fields tell an analyst how the file was made and whether it was touched afterward. Cotality's 2026 Mortgage Fraud Report estimates 1 in 119 mortgage applications carried indications of fraud in Q2 2026 and notes that AI-altered documents are harder than ever to identify visually. Metadata is strippable, so it is one check of six, not a verdict.
What PDF Metadata Actually Is
PDF metadata is the set of fields a PDF carries about itself, separate from the text and images a viewer renders, held in two stores. The document information dictionary, or Info dictionary, is a small block of key-value pairs: Title, Author, Subject, Keywords, Creator, Producer, CreationDate, ModDate. The XMP packet is an XML stream introduced with PDF 1.4 that repeats those fields under different names (CreatorTool, Producer, CreateDate, ModifyDate) and adds DocumentID, InstanceID, and an optional edit history. According to the PDF Association's 2025 presentation on PDF forensics, ISO 32000-2 deprecated most of the Info dictionary in favor of XMP, keeping only CreationDate and ModDate, and none of it is required.
The file structure carries a second layer of evidence. Creator names the application that authored the content; Producer names the library that wrote the PDF bytes. Under ISO 32000-1, a PDF can be modified by incremental update, which appends changed objects and a new cross-reference section to the end of the file while leaving the original bytes intact, so each save is recorded and the earlier version stays recoverable. Object streams, font subsets, embedded image encodings, and the PDF version in the header all carry the fingerprint of the tool that produced them.
What Each Signal Can and Cannot Prove
Every metadata field is evidence of a production step, not proof of intent.
Timestamps and incremental updates are the rows analysts misread most. A ModDate later than CreationDate is normal for a file printed to PDF before upload, and one incremental update is normal for a signed file. The finding is never "the file was saved twice." The finding is "the second save replaced the balance object on page one, and the earlier object is still in the file with a different value."
When Missing Metadata Is the Finding
The absence of metadata is itself a signal, because institutional PDF generators are consistent. A bank's core system, a payroll provider, or a company registry emits the same Producer string, PDF version, and structural pattern on every document, for years. The PDF Association presentation is direct: minimal metadata cannot prove a file is fraudulent, but its absence can undermine the rest of the investigation. A statement with no Producer, no CreationDate, and no XMP packet did not leave a banking system in that condition; something processed it afterward. The same applies to a Producer naming a bare PDF library with no application in front of it. Neither is a verdict. Both shift the burden toward requesting the native export from the issuer.
Producer Fingerprints, by Class
Producer strings sort into a handful of classes, and the useful question is whether the class fits the issuer. Bank core systems and enterprise document platforms write server-side libraries as Producer, often with a version fixed across a customer's entire statement history. Word processors and spreadsheets write themselves as Creator and a print or export engine as Producer, a normal fingerprint for a memo and an abnormal one for a bank statement. Browser-based design tools and online PDF editors write a web export engine, frequently with no Creator at all, and typically re-encode every embedded image. Desktop PDF editors tend to leave the original Producer in place and append an incremental update whose objects carry their fingerprint while the rest of the file carries the issuer's.
That last case matters most, because it is the signature of a genuine document edited after creation rather than a fake built from scratch. A file whose base objects say bank core and whose final update says desktop editor is a real statement with something changed, a different finding and a different customer conversation than a file that never touched a bank. Screenshots and image-only PDFs form their own class: no text objects, no fonts, one raster per page, a Producer from a phone or browser. Detecting an AI-generated PDF leans on this distinction, because generative pipelines usually deliver a flattened image or a file carrying a rendering library's fingerprint rather than an issuer's.
Reading the Timestamp Trail Against the Stated Date
The timestamp trail is the sequence of CreationDate, ModDate, any XMP history entries, and the release dates implied by recorded tool versions, read against the date printed on the document. A statement dated March 31 with a CreationDate of April 2 is ordinary; banks generate statements after the period closes. The same statement with a CreationDate in September, a ModDate the morning of the application, and a Producer version that did not exist in March needs an explanation.
The 2024 High Court judgment in COPA v Wright is the clearest public record of this test at work. Mr Justice Mellor found that Dr Craig Wright had forged documents on a grand scale in support of his claim to be Satoshi Nakamoto, and the appendix walks the evidence file by file. One document carried internal timestamps dating its creation to June 2007 but recorded a word processor build not released until September 2007. A drive image presented as untouched since 2007 held folder metadata showing a 2023 modification later overwritten with a 2007 date, made on a system clock set back sixteen years. None of those findings came from reading the page; each came from asking whether the stated date, the timestamps, and the release date of the recorded software form a sequence that could have happened. When they do not, the file is not yet a forgery, but the applicant owns the explanation.
Why Metadata Is One Check of Six
PDF metadata is trivially strippable and trivially writable. Any PDF library can rewrite the Info dictionary, drop the XMP packet, and flatten incremental updates into one clean file, so a reviewer who treats a tidy Producer string and matching dates as clearance is trusting the layer a fraudster controls most easily. That is why explainable document forensics treats metadata as the first breadcrumb rather than the conclusion, and why five other checks exist.
Production method looks past the labels to the byte-level structure tooling leaves behind whether or not it records its name. Issuer matching compares the file's fingerprint to how the claimed institution actually produces documents, which catches a stripped file as readily as a mislabeled one. Consistency tests whether numbers, dates, and layout agree with themselves and with account history. Model artifacts looks for the rendering traces image and language models leave in fonts, kerning, and pixel structure. Recycled patterns asks whether this skeleton or this exact file has been seen before, which is how template-farm documents get caught after their metadata is scrubbed. Metadata alone misses a stripped template; template matching alone misses a genuine statement with one edited field.
The Cotality 2026 Mortgage Fraud Report frames the stakes. Income misrepresentation remained Fannie Mae's top fraud finding at 46 percent of investigative findings for 2025, and the report attributes the growing difficulty to AI-altered income documents replacing whited-out pay stub figures, the edited-after-creation class that metadata surfaces first and structure confirms.
Where Sphinx Fits
Watchdoc reads the file's production story rather than its appearance. It parses the Info dictionary, XMP packet, every incremental update, fonts, and embedded images, runs the six checks — production method, timestamp trail, issuer matching, consistency, model artifacts, and recycled patterns — and returns a verdict with the objects that drove it highlighted, so an analyst can see which save changed which value. Published figures: 94.3% correct verdict, 2.8x more forgeries caught, clean files cleared in under 28 seconds, 1 million documents processed, $0.45 per document, no seats or platform fee. The Watchdoc playground runs the same x-ray on any PDF, free, no email required for the first file.
Frequently Asked Questions
What is PDF metadata analysis in fraud detection?
PDF metadata analysis is the examination of a PDF's Info dictionary, XMP packet, timestamps, producer and creator strings, and incremental update structure to establish how the file was made and whether it was modified after creation. Analysts compare those fields against the stated document date and the claimed issuer's normal output to judge whether the file's history is plausible.
Can PDF metadata prove a document is fake?
No. Metadata can show that a file's recorded history is impossible, such as a creation date earlier than the software that produced it, or that a genuine file was saved again by a different tool. Because every field can be rewritten or removed, metadata is corroborating evidence that structural analysis, issuer matching, and consistency checks must confirm.
What does it mean when a PDF has no metadata?
A PDF with no Producer, no CreationDate, and no XMP packet did not leave an institutional generator in that state, because bank cores, payroll providers, and registries emit consistent metadata on every file. Something processed it after issuance, and the reviewer should request the native export from the issuer.
Is a modification date later than the creation date a red flag?
Not by itself. Printing to PDF, signing, or filling a form all update ModDate and often append an incremental update. The signal is what the later save changed: a rewritten balance object with the earlier value still recoverable is a finding, and an added signature field is not.
What is a PDF incremental update?
An incremental update is the mechanism defined in ISO 32000 for saving changes to a PDF by appending modified objects and a new cross-reference section to the end of the file while leaving the original bytes in place. Because earlier objects remain in the file, an examiner can often recover what a page said before it was changed.

.png)