What metadata a PDF carries
A PDF stores its own description in two places. The document information dictionary holds the classic fields — Title, Author, Subject, Keywords, Creator, Producer, CreationDate and ModDate. Newer files also carry an XMP packet, an XML block that can hold far more, including editing history.
The two most revealing fields are the least obvious. Creator records the application the content was authored in; Producer records the library that actually wrote the PDF bytes. A document claiming to be a scanned contract but produced by a headless browser's print-to-PDF was generated from a web page, not scanned. That mismatch is the single most useful thing in the record.
How to read a PDF's metadata
- Open the PDF. Drop the file onto the box above, or click to browse. The document information dictionary and page geometry are read locally.
- Read the record. Author, title, creation and modification dates, the creating application and the producing library all appear together, with dates converted from PDF's own format into readable ones.
- Check the structural findings. Page size and rotation, whether page sizes are mixed, whether a text layer exists, and whether the file carries permission restrictions, bookmarks, attachments or form fields.
- Export. Copy the details to the clipboard or download the whole record as a text file.
What the tool reports
- Identity fields — title, author, subject and keywords, as recorded in the file.
- Dates — created and modified, decoded from PDF's
D:YYYYMMDDHHmmSSformat including its timezone offset. - Software trail — the creating application and the producing library, plus the PDF format version.
- Page geometry — dimensions in millimetres, the matching standard size where there is one (A4, Letter, Legal and others), rotation, and a warning when pages are not all the same size.
- Text layer check — whether the pages contain extractable text or are images requiring OCR, sampled across several pages rather than assumed from the first.
- Structure — permission restrictions, bookmark count, embedded attachments with their filenames, form field count, and how many XMP properties are present.
Which PDFs work
Every unencrypted PDF works, at any version from 1.0 to 2.0, and page count is not limited. To keep things fast on long documents, page sizes are compared across the first 25 pages and the text-layer check samples the first 5 — enough to catch a mixed or scanned document without reading a 900-page file in full.
PDFs protected by an open password cannot be read at all until the password is supplied; that is reported clearly. Files carrying only permission flags — restricting printing or copying rather than opening — are read normally, and the presence of those flags is itself reported.
Why reading metadata locally matters more than usual
Parsing runs through PDF.js inside your browser, so the file is never transmitted.
For this tool specifically, that is close to the whole point. Metadata is exactly the data people want to inspect because it may be sensitive — an author's real name, an internal file path in the title field, a creation date that contradicts a claim, a company's licensed software in the Producer string. Uploading a document to a website to find out what it discloses would disclose it to that website first.
What metadata cannot tell you
Metadata is a record the file makes about itself, and records can be wrong or absent:
- Every field is editable. Author and dates can be set to anything. Metadata is evidence, not proof.
- Many files carry almost none. Stripping metadata is a normal privacy step, so blanks mean "not recorded", not "nothing to hide".
- Dates have no timezone guarantee. The offset is recorded but may be the generating machine's local setting, which can be wrong.
- Full XMP history is not expanded. The property count is reported, but individual edit-history entries are not parsed out.
This tool also reads nothing about content — it reports whether text exists, not what the text says.
Who reads PDF metadata
- Checking a document before you trust it — confirming a supposed scan was actually scanned, or that an invoice came from the accounting system it claims.
- Privacy review before publishing — finding your own name, username or internal file path left in a document you are about to put online.
- Print production — verifying page size, rotation and consistency before sending to a printer.
- Diagnosing why a PDF will not behave — mixed page sizes, missing text layer or permission flags explain most complaints.
- Archiving — recording provenance for a document collection.
Reading metadata compared with extracting content
This tool answers questions about the file. To get what is inside it, use a different one: the PDF text extractor for the words, and the PDF table extractor for tabular data as CSV.
Run this first when a PDF is misbehaving, because it tells you which of those will work. "Text layer: none found" means text extraction will return nothing and the pages need image to text OCR instead — which is far quicker to learn here than by extracting an empty file and wondering why.
For a full inventory of every field a PDF can carry and what each reveals, read what metadata is stored in a PDF.
PDF metadata format standards and edge cases
PDF has two parallel metadata systems: the document information dictionary (PDF object
with keys /Title, /Author, /Producer,
/Creator, /CreationDate and /ModDate) and XMP metadata
(an XML stream in ISO 16684 format in the document's /Metadata stream). When
both exist, XMP takes precedence per ISO 32000-2. PDF date strings follow the format
D:YYYYMMDDHHmmSSOHH'mm'.
Three edge cases: (a) many PDF creators — including macOS Preview and web browsers — leave
/Author empty but set xmp:CreatorTool to their own application name;
(b) the /Producer field records the library that serialised the bytes
(Ghostscript, iText, PDF.js) rather than the application the user saw — a Canva export
shows iText as Producer, not Canva; (c) creation and modification dates can be identical
(the library wrote them at the same instant) or decades apart (an InDesign file from 2008
last opened in 2024).
Frequently asked questions
How do I see who created a PDF?
Drop it onto this page. The Author, Creator and Producer fields are read from the document's own information dictionary — Creator is the authoring application, Producer is the library that wrote the file.
Is my PDF uploaded to a server?
No. It is parsed by JavaScript in your browser and never transmitted, which matters especially for a tool whose purpose is finding sensitive metadata.
Can PDF metadata be trusted?
Treat it as evidence rather than proof. Every field is editable, and blank fields usually mean metadata was deliberately stripped.
How do I know if a PDF is scanned?
The Text layer row answers it. “None found” means the pages are images and OCR is required; “present” means the text can be extracted exactly.
Why does it say my pages are different sizes?
The first 25 pages were compared and did not match. Mixed page sizes commonly come from merging documents, and they cause most unexpected printing results.
What does the Restrictions row mean?
The PDF carries permission flags limiting actions such as printing or copying. These are honoured by convention rather than enforced by encryption, and they are separate from an open password.
Can it read a password-protected PDF?
Not one that needs a password to open. Remove the password in a PDF reader first. Files with only permission flags are read normally.
Does it show the XMP metadata in full?
It reports how many XMP properties exist and names the first few. Individual edit-history entries are not expanded.