Extract the Metadata From a PDF

Every PDF records who made it, when, and which software wrote it — and that record frequently contradicts what the document appears to be. Drop a file below to read its author, dates, producing software, page geometry, restrictions and whether it contains a real text layer. The PDF is read in your browser and never uploaded.

Drop a PDF here
or click to choose a file · nothing is uploaded

What metadata a PDF carries

A PDF stores its own description in two places. The document information dictionary holds the classic fields — Title, Author, Subject, Keywords, Creator, Producer, CreationDate and ModDate. Newer files also carry an XMP packet, an XML block that can hold far more, including editing history.

The two most revealing fields are the least obvious. Creator records the application the content was authored in; Producer records the library that actually wrote the PDF bytes. A document claiming to be a scanned contract but produced by a headless browser's print-to-PDF was generated from a web page, not scanned. That mismatch is the single most useful thing in the record.

How to read a PDF's metadata

  1. Open the PDF. Drop the file onto the box above, or click to browse. The document information dictionary and page geometry are read locally.
  2. Read the record. Author, title, creation and modification dates, the creating application and the producing library all appear together, with dates converted from PDF's own format into readable ones.
  3. Check the structural findings. Page size and rotation, whether page sizes are mixed, whether a text layer exists, and whether the file carries permission restrictions, bookmarks, attachments or form fields.
  4. Export. Copy the details to the clipboard or download the whole record as a text file.

What the tool reports

Which PDFs work

Every unencrypted PDF works, at any version from 1.0 to 2.0, and page count is not limited. To keep things fast on long documents, page sizes are compared across the first 25 pages and the text-layer check samples the first 5 — enough to catch a mixed or scanned document without reading a 900-page file in full.

PDFs protected by an open password cannot be read at all until the password is supplied; that is reported clearly. Files carrying only permission flags — restricting printing or copying rather than opening — are read normally, and the presence of those flags is itself reported.

Why reading metadata locally matters more than usual

Parsing runs through PDF.js inside your browser, so the file is never transmitted.

For this tool specifically, that is close to the whole point. Metadata is exactly the data people want to inspect because it may be sensitive — an author's real name, an internal file path in the title field, a creation date that contradicts a claim, a company's licensed software in the Producer string. Uploading a document to a website to find out what it discloses would disclose it to that website first.

What metadata cannot tell you

Metadata is a record the file makes about itself, and records can be wrong or absent:

This tool also reads nothing about content — it reports whether text exists, not what the text says.

Who reads PDF metadata

Reading metadata compared with extracting content

This tool answers questions about the file. To get what is inside it, use a different one: the PDF text extractor for the words, and the PDF table extractor for tabular data as CSV.

Run this first when a PDF is misbehaving, because it tells you which of those will work. "Text layer: none found" means text extraction will return nothing and the pages need image to text OCR instead — which is far quicker to learn here than by extracting an empty file and wondering why.

For a full inventory of every field a PDF can carry and what each reveals, read what metadata is stored in a PDF.

PDF metadata format standards and edge cases

PDF has two parallel metadata systems: the document information dictionary (PDF object with keys /Title, /Author, /Producer, /Creator, /CreationDate and /ModDate) and XMP metadata (an XML stream in ISO 16684 format in the document's /Metadata stream). When both exist, XMP takes precedence per ISO 32000-2. PDF date strings follow the format D:YYYYMMDDHHmmSSOHH'mm'.

Three edge cases: (a) many PDF creators — including macOS Preview and web browsers — leave /Author empty but set xmp:CreatorTool to their own application name; (b) the /Producer field records the library that serialised the bytes (Ghostscript, iText, PDF.js) rather than the application the user saw — a Canva export shows iText as Producer, not Canva; (c) creation and modification dates can be identical (the library wrote them at the same instant) or decades apart (an InDesign file from 2008 last opened in 2024).

Frequently asked questions

How do I see who created a PDF?

Drop it onto this page. The Author, Creator and Producer fields are read from the document's own information dictionary — Creator is the authoring application, Producer is the library that wrote the file.

Is my PDF uploaded to a server?

No. It is parsed by JavaScript in your browser and never transmitted, which matters especially for a tool whose purpose is finding sensitive metadata.

Can PDF metadata be trusted?

Treat it as evidence rather than proof. Every field is editable, and blank fields usually mean metadata was deliberately stripped.

How do I know if a PDF is scanned?

The Text layer row answers it. “None found” means the pages are images and OCR is required; “present” means the text can be extracted exactly.

Why does it say my pages are different sizes?

The first 25 pages were compared and did not match. Mixed page sizes commonly come from merging documents, and they cause most unexpected printing results.

What does the Restrictions row mean?

The PDF carries permission flags limiting actions such as printing or copying. These are honoured by convention rather than enforced by encryption, and they are separate from an open password.

Can it read a password-protected PDF?

Not one that needs a password to open. Remove the password in a PDF reader first. Files with only permission flags are read normally.

Does it show the XMP metadata in full?

It reports how many XMP properties exist and names the first few. Individual edit-history entries are not expanded.

• Specialist file parsing & security engineer • Verified: in our experience, our hands-on testing measured and verified private in-browser execution with zero file uploads • Last reviewed July 2026.