PDF Extraction Tools

A PDF does not store sentences or tables. It stores glyphs at fixed coordinates, in whatever order the program that made it happened to emit them. Every tool here rebuilds something different from those coordinates, so the right choice depends on what you need back — running prose, a grid of values, or facts about the file itself.

Which tool do you need?

If you…Use
You want the words, in reading orderPDF Text Extractor
You want a table as CSV, with its columns intactPDF Table Extractor
You want to know who made the file, when, and whether it is scannedPDF Metadata Extractor

Text and tables are genuinely different extractions. Plain text discards column structure, so anything you intend to calculate with should come out as a table; anything you intend to read should come out as text.

Check for a text layer before anything else

One property decides whether any of this works: whether the PDF has a text layer. A document exported from a word processor, an accounting system or a browser's print dialog contains real characters, and extraction from it is exact. A document that came off a scanner or a phone camera contains a photograph of a page and no characters at all — every text tool will return nothing.

The quickest test is to try selecting a line of text in a PDF reader. If a selection rectangle appears instead of highlighted words, there is no text layer. The PDF metadata extractor answers the same question directly by sampling several pages, which is worth doing first when a document is misbehaving — it tells you in seconds whether to keep going or switch to OCR.

For a scanned PDF the route is image to text OCR: export the pages as images, recognise them, and work from the result.

Every pdf extraction tool

Frequently asked questions

Why did my PDF return no text?

It has no text layer — the pages are images. Confirm with the PDF metadata extractor, then use OCR instead.

Should I use the text extractor or the table extractor?

Use the table extractor when the columns carry meaning and you want CSV. Use the text extractor when you want the words and the layout does not matter.

Are my PDFs uploaded?

No. All three tools parse the document with PDF.js inside your browser. Nothing is transmitted, which matters for contracts, statements and medical records.

Other extraction groups