How table extraction from a PDF works
A PDF stores text as characters placed at fixed coordinates, not as a table. There are no rows, no columns and no cells in the file — only positions. Extraction rebuilds the structure by grouping every character run that shares a vertical position into a row, then splitting each row into cells wherever the horizontal gap between runs exceeds the width of a normal space.
This is why results depend on the document. A cleanly typeset invoice reconstructs almost perfectly. A table whose columns are separated by a single space is ambiguous even to a human reader, and the tool will merge those columns.
How to extract a table from a PDF
- Open your PDF. Drag the file onto the box above, or click to browse. Parsing starts immediately and runs locally.
- Pick the page. Use the page selector to move through the document. The detected row and column count is shown for each page.
- Check the preview. Confirm the columns split where you expect. The preview shows the first 200 rows; the download contains every row.
- Download the CSV. Export the current page, every page in one file, or the raw text with tab-separated columns.
What you get out
- CSV for the current page — properly quoted, so values containing commas survive the round trip into Excel or Google Sheets.
- CSV for the whole document — every page appended in order, separated by a blank line.
- Tab-separated plain text — useful when you want to paste straight into a spreadsheet without importing a file.
Row and column counts are reported per page, so a page where detection went wrong is obvious before you download anything.
Which PDFs work
Digitally generated PDFs work: invoices, bank statements, financial reports, exported spreadsheets and anything produced by a word processor or accounting system. These files contain a real text layer, which is what the extractor reads.
Scanned PDFs and photographed pages contain no text layer — only an image of one. Those return no rows. Run the page through image to text OCR first to create text, then work from that. Password-protected PDFs must be unlocked before extraction.
Why your document never leaves your device
Parsing runs in your browser through PDF.js, so the file is read from local memory and never uploaded. Competing table extractors post your PDF to a server, where it is processed and often retained for a period. For an invoice, a payroll report or a bank statement, that is a meaningful difference.
No account is required and no copy of the document exists after you close the tab.
Where table detection struggles
Three document patterns produce poor results, and recognising them saves you from trusting bad data:
- Columns separated by one space. The gap is too small to distinguish from a space inside a sentence, so neighbouring columns merge.
- Cells that wrap onto two lines. Each visual line becomes its own row, because the wrapped text sits at a different vertical position.
- Merged headers spanning several columns. These produce one wide cell rather than a repeated label.
Always check the preview against the original page before using the CSV in a calculation.
Who extracts tables from PDFs
- Bookkeeping and finance — moving invoice line items and bank statement rows into a ledger without retyping.
- Analysts and researchers — lifting published data tables out of reports and papers for their own calculations.
- Procurement — comparing supplier price lists that only arrive as PDFs.
- Operations — reconciling a PDF report against a system export.
Table extraction compared with text extraction
Use table extraction when the layout carries meaning and you need columns preserved. Use the PDF text extractor when you want the words in reading order and the column structure does not matter — for example to search a contract or copy a paragraph.
To pull pictures and diagrams out of the same document, use the PDF image extractor. For scanned pages with no text layer at all, start with image to text OCR.
For tips on getting cleaner results from difficult tables, read how to improve PDF table extraction accuracy.
PDF coordinate algorithm and table-detection edge cases
PDF stores text as character-positioning operators: Td/TD
for line offsets, Tm for the full transformation matrix, and
Tj/TJ for character streams. Table detection collects each
character's computed x/y position, groups characters that share a vertical position (within
a tolerance band) into rows, then splits rows into cells wherever the horizontal gap between
character runs exceeds the expected word-space width for the current font size.
Three patterns cause consistent problems: (a) rotated table headers — the rotation is encoded in the transformation matrix, so the row-grouping algorithm can place them at a different y-coordinate than their body cells; (b) right-aligned number columns in accounting tables often place the decimal point close to a text column's right edge, narrowing the gap below the word-space threshold and merging two adjacent columns; (c) tables in two-column page layouts share the same y-coordinate band with both columns interleaving in the output unless they are separated by a gap wide enough to register as a column boundary. Checking the row and column counts in the preview before downloading catches all three quickly.
Frequently asked questions
How do I convert a PDF table to Excel?
Drop the PDF above, check the preview, then click Download this page as CSV. Open the CSV in Excel or Google Sheets — CSV is imported natively by both.
Is my PDF uploaded to a server?
No. The document is parsed by JavaScript in your browser and never transmitted.
Why did my table come out as a single column?
The columns in that PDF are separated by a gap too narrow to detect, usually a single space. Extraction cannot distinguish that from a space inside a sentence.
Does it work on scanned PDFs?
No. A scanned page contains an image, not text. Run it through image to text OCR first to produce a text layer.
Is there a page limit?
No fixed limit. Every page is processed locally, so very long documents are bounded only by your device's memory. Progress is shown while parsing.
Can it handle password-protected PDFs?
No. Remove the password in a PDF reader first, then extract.