Why PDF Table Extraction Isn’t Always Accurate
PDF table extraction isn’t always accurate because a PDF has no real table underneath — no rows, no columns, no cells. It only has text placed at x/y coordinates, and an extractor has to infer the grid from where the characters sit. Anything that blurs those positions — merged cells, missing borders, multi-line cells, or a scanned page — makes the guess harder. Understanding what the tool is actually doing tells you which tables extract cleanly and how to help the ones that don’t.
This guide explains why accuracy varies, what specifically breaks it, and how to get the best result from the PDF table extractor.
There’s no table in the file — only positions
When you look at a PDF table you see a grid, but the file doesn’t store one. It stores instructions like “draw ‘£1,240’ at x=310, y=560”. The lines you see may be separate drawing commands, or may not exist at all. To rebuild the table, an extractor clusters text by its horizontal position into columns and its vertical position into rows. When the layout is clean and regular, this works very well. When it isn’t, the inference slips.
What breaks accuracy
- No visible borders. Borderless tables give the extractor no ruling lines to lean on, so it relies entirely on spacing. Columns that are close together can merge, or a wide gap inside a cell can be read as a new column.
- Merged and spanning cells. A heading that spans three columns, or a cell merged down two rows, has no equivalent in a flat grid — so it lands in one cell and leaves blanks beside it.
- Multi-line cells. A cell whose text wraps onto two lines can be split into two rows, because each line sits at a different vertical position.
- Irregular alignment. Numbers aligned right and text aligned left within the same column can confuse column boundaries.
- Scanned tables. If the page is an image, there’s no text to position at all — you need OCR first, and OCR adds its own error rate on top.
How to get a cleaner result
- Check the source type. If you can’t select the table text, it’s a scan — run it through OCR first, then extract.
- Extract, then fix in a spreadsheet. Treat extraction as getting you 90% of the way. Pull the table to CSV with the PDF table extractor, open it, and repair the few merged-cell or wrapped-line rows by hand — far faster than retyping.
- Extract one table at a time when a page has several, so the column inference isn’t fighting two different grids.
- Prefer the digital original. If the same table exists as a real PDF and a scan, always use the digital one — its text positions are exact.
Set your expectations by table type
| Table | Typical accuracy |
|---|---|
| Ruled, single-line cells, digital PDF | Very high — usually clean |
| Borderless but well-spaced, digital | Good — occasional column slip |
| Merged/spanning or wrapped cells | Partial — needs light cleanup |
| Scanned table (needs OCR) | Variable — depends on scan quality |
Once you have the CSV, the CSV column extractor can pull just the columns you need. For the full path from document to spreadsheet, see the PDF to CSV workflow.
Frequently asked questions
Why is my PDF table extraction inaccurate?
Because a PDF has no real table structure — an extractor infers rows and columns from text positions. Merged cells, missing borders, wrapped text and scans all make that inference harder.
How do I extract a table from a PDF accurately?
Use the digital PDF (not a scan), extract to CSV with a PDF table extractor, then fix any merged or wrapped rows in a spreadsheet. Run scans through OCR first.
Why did one cell split into two rows?
The cell’s text wrapped onto two lines, and each line sits at a different vertical position, so the extractor read them as separate rows. Rejoin them in your spreadsheet.
Can I extract a table from a scanned PDF?
Only after OCR, which recognises the text from the image. Accuracy then depends on the scan quality.
Why are there blank cells next to a heading?
A heading that spans several columns can’t map to a flat grid, so it lands in one cell and leaves the neighbours blank.
Related reading
Last updated: 16 August 2026.