How text extraction from a PDF works
A PDF file structure is specified by ISO 32000-1:2008. It does not store sentences. It stores runs of glyphs, each placed at an exact coordinate on the page, in whatever order the generating program emitted them. That order frequently has nothing to do with reading order.
Extraction rebuilds the reading order from geometry. Runs sharing a baseline are grouped into a line, each line is ordered left to right, and a space is inserted wherever the horizontal gap between two runs exceeds normal letter spacing. Lines are then stacked top to bottom. The result is the text as a human reads it, rather than the order the file happened to store it in.
How to extract text from a PDF
- Open your PDF. Drop the file onto the box above, or click to browse. Every page is read locally, with progress shown as it goes.
- Choose how lines are joined. Keep paragraphs joined for prose you intend to read or reuse. Keep the exact lines when the line breaks themselves matter, such as in addresses, code or poetry.
- Add page markers if you need them. Switching on page markers inserts a labelled separator before each page, which makes it easy to cite or locate a passage later.
- Copy or download the text. The result appears in an editable box. Copy it to the clipboard, or download it as a .txt file.
What you get out
- All text from every page, in reading order, as one editable block.
- Paragraph mode โ wrapped lines are joined back into paragraphs, so a sentence broken across three printed lines becomes one sentence again.
- Exact-line mode โ every line break on the page is preserved, which matters for addresses, tables of contents, code listings and verse.
- Optional page markers โ a labelled separator before each page.
- Counts โ pages, words and characters, so you can confirm nothing was lost.
Which PDFs work
Any PDF containing a text layer extracts fully: documents exported from word processors, spreadsheets, accounting systems, design tools and web-page print dialogs. Page count is not limited, and a 500-page report processes in a few seconds because parsing runs locally.
Scanned documents contain no text layer, only a photograph of one, and return nothing. The tool detects this and says so explicitly, including the case where only some pages are scanned. Encrypted PDFs must be unlocked in a PDF reader first.
Why the document is never uploaded
Parsing runs through PDF.js inside your browser, so the file is read from local memory and never transmitted. Once the page has loaded, extraction works with the network disconnected.
Most online PDF-to-text converters upload the document, convert it server-side and return a link โ with your file resident on their infrastructure in the meantime. For a contract, a medical record, a payroll report or anything under NDA, local extraction removes that exposure entirely.
What extraction does not preserve
Plain text carries no formatting, so four things are lost by design:
- Styling. Bold, italics, fonts, sizes and colour do not survive.
- Column structure. A two-column layout is read as lines that span both columns where they share a baseline. Extract each column separately for a clean result.
- Table geometry. Cells arrive as text on a line. For CSV output with columns preserved, use the PDF table extractor.
- Images and figures. Only text is returned. Pull pictures out with the PDF image extractor.
Ligatures and unusual font encodings occasionally produce odd characters, which is a property of how the PDF embedded its fonts rather than of the extraction.
Why people extract PDF text
- Reusing content โ moving text out of a report or brochure into a document, email or CMS without retyping it.
- Searching an unsearchable file โ extracting the text so it can be grepped, indexed or read by a script.
- Feeding text to an AI tool โ producing clean plain text to paste into a model that cannot read PDFs directly.
- Translation and editing โ getting an editable version of a document supplied only as a PDF.
- Accessibility โ producing text a screen reader can voice reliably.
Text extraction compared with OCR
Use text extraction when the PDF has a text layer, and OCR only when it does not. Extraction reads character codes the file already contains, so it is exact, instant and lossless. OCR guesses characters from pixels, so it is slower and never perfect. Running OCR on a digital PDF makes the result worse, not better.
If this tool reports no text layer, the pages are images: use image to text OCR instead. To keep columns as columns, use the PDF table extractor. Once you have the text, the email and URL extractor pulls out the addresses and links inside it.
For a walkthrough of when a PDF needs OCR and when it already has a text layer, read scanned vs searchable PDF.
PDF text encoding and extraction edge cases
PDF text extraction has three encoding layers: the font's built-in encoding (WinAnsiEncoding, MacRomanEncoding or Identity-H for Unicode fonts), a ToUnicode CMap that maps glyph IDs to Unicode code points, and an ActualText attribute for ligature spans. When a PDF embeds a font with Identity-H encoding and a correct ToUnicode CMap, extraction is reliable. When neither is present โ common in PDFs from older InDesign versions and some Asian document workflows โ characters appear as hex codes rather than readable text.
Three extraction edge cases: (a) ligatures (fi, fl, ffl) are single glyphs with no 1:1 Unicode mapping unless ActualText is set โ they may appear as a single character, as missing or as an approximation; (b) right-to-left text (Arabic, Hebrew) is stored in visual order in most PDFs, not logical order โ the extractor returns characters in stream order, so RTL text can read reversed without post-processing; (c) Tagged PDF includes reading-order hints in its structure tree, which are used when present to improve column disambiguation and correct multi-column reading order.
Frequently asked questions
How do I extract text from a PDF for free?
Drop the PDF onto this page and the text appears in an editable box. Copy it or download it as a .txt file. No signup, no watermark and no page limit.
Is my PDF uploaded to a server?
No. The document is parsed by JavaScript in your browser and is never transmitted.
Why did my PDF return no text?
The document has no text layer, which means its pages are scanned images. Use image to text OCR to read them instead.
What is the difference between paragraph mode and exact-line mode?
Paragraph mode rejoins lines that were wrapped by the page layout, giving continuous prose. Exact-line mode keeps every line break as printed, which matters for addresses, code and verse.
Why is the text from a two-column page interleaved?
Lines are grouped by vertical position, so two columns sharing a baseline are read as one line. Extract one column at a time, or use a PDF reader's column-aware selection.
Can it extract text from a password-protected PDF?
No. Remove the password in a PDF reader first, then extract.
Is there a page limit?
No. Parsing happens locally, so long documents are limited only by your device's memory. Progress is shown while pages are read.