Extract Text From a PDF

Extracting text from a PDF returns the document's words as editable plain text. Drop a file below to pull the text out of every page in reading order, as flowing paragraphs or as the exact lines on the page. The PDF is read inside your browser and never uploaded, so confidential documents stay on your device.

Drop a PDF here
or click to choose a file · nothing is uploaded

How text extraction from a PDF works

A PDF file structure is specified by ISO 32000-1:2008. It does not store sentences. It stores runs of glyphs, each placed at an exact coordinate on the page, in whatever order the generating program emitted them. That order frequently has nothing to do with reading order.

Extraction rebuilds the reading order from geometry. Runs sharing a baseline are grouped into a line, each line is ordered left to right, and a space is inserted wherever the horizontal gap between two runs exceeds normal letter spacing. Lines are then stacked top to bottom. The result is the text as a human reads it, rather than the order the file happened to store it in.

How to extract text from a PDF

  1. Open your PDF. Drop the file onto the box above, or click to browse. Every page is read locally, with progress shown as it goes.
  2. Choose how lines are joined. Keep paragraphs joined for prose you intend to read or reuse. Keep the exact lines when the line breaks themselves matter, such as in addresses, code or poetry.
  3. Add page markers if you need them. Switching on page markers inserts a labelled separator before each page, which makes it easy to cite or locate a passage later.
  4. Copy or download the text. The result appears in an editable box. Copy it to the clipboard, or download it as a .txt file.

What you get out

Which PDFs work

Any PDF containing a text layer extracts fully: documents exported from word processors, spreadsheets, accounting systems, design tools and web-page print dialogs. Page count is not limited, and a 500-page report processes in a few seconds because parsing runs locally.

Scanned documents contain no text layer, only a photograph of one, and return nothing. The tool detects this and says so explicitly, including the case where only some pages are scanned. Encrypted PDFs must be unlocked in a PDF reader first.

Why the document is never uploaded

Parsing runs through PDF.js inside your browser, so the file is read from local memory and never transmitted. Once the page has loaded, extraction works with the network disconnected.

Most online PDF-to-text converters upload the document, convert it server-side and return a link โ€” with your file resident on their infrastructure in the meantime. For a contract, a medical record, a payroll report or anything under NDA, local extraction removes that exposure entirely.

What extraction does not preserve

Plain text carries no formatting, so four things are lost by design:

Ligatures and unusual font encodings occasionally produce odd characters, which is a property of how the PDF embedded its fonts rather than of the extraction.

Why people extract PDF text

Text extraction compared with OCR

Use text extraction when the PDF has a text layer, and OCR only when it does not. Extraction reads character codes the file already contains, so it is exact, instant and lossless. OCR guesses characters from pixels, so it is slower and never perfect. Running OCR on a digital PDF makes the result worse, not better.

If this tool reports no text layer, the pages are images: use image to text OCR instead. To keep columns as columns, use the PDF table extractor. Once you have the text, the email and URL extractor pulls out the addresses and links inside it.

For a walkthrough of when a PDF needs OCR and when it already has a text layer, read scanned vs searchable PDF.

PDF text encoding and extraction edge cases

PDF text extraction has three encoding layers: the font's built-in encoding (WinAnsiEncoding, MacRomanEncoding or Identity-H for Unicode fonts), a ToUnicode CMap that maps glyph IDs to Unicode code points, and an ActualText attribute for ligature spans. When a PDF embeds a font with Identity-H encoding and a correct ToUnicode CMap, extraction is reliable. When neither is present โ€” common in PDFs from older InDesign versions and some Asian document workflows โ€” characters appear as hex codes rather than readable text.

Three extraction edge cases: (a) ligatures (fi, fl, ffl) are single glyphs with no 1:1 Unicode mapping unless ActualText is set โ€” they may appear as a single character, as missing or as an approximation; (b) right-to-left text (Arabic, Hebrew) is stored in visual order in most PDFs, not logical order โ€” the extractor returns characters in stream order, so RTL text can read reversed without post-processing; (c) Tagged PDF includes reading-order hints in its structure tree, which are used when present to improve column disambiguation and correct multi-column reading order.

Frequently asked questions

How do I extract text from a PDF for free?

Drop the PDF onto this page and the text appears in an editable box. Copy it or download it as a .txt file. No signup, no watermark and no page limit.

Is my PDF uploaded to a server?

No. The document is parsed by JavaScript in your browser and is never transmitted.

Why did my PDF return no text?

The document has no text layer, which means its pages are scanned images. Use image to text OCR to read them instead.

What is the difference between paragraph mode and exact-line mode?

Paragraph mode rejoins lines that were wrapped by the page layout, giving continuous prose. Exact-line mode keeps every line break as printed, which matters for addresses, code and verse.

Why is the text from a two-column page interleaved?

Lines are grouped by vertical position, so two columns sharing a baseline are read as one line. Extract one column at a time, or use a PDF reader's column-aware selection.

Can it extract text from a password-protected PDF?

No. Remove the password in a PDF reader first, then extract.

Is there a page limit?

No. Parsing happens locally, so long documents are limited only by your device's memory. Progress is shown while pages are read.

• Specialist file parsing & security engineer • Verified: in our experience, our hands-on testing measured and verified private in-browser execution with zero file uploads • Last reviewed September 2026.