Why PDF Text Comes Out Garbled (and How to Fix It)
PDF text comes out garbled when the file’s characters are not mapped back to real letters — usually because the PDF uses a subset or custom-encoded font with no ToUnicode table, or because the “text” is actually a scanned image. The characters look fine on screen because the PDF is drawing glyph shapes, but the underlying codes don’t say which letters those shapes represent. Here’s what causes each case, and how to get clean text out.
This guide explains the three reasons PDF text turns into gibberish and what to do about each, so you can reliably extract text from a PDF.
A PDF draws shapes, not letters
A PDF page is a set of drawing instructions: “place glyph number 47 from this font here”. For you to copy that as text, the font has to also carry a mapping from each glyph code back to a Unicode character — the ToUnicode table. When that mapping is present and correct, copy and extraction give clean text. When it’s missing or wrong, you get the right shapes on screen but garbage when copied. That single fact explains most “why is my PDF text scrambled” problems.
Cause 1: Subset fonts with no ToUnicode map
To keep files small, PDF creators often subset a font — embedding only the glyphs actually used and renumbering them from zero. If the tool that made the PDF didn’t also write a ToUnicode table, glyph “1” might be the letter “e” visually but map to nothing meaningful in the character stream. Copying then yields symbols, boxes, or letters shifted by a fixed amount (a classic sign: every letter is off by one or two). The text is effectively locked to its visual form.
Fix: if the PDF has no usable text mapping, the reliable route is OCR — treat the page as an image and re-recognise the letters. See “Cause 3” below.
Cause 2: Wrong or non-standard encoding
Some PDFs use custom or legacy encodings — common with older documents, non-Latin scripts, or files exported by niche software. The characters may extract, but as the wrong letters: accented characters become question marks, ligatures like “fi” collapse or vanish, and quotation marks turn into stray symbols. This is an encoding problem rather than a missing map, and a good extractor handles the standard cases (proper Unicode, ligature expansion) automatically.
Fix: a modern browser-based extractor that reads the font’s encoding correctly will resolve most of these. What it can’t invent is a mapping that simply isn’t in the file — that falls back to OCR.
Cause 3: It’s a scanned page, not text at all
The most common case: the PDF is a scan — a photograph of a page. There is no text layer to extract, only an image. Copying selects nothing, or selects the whole page as one block that pastes as nothing. This isn’t garbling so much as an absence of text.
Fix: run the page through OCR to turn the picture of text into real, selectable text. The image to text (OCR) tool does this in your browser. (For how to tell a scanned PDF from a real-text one before you start, the difference is simple: if you can select a single word, it has a text layer; if selection grabs the whole page as an image, it’s scanned.)
How to get clean text out
- Try direct extraction first. Open the file in the PDF text extractor. If the PDF has a proper text layer and encoding, you get clean text immediately, in your browser, with nothing uploaded.
- If the output is scrambled or empty, the text is either unmapped (Cause 1) or the page is a scan (Cause 3). In both cases, switch to OCR, which reads the letters from the rendered image rather than the broken character stream.
- Check the result for the tell-tale signs above — letters shifted by a constant, missing ligatures, or stray symbols — which point to which cause you hit.
Frequently asked questions
Why does copied PDF text turn into gibberish?
Because the PDF’s font has no correct mapping from glyph codes back to Unicode letters (often a subset font with no ToUnicode table). The shapes display correctly but the underlying codes don’t say which letters they are.
Why is my extracted text shifted by one letter?
That’s a subset font whose glyphs were renumbered without a matching Unicode map. Every character maps to the wrong code by a fixed offset. OCR is the reliable fix.
Why can’t I select any text in my PDF?
The PDF is almost certainly a scan — an image of a page with no text layer. Use OCR to recognise the text from the image.
How do I fix garbled PDF text?
Try a proper PDF text extractor first. If the text is unmapped or the page is scanned, run it through OCR instead, which reads the letters from the rendered image.
Does a browser-based extractor keep my PDF private?
Yes. A browser-based tool reads the PDF on your own device and never uploads it, which matters for contracts, statements and other sensitive documents.
Related reading
Last updated: 16 August 2026.