Why some PDFs produce garbled text
PDF 1.5 introduced object streams: a way to bundle many objects into a single compressed chunk. This reduces file size, but it means older readers — and many extraction tools — must decompress an entire chunk before they can read any single object inside it. When they cannot do that decompression, they read raw bytes and produce garbled characters, empty pages, or errors instead of text.
The fix is to re-save the PDF without object streams, which PDF tools call setting compatibility to PDF 1.4. Every object is then stored individually, uncompressed, and every reader can find and decode them directly. File size typically increases slightly, but compatibility becomes universal.
For a full explanation of why PDF text comes out garbled and the other causes behind it, read why PDF text comes out garbled.
How to repair a PDF
- Drop the PDF in. Drop the broken PDF onto the box above, or click to browse. The file is read inside your browser — nothing is sent to any server.
- Wait for repair. The PDF is loaded, its structure normalised, and the compressed object streams removed. For most files this takes under a second.
- Download the fixed file. Click Download repaired PDF to save the result. The repaired file is compatible with PDF 1.4 readers, Acrobat, every text extractor, and all print drivers.
- Try extraction again. Once you have the repaired PDF, run it through the PDF text extractor or PDF table extractor — the output should now be clean.
What the repaired PDF contains
- All pages, in the original order, with no content removed.
- All fonts, embedded exactly as stored in the original — text rendering is unchanged.
- All images, at original resolution and quality.
- All annotations, form fields, links and bookmarks.
- Standard cross-reference tables instead of compressed cross-reference streams, so the structure is readable by every conforming PDF reader.
What changes is only the internal storage format: object streams and cross-reference streams are expanded out. The visual appearance of the document is identical to the original.
What this tool can and cannot fix
Fixes:
- PDFs that extract as garbled or reversed characters because of compressed object streams
- PDFs that open blank or partially in older readers (Acrobat 5, PDF.js with limited support)
- PDFs that fail in extraction tools with a "cross-reference stream" or "object stream" error
- PDFs produced by LaTeX, InDesign, or other tools with PDF 1.5+ compression that needs downgrading
Does not fix:
- Font encoding issues. If the PDF uses a custom character map that maps standard codepoints to different glyphs, the text layer itself is garbled — no re-save fixes that.
- Scanned-only PDFs. A scan with no text layer contains no text to extract; use the image to text (OCR) tool instead.
- Corrupted files. A PDF truncated mid-download or with a damaged cross-reference table may not load at all — use Acrobat's native repair first.
- Password-protected PDFs. Encryption must be removed before repair.
Why the PDF is never uploaded
The PDF is loaded from your disk by your browser and processed by pdf-lib, a pure-JavaScript PDF library. No network request is made during the repair, and the fixed file is written directly to your downloads folder. No copy exists anywhere else after the tab closes.
PDFs sent for repair often contain exactly the kind of content that should not travel to a stranger’s server — contracts, invoices, medical records, signed agreements. Here there is nothing to trust a third party with, because there is no third party.
Limits
- File size. PDFs up to ~150 MB process in-browser. Larger files may exhaust the tab’s memory; split the PDF first with the PDF page splitter and repair each part.
- Encryption. Password-protected PDFs cannot be loaded. Remove the password in Acrobat or with a desktop tool, then repair.
- Digital signatures. Re-saving a signed PDF invalidates the signature, because the bytes of the file change. This is expected and unavoidable.
- DRM and copy protection. Some PDFs use rights-management restrictions beyond simple password encryption; these may prevent loading.
When to use PDF repair
- Garbled text from a PDF extractor — characters come out as boxes, reversed sequences, or meaningless strings despite the PDF looking fine on screen.
- PDF that will not open in an older system — a legacy reader, a printer, a document management system that requires PDF 1.4 compatibility.
- Extraction tool rejecting the file — the tool reports an "xref stream" or "object stream" error and refuses to process the PDF.
- Template PDFs from design tools — InDesign, Illustrator, and LaTeX all output PDF 1.5+ by default; a repair brings them down to universal compatibility.
- Before archiving — long-term archives (PDF/A) require no object streams; repair normalises the structure before a PDF/A validator runs.
PDF repair compared with other PDF fixes
This tool fixes structural compression issues. Other PDF problems need different tools:
- Scanned PDF with no text — use image to text OCR to create a text layer.
- PDF with the right structure but wrong page order — use the PDF page extractor to pull out the pages you need.
- Large PDF you want to split — use the PDF page splitter to divide it before repairing.
- PDF with hidden metadata — use the PDF metadata extractor to read what is embedded, and note that repair preserves all metadata fields.
If repair does not fix the garbled text, the cause is a custom font encoding rather than object streams. Read why PDF text comes out garbled for the full decision tree.
Frequently asked questions
Why does my PDF show garbled text when I extract it?
The most common cause is compressed object streams (a PDF 1.5+ feature). Extraction tools that cannot decompress them read raw bytes and output garbled characters. Drop the PDF here to re-save it without those streams — extraction should produce clean text after that.
Will repairing the PDF change how it looks?
No. The visual content — text, images, layout, fonts — is preserved exactly. Only the internal storage format changes: objects are stored individually instead of bundled in compressed streams.
Is the PDF uploaded to a server?
No. The file is processed entirely in your browser by pdf-lib, a JavaScript library. Nothing is sent over the network.
My PDF is password-protected. Can I repair it?
Not directly — encryption prevents loading. Remove the password in Acrobat or a desktop tool first, then repair.
The repair did not fix the garbled text. What else could cause it?
A custom font encoding. Some PDFs use a character map that assigns non-standard meaning to standard codepoints — the text layer itself records the wrong characters. Re-saving cannot fix this; only the original application can output correct text. Read the guide on why PDF text comes out garbled for the full diagnosis.
Will repairing a signed PDF break the signature?
Yes. Re-saving any PDF changes its bytes, which invalidates embedded digital signatures. If the signature matters, keep the original; if only the content matters, repair is fine.
How big a PDF can I repair?
Files up to roughly 150 MB process in-browser without issue. For very large PDFs, split the file first with the PDF page splitter, then repair each part individually.
Can I use this to make a PDF compatible with PDF 1.4?
Yes. Removing object streams is the main thing that differentiates a PDF 1.4-compatible file from a PDF 1.5+ one. After repair, your PDF will open in any reader that supports PDF 1.4, which covers everything from Acrobat 5 onward.