The PDF Text Layer and Why It Matters for Accessibility
A PDF’s text layer is the real, selectable text stored behind the page image — and it’s what makes a PDF both extractable and accessible. The same layer that lets you copy or extract the text is the one a screen reader speaks aloud. A scanned PDF has no text layer, so it fails both at once: you can’t extract it, and assistive technology can’t read it. The two problems are the same problem, which is why the fix is the same too.
This guide explains what the text layer is, how document tagging builds accessibility on top of it, and how to check whether a PDF has one using the PDF text extractor.
What the text layer is
A PDF page can carry two things: what you see (glyphs and images drawn on the page) and the text those glyphs represent, stored as real characters tied to a font. That second thing is the text layer. When it’s present, you can select words, search the document, and extract the text cleanly. When it’s absent — as in a scan — the page is only a picture, and there’s nothing to select, search, extract, or read aloud.
Extraction and accessibility are the same capability
It’s worth stating plainly because it’s so useful: if you can extract a PDF’s text, a screen reader can read it; if you can’t, it can’t. Both depend on the text layer being present and correctly encoded. This means the humble “can I select a word?” test tells you about accessibility as much as extractability. A document that returns clean text in an extractor is one a blind or low-vision reader can also use.
Tagging: accessibility built on top of the text layer
A text layer alone makes a PDF readable. To make it properly navigable, PDFs add a tag tree — structure markup that labels each part as a heading, paragraph, list, table cell, or image, and records the correct reading order. A tagged PDF lets a screen reader announce “heading level 2”, skip between sections, read a table by rows, and describe an image via its alt text. This is the basis of the PDF/UA accessibility standard.
Tagging sits on top of the text layer: without the text layer there’s nothing to tag, and with the text layer but no tags, the document is readable but hard to navigate.
Why scanned PDFs fail accessibility
A scan is an image of a page. To a screen reader it’s a blank — there are no words, headings or reading order, just pixels. This is one of the most common accessibility failures in real documents: a form or report distributed as a scan is completely opaque to assistive technology. The remedy is the same as for extraction: run OCR to create a text layer, then (ideally) tag the result.
How to check a PDF
- Text layer present? Open the file in the PDF text extractor. Clean text out means a real text layer — good for both extraction and reading aloud. Nothing, or gibberish, means it’s a scan or the text is unmapped (see why PDF text comes out garbled).
- Scanned? If there’s no text layer, add one with OCR before relying on the document for either purpose.
- Tagged? A PDF reader’s accessibility checker reports whether the file is tagged and flags missing alt text and reading-order issues.
Frequently asked questions
What is the text layer in a PDF?
The real, selectable text stored behind the visible page. It’s what lets you copy, search and extract the text — and what a screen reader reads aloud.
Why can’t a screen reader read my PDF?
Almost always because the PDF is a scan with no text layer — an image of a page. Running OCR adds a text layer so assistive technology (and text extraction) can work.
Is an extractable PDF also accessible?
It’s readable: if the text extracts cleanly, a screen reader can read it. Full accessibility also needs tags (a tagged PDF) for correct headings, reading order and alt text.
What is a tagged PDF?
A PDF with a structure tree labelling headings, paragraphs, lists, tables and images, plus reading order. It’s what lets a screen reader navigate the document, and the basis of the PDF/UA standard.
How do I check whether my PDF has a text layer?
Try to select a word, or open it in a PDF text extractor. Clean text means a text layer is present; nothing means it’s a scan.
Related reading
Last updated: 16 August 2026.