How an EPUB decides its own reading order
An EPUB is a ZIP containing XHTML documents, stylesheets, images and a package file called the OPF. The OPF holds two lists that matter: the manifest, naming every file in the book, and the spine, listing the ids of those files in reading order.
The spine is the book. Filenames are arbitrary โ publishers use
chapter01.xhtml, part0004.html, or numbered fragments that bear no relation to
the sequence a reader sees. Extraction that sorts filenames alphabetically will put chapter 10 before
chapter 2 and place the copyright page in the middle. This tool follows
META-INF/container.xml to the OPF, then the spine, in order.
How to extract text from an EPUB
- Open the ebook. Drop the .epub onto the box above, or click to browse. The container is read locally and the spine is followed to establish reading order.
- Choose whole book or one section. The selector lists every section using its own first heading as a label, so chapters are easy to find.
- Decide on section labels. Labels insert a separator naming each section, which makes a long extraction navigable. Switch them off for continuous prose.
- Copy or download. Copy the text to the clipboard, or download it as a .txt file.
What comes out
- The full text of every spine section, in reading order, with paragraph breaks preserved.
- Section labels taken from each section's own first heading, so the output reads as a table of contents.
- Book metadata โ title, author, publisher, language and publication date, read from the OPF.
- Per-section extraction, so you can pull a single chapter without the rest of the book.
- Word and character counts for whatever you have selected.
Which ebooks work
EPUB 2 and EPUB 3 both work, including books from publishers, Project Gutenberg, Standard Ebooks and files produced by Calibre, Sigil and Pandoc. Malformed XHTML is handled, because sections are parsed as HTML rather than strict XML โ real-world EPUBs frequently contain markup that a strict parser rejects.
Amazon's .mobi, .azw and .azw3 formats are unrelated and are
detected and reported rather than misread. DRM-protected books are encrypted, and no browser tool will read
them โ that is the point of the encryption.
Why the ebook is never uploaded
The archive is read from your disk by the browser and decompressed locally. No server is involved, so the file is never transmitted and nothing remains once the tab closes.
Beyond the usual privacy argument, this one matters because uploading a copyrighted ebook to a third-party service is a distribution question you probably do not want to have. Extracting locally keeps the book on your machine.
What is not preserved, and what to be careful about
Plain text drops everything an EPUB uses to look like a book:
- Formatting. Italics, small caps, drop caps, fonts and styling are gone.
- Images. Covers, illustrations and diagrams are not extracted here.
- Footnotes and endnotes. These extract as text where they are stored, which may be far from their reference.
- Tables. They flatten into lines and lose their columns.
- Page numbers. EPUB is reflowable; it has no fixed pages to number.
On copyright: extracting text from a book you own, for your own reading, study or accessibility, is ordinary use. Redistributing that text is not. The tool does not and cannot remove DRM.
Why people extract ebook text
- Accessibility โ producing plain text a screen reader or text-to-speech tool handles reliably.
- Study and research โ quoting accurately, or searching a book your reader will not search well.
- Text analysis โ word frequency, readability or corpus work on a book you own.
- Format conversion โ getting clean text as the starting point for another format.
- Reading on a device with no ebook app โ a plain .txt file opens anywhere.
- Feeding a book to an AI tool โ most models cannot read .epub directly.
EPUB extraction compared with PDF extraction
EPUB is the easier of the two by a wide margin. It is structured HTML with an explicit reading order, so extraction is exact and chapter boundaries are real. A PDF stores glyphs at fixed coordinates, so the PDF text extractor has to reconstruct reading order from geometry and can interleave columns. If a book exists in both formats, extract the EPUB.
For a scanned book with no text layer, start with image to text OCR. To see the raw files inside the ebook โ cover images, fonts, stylesheets โ open it with the ZIP extractor, since an EPUB is a ZIP container. Read our complete guide on how to extract text and chapters from EPUB files to learn more about OPF package parsing and structure.
EPUB format standards and extraction edge cases
An EPUB is a ZIP archive containing an OPF package document (EPUB 2: an
.opf file; EPUB 3: a <package> element). The
<spine> element lists content documents in reading order โ the order in
which the extractor processes them. EPUB 3 uses XHTML5 as its content format; EPUB 2 uses
XHTML 1.1.
Three edge cases: (a) EPUB 3 allows inline MathML and SVG โ MathML is skipped and
<svg:text> content is not extracted, because SVG text layout does not map
cleanly to document flow; (b) fixed-layout EPUBs โ used for picture books and comics โ are
often entirely graphical with no text layer; the extractor returns empty output and reports
the absence of readable content rather than failing silently; (c) font files embedded as
WOFF2 resources and the META-INF directory are valid ZIP entries but are not content
documents โ the extractor processes only OPS content documents and skips all binary
media.
Frequently asked questions
How do I convert an EPUB to text?
Drop the .epub onto this page and the text appears in an editable box, in reading order. Copy it or download it as a .txt file.
Is my ebook uploaded to a server?
No. The file is read from your disk and unzipped by your browser. It is never transmitted.
Are the chapters in the right order?
Yes. Order comes from the EPUB's spine, which is the reading order recorded in the book itself, not from filenames.
Can I extract just one chapter?
Yes. The section selector lists every section by its own first heading; choose one and only that section is exported.
Does it work on Kindle files?
No. .mobi, .azw and .azw3 are Amazon formats built on a different structure. Convert to EPUB with Calibre first, if the file is DRM-free.
Can it remove DRM?
No. DRM-protected books are encrypted and cannot be read by this or any other browser tool.
Why is the text missing italics and images?
Plain text carries characters only. Formatting and images are part of the book's presentation, not its text.
Is there a size limit?
No fixed limit. Extraction is local, so the constraint is your device's memory. A full-length novel processes in under a second.